The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
58,216 characters
Testing Hypotheses About Ratios of Linear Trend Slopes in Systems of Equations with a Focus on Tests of Equal Trend Ratios
\title{Testing Hypotheses About Ratios of Linear Trend Slopes in Systems of Equations
with a Focus on Tests of Equal Trend Ratios}
\author{Timothy J. Vogelsang\thanks{Tim Vogelsang, Department of Economics, 486 W.
Circle Drive, 110 Marshall-Adams Hall, Michigan State University East Lansing,
MI 48824-1038. Phone: 517-353-4582, Fax: 517--432-1068, email: [email removed]}}
\date{February 26, 2026}
\maketitle
\begin{abstract}
This paper develops inference methods for ratios of deterministic trend slopes
in systems of pairs of time series. Hypotheses based on linear cross-equation
restrictions are considered with particular interest in tests that trend
ratios are equal across pairs of trending series. Tests of equal ratios can be
used for the empirical assessment of climate models through comparisons of
trend ratios (amplification ratios) of model generated temperature series and
observed temperature series. The analysis in this paper builds on the
estimation and inference methods developed by \nocite{vogelsang-nawaz-JTSA}
Vogelsang and Nawaz (2017, \textit{Journal of Time Series Analysis}) for a
single pair of trending time series. Because estimators of ratios can have
poor finite sample properties when the trend slope are small relative to
variation around the trends, tests of equal trend ratios are restated in terms
of products of trend slopes leading to inference that is less affected by
small trend slopes. Asymptotic theory is developed that can be used to
generate critical values. For tests of equal trend ratios, finite sample
performance is assessed using simulations. Practical advice is provided for
empirical practitioners. An empirical application compares amplification
ratios (trend ratios) across a set of five groups of observed global
temperature series.
\end{abstract}
Keywords: Trend Stationary, Instrumental Variables Estimation, HAC Estimator,
Fixed-b Asymptotics, Amplification Ratio, Long Run Variance\bigskip
\setcounter{page}{0} \thispagestyle{empty} \baselineskip=16.0pt \pagestyle
{plain}
\section{Introduction}
This paper analyzes estimation and inference methods for parameters that
represent ratios of pairs of linear trend slopes in a system of linear
trending time series with covariance stationary fluctuations around the trend.
The proposed methods extend the results of \cite{vogelsang-nawaz-JTSA} where
the focus was on a single trend ratio parameter, and it was shown that
instrumental variables (IV) estimation using time as an instrument was
preferred to ordinary least squares (OLS) of the relevant estimating equation.
Here, there is a vector of trend ratio parameters estimated by IV. Inference
focuses on hypotheses that represent linear restrictions across trend ratio
parameters. Much of the intuition in \cite{vogelsang-nawaz-JTSA} for the
single pair case extends to the multi-pair case especially with respect to the
importance of the magnitude of trend slopes relative to noise
(variation/fluctuations) around the trends for estimation and inference.
Throughout the paper, magnitudes of trends slopes are always interpreted
relative to the magnitude of the noise.
Detailed attention is given to the special case where an empirical
practitioner wants to test the equality of trend ratio parameters between two
pairs of series. This special case is directly relevant for the study of
amplification ratios in the empirical climate literature;
\nocite{santer2005amplification}Santer \textit{et al} (2005),
\nocite{thorne2007tropical}Thorne \textit{et al} (2007),
\nocite{klotzbach2009alternative}\nocite{klotzbach2010correction}Klotzbach
\textit{et al} (2009, 2010), \nocite{christy2010observational}Christy
\textit{et al} (2010), \cite{po2012discrepancies}, \cite{vogelsang-nawaz-JTSA}
, \cite{VMCS_2026} (hereafter VMCS26), and references in those papers.
The amplification ratio is the ratio of a temperature trend in the troposphere
of the earth relative to a temperature trend at the surface of the earth. A
key assessment of theoretical climate models is determining whether estimated
amplification ratios of observed temperature series are aligned with estimated
amplification ratios of model generated temperatures. The null hypothesis of
interest is equality of trend ratios between two pairs of temperature series:
one pair for observed temperatures and a second pair for model generated
temperatures. Whereas the previous literature treats the amplification ratio
computed for model generated temperatures as fixed (ignoring that it is
estimated), the methods developed in this paper treat both the observed and
model amplification ratios as estimators. Inference takes into account the
joint sampling distributions of both estimated amplification ratios. VMCS26
use these methods to assess the alignment of amplification ratios in CMIP6
model generated temperature series with amplification ratios in observed
temperature series. VMCS26 find systematic misalignments especially for
amplification ratios in the lower troposphere in which case the models exhibit
substantially higher amplification (more warming relative to the surface) than
in observed temperatures. In the present paper amplification ratios are
compared across five sets of observed temperatures.
The remainder of the paper is organized as follows: Section 2 lays out the
system of estimating equations used to estimate trend ratios of pairs of
linear trending time series. The asymptotic properties of the IV estimator of
the trend slope ratios is provided. As shown by \cite{vogelsang-nawaz-JTSA},
IV estimation is used rather than OLS because of systematic correlation
between the regressors and regression errors. Section 3 provides a framework
for testing linear restrictions of trend ratios across equations in the
system. Test statistics are configured to be robust to serial correlation in
the fluctuations of the time series around their trends as well as correlation
across time series. Tests of equal trend ratios between pairs is obtained as a
special case. When the trend slopes are small relative to the variation of the
time series around their trend, the IV estimators, and tests built on them,
can behave very differently than when trend slopes are large. To help offset
this sensitivity to the magnitude of the trend slopes, Section 4 explores an
alternative approach to inference for tests of equal trends that is labeled
the "product approach". The idea is related to the linear-in-trend slopes
inference method (\nocite{fieller1954}Fieller 1954) used by
\cite{vogelsang-nawaz-JTSA} for a single trend ratio. Unlike the linear in
trend slopes approach, the product approach is not fully robust to very small,
or even zero, trend slopes. However, as the finite sample simulations show in
Section 5, the product approach for testing equal trend slopes is often less
sensitive to very small trend slopes relative to tests based on the IV
estimators. The simulations suggest that while the product approach can be
more robust to very small trend slopes, IV based tests tend to be less
sensitive to strong autocorrelation (tendency to over-reject under the null
hypothesis is less). Section 6 provides some practical recommendations for
empirical researchers. Section 7 uses the proposed tests of equal trend ratios
to compare amplification ratios across five sets of observed temperatures for
the tropics of the earth. Section 8 concludes, and proofs are provided in an appendix.
\section{The Model and Estimation}
\subsection{Statistical Model and Assumptions}
The setup consists of $i=1,2,\ldots,n$ \textit{pairs} of univariate linear
trending time series, $y_{1t}^{(i)}$ and $y_{2t}^{(i)}$, given by
\begin{equation}
y_{1t}^{(i)}=\mu_{1}^{(i)}+\beta_{1}^{(i)}t+u_{1t}^{(i)}, \label{1.1}
\end{equation}
\begin{equation}
y_{2t}^{(i)}=\mu_{2}^{(i)}+\beta_{2}^{(i)}t+u_{2t}^{(i)}, \label{1.2}
\end{equation}
where $u_{1t}^{(i)}$ and $u_{2t}^{(i)}$ are mean zero covariance stationary
processes and $t=1,2,\ldots,T$. The parameters of interest are the $n$ ratios
of trend slopes between each pairs of series given by
\[
\theta^{(i)}=\frac{\beta_{1}^{(i)}}{\beta_{2}^{(i)}},
\]
where $\beta_{2}^{(i)}\neq0$. Using simple algebra from
\cite{vogelsang-nawaz-JTSA}, estimating equations for the ratios can be
derived as
\begin{equation}
y_{1t}^{(i)}=\delta^{(i)}+\theta^{(i)}y_{2t}^{(i)}+\epsilon_{\theta t}^{(i)},
\label{y1y2reg}
\end{equation}
where
\[
\delta^{(i)}=\mu_{1}^{(i)}-\theta^{(i)}\mu_{2}^{(i)},\text{ \ \ \ \ }
\epsilon_{\theta t}^{(i)}=u_{1t}^{(i)}-\theta^{(i)}u_{2t}^{(i)}.
\]
Throughout the paper, the time series are assumed to be covariance stationary
around their respective linear trends and that sufficient weak dependence
holds so that a functional central limit theorem (FCLT) holds for the
$2n\times1$ vector:
\[
\mathbf{U}_{t}=\left[
\begin{array}
[c]{c}
\mathbf{U}_{1t}\\
\mathbf{U}_{2t}
\end{array}
\right] ,
\]
where $\mathbf{U}_{1t}=\left[ u_{1t}^{(1)},u_{1t}^{(2)},\ldots,u_{1t}
^{(n)}\right] ^{\prime}$ and $\mathbf{U}_{2t}=\left[ u_{2t}^{(1)}
,u_{2t}^{(2)},\ldots,u_{2t}^{(n)}\right] ^{\prime}$. Specifically, it is
assumed that
\begin{equation}
T^{-1/2}
{\displaystyle\sum_{t=1}^{[rT]}}
\mathbf{U}_{t}=T^{-1/2}
{\displaystyle\sum_{t=1}^{[rT]}}
\left[
\begin{array}
[c]{c}
\mathbf{U}_{1t}\\
\mathbf{U}_{2t}
\end{array}
\right] \Rightarrow\left[
\begin{array}
[c]{c}
\mathbf{B}_{\mathbf{u}1}(r)\\
\mathbf{B}_{\mathbf{u}2}(r)
\end{array}
\right] \equiv\mathbf{B}_{\mathbf{u}}\mathbf{(r),} \label{fclt}
\end{equation}
where $r\in\lbrack0,1]$ and $[rT]$ is the integer part of $rT$. The elements
of the $n\times1$ vectors of Brownian motions, $\mathbf{B}_{\mathbf{u}1}(r)$
and $\mathbf{B}_{\mathbf{u}2}(r)$, are given by
\[
\mathbf{B}_{\mathbf{u}1}(r)=\left[ B_{u1}^{(1)}(r),B_{u1}^{(2)}
(r)\ldots,B_{u1}^{(n)}(r)\right] ^{\prime},\text{ \ \ \ \ }\mathbf{B}
_{\mathbf{u}2}(r)=\left[ B_{u2}^{(1)}(r),B_{u2}^{(2)}(r)\ldots,B_{u2}
^{(n)}(r)\right] ^{\prime}.
\]
The vector of Brownian motions, $\mathbf{B}_{\mathbf{u}}\mathbf{(r)}$, can be
written as $\mathbf{\Lambda}_{\mathbf{u}}\mathbf{W}_{\mathbf{u}}(r)$ where
$\mathbf{W}_{\mathbf{u}}(r)$ is a $2n\times1$ vector of independent standard
Wiener processes and $\mathbf{\Omega}_{\mathbf{u}}=\mathbf{\Lambda
}_{\mathbf{u}}\mathbf{\Lambda}_{\mathbf{u}}^{\prime}$ is the long run variance
of $\mathbf{U}_{t}$. It is\textit{\ not} assumed that $\mathbf{\Omega
}_{\mathbf{u}}$ is diagonal allowing for correlation across elements
of\ $\mathbf{U}_{t}$ (within and across pairs). In addition to (\ref{fclt}),
it is assumed that $\mathbf{U}_{t}$ is ergodic for the first and second moments.
Stacking the individual estimation equation errors, $\epsilon_{\theta t}
^{(i)}$, gives the $n\times1$ vector, $\scalebox{1.5}{$\bm{\epsilon}$}_{\theta
t}$, that can be written as
\[
\scalebox{1.5}{$\bm{\epsilon}$}_{\theta t}=\mathbf{U}_{1t}-\mathbf{D}
_{\mathbf{\theta}}\mathbf{U}_{2t},
\]
where
\[
\mathbf{\theta=}\left[ \theta^{(1)},\theta^{(2)},\ldots,\theta^{(n)}\right]
^{\prime},
\]
and $\mathbf{D}_{\mathbf{\theta}}$ is an $n\times n$ diagonal matrix with
$i^{th}$ diagonal elements $\theta^{(i)}$. Notice that
$\scalebox{1.5}{$\bm{\epsilon}$}_{\theta t}$ can be written as a linear
function of $\mathbf{U}_{t}$ through the relationship
\[
\scalebox{1.5}{$\bm{\epsilon}$}_{\theta t}=\left[ \mathbf{I}_{n}
,-\mathbf{D}_{\mathbf{\theta}}\right] \mathbf{U}_{t},
\]
where $\mathbf{I}_{n}$ is an $n\times n$ identity matrix. It immediately
follows from (\ref{fclt}) that
\begin{equation}
T^{-1/2}
{\displaystyle\sum_{t=1}^{[rT]}}
\scalebox{1.5}{$\bm{\epsilon}$}_{\theta t}\Rightarrow\left[ \mathbf{I}
_{n},-\mathbf{D}_{\mathbf{\theta}}\right] \mathbf{B}_{\mathbf{u}
}\mathbf{(r)\sim\Lambda}_{\mathbf{\epsilon}}\mathbf{W}_{\mathbf{\epsilon}}(r),
\label{fclteps}
\end{equation}
where $\mathbf{W}_{\mathbf{\epsilon}}(r)$ is an $n\times1$ vector of
independent Wiener processes and
\[
\mathbf{\Omega}_{\mathbf{\epsilon}}=\mathbf{\Lambda}_{\mathbf{\epsilon}
}\mathbf{\Lambda}_{\mathbf{\epsilon}}^{\prime}=\left[ \mathbf{I}
_{n},-\mathbf{D}_{\mathbf{\theta}}\right] \mathbf{\Lambda}_{\mathbf{u}
}\mathbf{\Lambda}_{\mathbf{u}}^{\prime}\left[ \mathbf{I}_{n},-\mathbf{D}
_{\mathbf{\theta}}\right] ^{\prime},
\]
is the long run variance of $\scalebox{1.5}{$\bm{\epsilon}$}_{\theta t}$.
\subsection{Estimation of the Trend Slope Ratios}
The trend slope ratios $\theta^{(i)}$ can be estimated using the estimating
equations (\ref{y1y2reg}). While ordinary least squares (OLS) applied equation
by equation for each $i$ would seem natural, \cite{vogelsang-nawaz-JTSA}
showed that OLS applied to (\ref{y1y2reg}), for a given $i$, yields a biased
estimator of $\theta^{(i)}$. The bias is caused by correlation between
$y_{2t}^{(i)}$ and $\epsilon_{\theta t}^{(i)}$ through the common term
$u_{2t}^{(i)}$. Instead, \cite{vogelsang-nawaz-JTSA} recommend using
instrumental variables (IV) estimation using $t$ as the instrument for
$y_{2t}^{(i)}$. The IV estimators are defined as
\begin{equation}
\widehat{\theta}^{(i)}=\left(
{\displaystyle\sum\limits_{t=1}^{T}}
(t-\overline{t})(y_{2t}^{(i)}-\overline{y}_{2}^{(i)})\right) ^{-1}
{\displaystyle\sum\limits_{t=1}^{T}}
(t-\overline{t})(y_{1t}^{(i)}-\overline{y}_{1}^{(i)}), \label{iv}
\end{equation}
where $\overline{y}_{1}^{(i)}=T^{-1}\sum_{t=1}^{T}y_{1t}^{(i)}$, $\overline
{y}_{2}^{(i)}=T^{-1}\sum_{t=1}^{T}y_{2t}^{(i)}$ and $\overline{t}=T^{-1}
\sum_{t=1}^{T}t$ are sample averages. Standard algebra gives the relationship
\[
\widehat{\theta}^{(i)}-\theta^{(i)}=\left(
{\displaystyle\sum\limits_{t=1}^{T}}
(t-\overline{t})(y_{2t}^{(i)}-\overline{y}_{2}^{(i)})\right) ^{-1}
{\displaystyle\sum\limits_{t=1}^{T}}
(t-\overline{t})\epsilon_{\theta t}^{(i)}.
\]
Notice that the IV estimator can be equivalently written as
\[
\widehat{\theta}^{(i)}=\frac{\widehat{\beta}_{1}^{(i)}}{\widehat{\beta}
_{2}^{(i)}},
\]
where $\widehat{\beta}_{1}^{(i)}$ and $\widehat{\beta}_{2}^{(i)}$ are the OLS
estimators of $\beta_{1}^{(i)}$ and $\beta_{2}^{(i)}$ based on regressions
(\ref{1.1}) and (\ref{1.2}):
\begin{equation}
\widehat{\beta}_{1}^{(i)}=\left( \sum_{t=1}^{T}(t-\overline{t})^{2}\right)
^{-1}\sum_{t=1}^{T}(t-\overline{t})(y_{1t}^{(i)}-\overline{y}_{1}^{(i)}),
\label{beta1ols}
\end{equation}
\begin{equation}
\widehat{\beta}_{2}^{(i)}=\left( \sum_{t=1}^{T}(t-\overline{t})^{2}\right)
^{-1}\sum_{t=1}^{T}(t-\overline{t})(y_{2t}^{(i)}-\overline{y}_{2}^{(i)}).
\label{beta2ols}
\end{equation}
Stack the $\widehat{\theta}^{(i)}$ into the vector
\[
\widehat{\mathbf{\theta}}\mathbf{=}\left[ \widehat{\theta}^{(1)}
,\widehat{\theta}^{(2)},\ldots,\widehat{\theta}^{(n)}\right] ^{\prime}.
\]
Using (\ref{iv}), $\widehat{\mathbf{\theta}}$ can be written as
\[
\widehat{\mathbf{\theta}}\text{ }\mathbf{=}\text{ }\widehat{\mathbf{D}}
_{2}^{-1}\sum_{t=1}^{T}\left( \mathbf{y}_{1t}-\overline{\mathbf{y}}
_{1}\right) \left( t-\overline{t}\right) ,
\]
where $\mathbf{y}_{1t}=\left[ y_{1t}^{(1)},y_{1t}^{(2)},\ldots,y_{1t}
^{(n)}\right] ^{\prime}$ and $\widehat{\mathbf{D}}_{2}$ is an $n\times n$
diagonal matrix with $i^{th}$ diagonal element given by $\sum_{t=1}^{T}\left(
t-\overline{t}\right) \left( y_{2t}^{(i)}-\overline{y}_{2}^{(i)}\right) $.
Standard calculations give
\begin{equation}
\widehat{\mathbf{\theta}}-\mathbf{\theta=}\text{ }\widehat{\mathbf{D}}
_{2}^{-1}\sum_{t=1}^{T}\left( t-\overline{t}\right)
\scalebox{1.5}{$\bm{\epsilon}$}_{\theta t}, \label{thetahat_centered}
\end{equation}
which holds as long as the trend slopes are nonzero. When trends slopes are
zero, $\mathbf{\theta}$ is not defined and it follows that
\begin{equation}
\widehat{\theta}^{(i)}=\frac{
{\textstyle\sum\nolimits_{t=1}^{T}}
(t-\overline{t})(u_{1t}^{(i)}-\overline{u}_{1}^{(i)})}{
{\textstyle\sum\nolimits_{t=1}^{T}}
(t-\overline{t})(u_{2t}^{(i)}-\overline{u}_{2}^{(i)})}=\frac{
{\textstyle\sum\nolimits_{t=1}^{T}}
(t-\overline{t})u_{1t}^{(i)}}{
{\textstyle\sum\nolimits_{t=1}^{T}}
(t-\overline{t})u_{2t}^{(i)}}. \label{thetahat_zero_slopes}
\end{equation}
For the rest of the paper, the focus is on the IV estimator,
$\widehat{\mathbf{\theta}}$, and tests of linear restrictions regarding
$\mathbf{\theta}$.
\subsection{Asymptotic Properties of the IV Estimator}
\noindent The following Theorem gives the asymptotic properties of
$\widehat{\mathbf{\theta}}$ under the assumption that the FCLT (\ref{fclt})
holds. The asymptotic limit of $\widehat{\mathbf{\theta}}$ depends on the
magnitude of the trend slope parameters, $\beta_{1}^{(i)},\beta_{2}^{(i)}$
relative to the variation in the random components (noise), $u_{1t}^{(i)}$ and
$u_{2t}^{(i)}$.
\begin{theorem}
Suppose that (\ref{fclt}) holds which implies that (\ref{fclteps}) holds. Let
$\overline{\beta}_{1}^{(i)},\overline{\beta}_{2}^{(i)}$ be fixed with respect
to $T$. Let $\mathbf{D}_{\overline{\mathbf{\beta}}_{2}}$ be an $n\times n$
diagonal matrix with $i^{th}$ diagonal element $\overline{\beta}_{2}^{(i)}$,
and let $\mathbf{D}_{\mathbf{B}_{u2}}$be an $n\times n$ diagonal matrix with
$i^{th}$ diagonal element $\int_{0}^{1}\left( s-\frac{1}{2}\right)
dB_{u2}^{(i)}(s)$. The following hold as $T\rightarrow\infty$:
\[
\]
Case 1 (large to small slopes): For $\beta_{1}^{(i)}=T^{-\kappa}
\overline{\beta}_{1}^{(i)}$, $\beta_{2}^{(i)}=T^{-\kappa}\overline{\beta}
_{2}^{(i)}$ with $0\leq\kappa<\frac{3}{2}$,
\[
T^{3/2-\kappa}\left( \widehat{\mathbf{\theta}}-\mathbf{\theta}\right)
\Rightarrow\left( \frac{1}{12}\mathbf{D}_{\overline{\mathbf{\beta}}_{2}
}\right) ^{-1}\mathbf{\Lambda}_{\mathbf{\epsilon}}\int_{0}^{1}\left(
s-\frac{1}{2}\right) d\mathbf{W}_{\mathbf{\epsilon}}(s)\sim N\left(
\mathbf{0},12\mathbf{D}_{\overline{\mathbf{\beta}}_{2}}^{-1}\mathbf{\Omega
}_{\mathbf{\epsilon}}\mathbf{D}_{\overline{\mathbf{\beta}}_{2}}^{-1}\right)
,
\]
Case 2 (very small slopes): For $\beta_{1}^{(i)}=T^{-3/2}\overline{\beta}
_{1}^{(i)},$ $\beta_{2}^{(i)}=T^{-3/2}\overline{\beta}_{2}^{(i)}$
($\kappa=\frac{3}{2}$),
\[
T^{3/2}\left( \widehat{\mathbf{\theta}}-\mathbf{\theta}\right) =\left(
\widehat{\mathbf{\theta}}-\mathbf{\theta}\right) \Rightarrow\left( \frac
{1}{12}\mathbf{D}_{\overline{\mathbf{\beta}}_{2}}+\mathbf{D}_{\mathbf{B}_{u2}
}\right) ^{-1}\mathbf{\Lambda}_{\mathbf{\epsilon}}\int_{0}^{1}\left(
s-\frac{1}{2}\right) d\mathbf{W}_{\mathbf{\epsilon}}(s).
\]
Case 3 (zero slopes): For $\beta_{1}^{(i)}=0$, $\beta_{2}^{(i)}=0$,
\[
\widehat{\mathbf{\theta}}\Rightarrow\mathbf{D}_{\mathbf{B}_{u2}}^{-1}\int
_{0}^{1}\left( s-\frac{1}{2}\right) d\mathbf{B}_{\mathbf{u}1}(s).
\]
\end{theorem}
For Cases 1 and 2 the limits given in Theorem 1 are multivariate versions of
the limits obtained by \cite{vogelsang-nawaz-JTSA} and are identical to the
limits in \cite{vogelsang-nawaz-JTSA} when $n=1$. In Case 1,
$\widehat{\mathbf{\theta}}$ consistently estimates $\mathbf{\theta}$ and the
precision of $\widehat{\mathbf{\theta}}$ depends on the magnitudes of the
trends slopes in the denominator series. In Case 2, $\widehat{\mathbf{\theta}
}$ becomes inconsistent. This is not surprising because the case of very small
slopes means that the trend component of $y_{2t}^{(i)}$ is dominated by the
noise, $u_{2t}^{(i)}$, in which case $t$ is a weak instrument (Staiger and
Stock 1997\nocite{staiger-stock}) for $y_{2t}^{(i)}$. When all trend slopes
are zero (Case 3), trend ratios are not defined and $\widehat{\mathbf{\theta}
}$ converges to a random vector that depends on $\mathbf{B}_{\mathbf{u}
}\mathbf{(r)}$. It is obvious that for very small and zero slopes, inference
will be affected by the different behavior of $\widehat{\mathbf{\theta}}$.
\section{Testing Linear Restrictions Across Trend Ratios}
\noindent Suppose one is interested in testing linear restrictions across the
trend slopes, $\theta^{(i)}$, using the IV estimators $\widehat{\theta}^{(i)}
$. More formally, consider testing the null hypothesis
\[
H_{0}:\mathbf{R\theta=r,}
\]
against the alternative
\[
H_{1}:\mathbf{R\theta\neq r,}
\]
where $\mathbf{R}$ is a known $q\times n$ matrix with $rank(\mathbf{R)=}q$ and
$\mathbf{r}$ is a known $q\times1$ vector.
Using the Case 1 limit from Theorem 1, the Wald statistic for testing $H_{0}$
against $H_{1}$ is given by
\[
Wald_{IV}=\left( \mathbf{R}\widehat{\mathbf{\theta}}-\mathbf{r}\right)
^{\prime}\left[ \mathbf{R}\widehat{\mathbf{V}}_{IV}\mathbf{R}^{\prime
}\right] ^{-1}\left( \mathbf{R}\widehat{\mathbf{\theta}}-\mathbf{r}\right)
\]
where
\[
\widehat{\mathbf{V}}_{IV}=\left( \sum\nolimits_{t=1}^{T}(t-\overline{t}
)^{2}\right) \widehat{\mathbf{D}}_{2}^{-1}\widehat{\mathbf{\Omega}
}_{\mathbf{\epsilon}}\widehat{\mathbf{D}}_{2}^{-1}.
\]
The middle term of $\widehat{\mathbf{V}}_{IV}$ is an estimator of
$\mathbf{\Omega}_{\mathbf{\epsilon}}$ given by
\[
\widehat{\mathbf{\Omega}}_{\mathbf{\epsilon}}=\widehat{\mathbf{\Gamma}
}_{\mathbf{\epsilon}0}+
{\displaystyle\sum\limits_{j=1}^{T-1}}
k\left( \frac{j}{M}\right) \left( \widehat{\mathbf{\Gamma}}
_{\mathbf{\epsilon}j}+\widehat{\mathbf{\Gamma}}_{\mathbf{\epsilon}j}^{\prime
}\right) \text{, \ \ }\widehat{\mathbf{\Gamma}}_{\mathbf{\epsilon}j}=T^{-1}
{\displaystyle\sum_{t=j+1}^{T}}
\widehat{\scalebox{1.5}{$\bm{\epsilon}$}}_{\theta t}
\widehat{\scalebox{1.5}{$\bm{\epsilon}$}}_{\theta t-j}^{\prime},
\]
where $\widehat{\scalebox{1.5}{$\bm{\epsilon}$}}_{\theta t}$ is the vector of
IV residuals given by
\[
\widehat{\scalebox{1.5}{$\bm{\epsilon}$}}_{\theta t}=\left[ \widehat{\epsilon
}_{\theta t}^{(1)},\widehat{\epsilon}_{\theta t}^{(2)},\ldots
,\widehat{\epsilon}_{t}^{(n)}\right] ^{\prime},
\]
with
\[
\widehat{\epsilon}_{t}^{(i)}=y_{1t}^{(i)}-\overline{y}_{1}^{(i)}
-\widehat{\theta}^{(i)}\left( y_{2t}^{(i)}-\overline{y}_{2}^{(i)}\right) .
\]
The long run variance estimator, $\widehat{\mathbf{\Omega}}_{\mathbf{\epsilon
}}$, is of the well known kernel form where $k(x)$ is the kernel
(downweighting function) and $M$ is the bandwidth tuning parameter that
controls the extent of downweighting by the kernel. For the case where one
restriction is being tested, $q=1$, a $t$-statistic can be used to test
one-sided hypotheses:
\[
t_{IV}=\frac{\mathbf{R}\widehat{\mathbf{\theta}}-\mathbf{r}}{\sqrt
{\mathbf{R}\widehat{\mathbf{V}}_{IV}\mathbf{R}^{\prime}}}.
\]
In order to understand the power properties of $Wald_{IV}$ and $t_{IV}$, their
asymptotic limits are derived for local alternatives, $H_{1L}$, of the form
\[
H_{1L}:\mathbf{R\theta=r+}\overline{\mathbf{\Delta}}T^{-3/2+\kappa}\mathbf{,}
\]
where $\kappa$ is the same parameter used in Theorem 1 to model the trend
slopes as local to zero. In deriving the asymptotic results, the bandwidth
parameter for $\widehat{\mathbf{\Omega}}_{\mathbf{\epsilon}}$ is assumed to be
a fixed proportion of the sample size, $b\in(0,1]$, i.e. $M=bT$. Modeling
$M/T$ as a fixed constant gives the fixed-$b$ (or fixed-smoothing) limit of
$\widehat{\mathbf{\Omega}}_{\mathbf{\epsilon}}$ (and $t_{IV}$). The advantage
of the fixed-$b$ approach is that it delivers an asymptotic random variable
and associated critical values that depend on the bandwidth and kernel. This
is in contrast to appealing to a consistency result for
$\widehat{\mathbf{\Omega}}_{\mathbf{\epsilon}}$ which would not depend on the
bandwidth or kernel. For more details on the fixed-$b$ approach see
\cite{bunzel-vogelsang}, \cite{bunzel-vogelsang}, \cite{jansson-kvb},
\cite{kvnewasym}, \cite{phillips-sun-jin-higherpower}, \cite{zhang2013fixed},
\cite{sun-twostepgmm}, \cite{LLSW18}, and \cite{LLS21}.
Because the form of the fixed-$b$ limit of the test statistics depends on the
type of kernel function used to compute $\widehat{\mathbf{\Omega}
}_{\mathbf{\epsilon}}$, definitions from \cite{kvnewasym} are used. A kernel
is labelled Type 1 if $k\left( x\right) $ is twice continuously
differentiable everywhere and as a Type 2 kernel if $k\left( x\right) $ is
continuous, $k\left( x\right) =0$ for $\left\vert x\right\vert \geq1$ and
$k\left( x\right) $ is twice continuously differentiable everywhere except
at $\left\vert x\right\vert =1.$ The Bartlett kernel (which is neither Type 1
or 2) is considered separately.
The fixed-$b$ limiting distributions are expressed in terms of the following
stochastic functions. Let $\mathbf{Q}(r)$ be a generic vector stochastic
process. Define the random variable $\mathbf{P}_{b}(\mathbf{Q}(r))$ as
\[
\mathbf{P}_{b}(\mathbf{Q}(r))=\left\{
\begin{array}
[c]{cc}
\int_{0}^{1}\int_{0}^{1}-k^{\ast\prime\prime}\left( r-s\right)
\mathbf{Q}(r)\mathbf{Q}(s)^{\prime}drds & \text{if }k\left( x\right) \text{
is Type 1}\\
& \\
\int\int_{\left\vert r-s\right\vert <b}-k^{\ast\prime\prime}\left(
r-s\right) \mathbf{Q}\left( r\right) \mathbf{Q}\left( s\right) ^{\prime
}drds & \\
+k_{-}^{\ast\prime}\left( b\right) \int_{0}^{1-b}\left( \mathbf{Q}\left(
r+b\right) \mathbf{Q}\left( r\right) ^{\prime}+\mathbf{Q}\left( r\right)
\mathbf{Q}\left( r+b\right) ^{\prime}\right) dr & \text{if }k\left(
x\right) \text{ is Type 2}\\
& \\
\frac{2}{b}\int_{0}^{1}\mathbf{Q}\left( r\right) \mathbf{Q}\left( r\right)
^{\prime}dr-\frac{1}{b}\int_{0}^{1-b}\left( \mathbf{Q}\left( r+b\right)
\mathbf{Q}\left( r\right) ^{\prime}+\mathbf{Q}\left( r\right)
\mathbf{Q}\left( r+b\right) ^{\prime}\right) dr & \text{if }k\left(
x\right) \text{ is Bartlett}
\end{array}
\right.
\]
where $k^{\ast}(x)=k\left( \frac{x}{b}\right) $ and $k_{-}^{\ast^{\prime}}$
is the first derivative of $k^{\ast}$ from below (left).
\begin{theorem}
Suppose that (\ref{fclt}) holds which implies that (\ref{fclteps}) holds. Let
$\overline{\beta}_{1}^{(i)},\overline{\beta}_{2}^{(i)}$ be fixed with respect
to $T$. Let $\mathbf{D}_{\overline{\mathbf{\beta}}_{2}}$ and $\mathbf{D}
_{\mathbf{B}_{u2}}$ be defined as in Theorem 1. Let $\mathbf{\Lambda
}_{\mathbf{\epsilon}}^{\ast}$ be the matrix square root of $\mathbf{\Omega
}_{\epsilon}^{\ast}=\mathbf{RA}_{\mathbf{B}_{2}}^{-1}\mathbf{\Lambda
}_{\mathbf{\epsilon}}\mathbf{\Lambda}_{\mathbf{\epsilon}}^{\prime}
\mathbf{A}_{\mathbf{B}_{2}}^{-1}\mathbf{R}^{\prime}$ ($\mathbf{\Lambda
}_{\mathbf{\epsilon}}^{\ast}\mathbf{\Lambda}_{\mathbf{\epsilon}}^{\ast\prime
}=\mathbf{\Omega}_{\epsilon}^{\ast}$). Suppose $\mathbf{R\theta=r+}
\overline{\mathbf{\Delta}}T^{-3/2+\kappa}$. The following hold as
$T\rightarrow\infty$.
\[
\]
Case 1 (large to small slopes): For $\beta_{1}^{(i)}=T^{-\kappa}
\overline{\beta}_{1}^{(i)}$, $\beta_{2}^{(i)}=T^{-\kappa}\overline{\beta}
_{2}^{(i)}$ with $0\leq\kappa<\frac{3}{2}$,
\[
Wald_{IV}\Rightarrow\left( \mathbf{Z}_{\epsilon}^{\ast}+\frac{1}{\sqrt{12}
}\mathbf{\Lambda}_{\mathbf{\epsilon}}^{\ast-1}\overline{\mathbf{\Delta}
}\right) ^{\prime}\mathbf{P}_{b}(\widetilde{\mathbf{W}}_{\mathbf{\epsilon}
}^{\ast}(r))^{-1}\left( \mathbf{Z}_{\epsilon}^{\ast}+\frac{1}{\sqrt{12}
}\mathbf{\Lambda}_{\mathbf{\epsilon}}^{\ast-1}\overline{\mathbf{\Delta}
}\right) ,
\]
for $q=1$
\[
t_{IV}\Rightarrow\frac{\mathbf{Z}_{\epsilon}^{\ast}+\frac{1}{\sqrt{12}
}\mathbf{\Lambda}_{\mathbf{\epsilon}}^{\ast-1}\overline{\mathbf{\Delta}}
}{\sqrt{\mathbf{P}_{b}(\widetilde{\mathbf{W}}_{\mathbf{\epsilon}}^{\ast}(r))}
},
\]
where $\mathbf{Z}_{\epsilon}^{\ast}=\sqrt{12}\int_{0}^{1}\left( s-\frac{1}
{2}\right) d\mathbf{W}_{\mathbf{\epsilon}}^{\ast}(s)$, $\widetilde{\mathbf{W}
}_{\mathbf{\epsilon}}^{\ast}(r)=\mathbf{W}_{\mathbf{\epsilon}}^{\ast
}(r)-r\mathbf{W}_{\mathbf{\epsilon}}^{\ast}(1)-12L(r)\int_{0}^{1}\left(
s-\frac{1}{2}\right) d\mathbf{W}_{\mathbf{\epsilon}}^{\ast}(s)$,
$L(r)=\int_{0}^{r}\left( s-\frac{1}{2}\right) ds$, and $\mathbf{W}
_{\mathbf{\epsilon}}^{\ast}(r)$ is a $q\times1$ vector of independent Wiener
processes. Note that $\mathbf{Z}_{\epsilon}^{\ast}\sim\mathbf{N}
(0,\mathbf{I}_{q})$ and is independent of $\widetilde{\mathbf{W}
}_{\mathbf{\epsilon}}^{\ast}(r)$.
\[
\]
Case 2 (very small slopes): For $\beta_{1}^{(i)}=T^{-3/2}\overline{\beta}
_{1}^{(i)},$ $\beta_{2}^{(i)}=T^{-3/2}\overline{\beta}_{2}^{(i)}$
($\kappa=\frac{3}{2}$),
\begin{align*}
Wald_{IV} & \Rightarrow\left( \mathbf{R}\left( \frac{1}{12}\mathbf{D}
_{\overline{\mathbf{\beta}}_{2}}+\mathbf{D}_{\mathbf{B}_{u2}}\right)
^{-1}\mathbf{\Lambda}_{\mathbf{\epsilon}}\int_{0}^{1}\left( s-\frac{1}
{2}\right) d\mathbf{W}_{\mathbf{\epsilon}}(s)+\overline{\mathbf{\Delta}
}\right) ^{\prime}\\
& \times\left[ \mathbf{R}\left( \frac{1}{12}\mathbf{D}_{\overline
{\mathbf{\beta}}_{2}}+\mathbf{D}_{\mathbf{B}_{u2}}\right) ^{-1}\mathbf{P}
_{b}(\mathbf{H}_{1}(r))\left( \frac{1}{12}\mathbf{D}_{\overline
{\mathbf{\beta}}_{2}}+\mathbf{D}_{\mathbf{B}_{u2}}\right) ^{-1}
\mathbf{R}^{\prime}\right] ^{-1}\\
& \times\left( \mathbf{R}\left( \frac{1}{12}\mathbf{D}_{\overline
{\mathbf{\beta}}_{2}}+\mathbf{D}_{\mathbf{B}_{u2}}\right) ^{-1}
\mathbf{\Lambda}_{\mathbf{\epsilon}}\int_{0}^{1}\left( s-\frac{1}{2}\right)
d\mathbf{W}_{\mathbf{\epsilon}}(s)+\overline{\mathbf{\Delta}}\right) ,
\end{align*}
for $q=1$
\[
t_{IV}\Rightarrow\frac{\mathbf{R}\left( \frac{1}{12}\mathbf{D}_{\overline
{\mathbf{\beta}}_{2}}+\mathbf{D}_{\mathbf{B}_{u2}}\right) ^{-1}
\mathbf{\Lambda}_{\mathbf{\epsilon}}\int_{0}^{1}\left( s-\frac{1}{2}\right)
d\mathbf{W}_{\mathbf{\epsilon}}(s)+\overline{\mathbf{\Delta}}}{\sqrt{\frac
{1}{12}\mathbf{R}\left( \frac{1}{12}\mathbf{D}_{\overline{\mathbf{\beta}}
_{2}}+\mathbf{D}_{\mathbf{B}_{u2}}\right) ^{-1}\mathbf{P}_{b}(\mathbf{H}
_{1}(r))\left( \frac{1}{12}\mathbf{D}_{\overline{\mathbf{\beta}}_{2}
}+\mathbf{D}_{\mathbf{B}_{u2}}\right) ^{-1}\mathbf{R}^{\prime}}},
\]
where $\mathbf{D}_{\overline{\mathbf{\beta}}_{2}}$ and $\mathbf{D}
_{\mathbf{B}_{u2}}$ are defined in Theorem 1,
\[
\mathbf{H}_{1}(r)=\mathbf{\Lambda}_{\mathbf{\epsilon}}\left( \mathbf{W}
_{\mathbf{\epsilon}}(r)-r\mathbf{W}_{\mathbf{\epsilon}}(1)\right) -\left(
L(r)\mathbf{D}_{\overline{\mathbf{\beta}}_{2}}+\mathbf{D}_{\widehat{\mathbf{B}
}_{u2}(r)}\right) \left( \left( \frac{1}{12}\mathbf{D}_{\overline
{\mathbf{\beta}}_{2}}+\mathbf{D}_{\mathbf{B}_{u2}}\right) ^{-1}
\mathbf{\Lambda}_{\mathbf{\epsilon}}\int_{0}^{1}\left( s-\frac{1}{2}\right)
d\mathbf{W}_{\mathbf{\epsilon}}(s)\right) ,
\]
and $\mathbf{D}_{\widehat{\mathbf{B}}_{u2}(r)}$ is an $n\times n$ diagonal
matrix with $i^{th}$ diagonal element $\widehat{B}_{u2}^{(i)}(r)=B_{u2}
^{(i)}(r)-rB_{u2}^{(i)}(1)$.
\[
\]
Case 3 (zero slopes): For $\beta_{1}^{(i)}=0$, $\beta_{2}^{(i)}=0$,
\begin{align*}
Wald_{IV} & \Rightarrow\left( \mathbf{RD}_{\mathbf{B}_{u2}}^{-1}\int_{0}
^{1}\left( s-\frac{1}{2}\right) d\mathbf{B}_{\mathbf{u}1}(s)-\mathbf{r}
\right) ^{\prime}\left[ \frac{1}{12}\mathbf{RD_{\mathbf{B}_{u2}}
^{-1}\mathbf{P}_{b}(\mathbf{H}_{2}(r))D_{\mathbf{B}_{u2}}^{-1}R}^{\prime
}\right] ^{-1}\\
& \text{ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ }\times\left(
\mathbf{RD}_{\mathbf{B}_{u2}}^{-1}\int_{0}^{1}\left( s-\frac{1}{2}\right)
d\mathbf{B}_{\mathbf{u}1}(s)-\mathbf{r}\right) ,
\end{align*}
for $q=1$
\[
t_{IV}\Rightarrow\frac{\mathbf{RD}_{\mathbf{B}_{u2}}^{-1}\int_{0}^{1}\left(
s-\frac{1}{2}\right) d\mathbf{B}_{\mathbf{u}1}(s)-\mathbf{r}}{\sqrt{\frac
{1}{12}\mathbf{RD_{\mathbf{B}_{u2}}^{-1}P}_{b}(\mathbf{H}_{2}
(r))\mathbf{D_{\mathbf{B}_{u2}}^{-1}R}^{\prime}}},
\]
where
\[
\mathbf{H}_{2}(r)=\mathbf{B}_{\mathbf{u}1}(s)-r\mathbf{B}_{\mathbf{u}
1}(1)-\mathbf{D}_{\widehat{\mathbf{B}}_{u2}(r)}\mathbf{D}_{\mathbf{B}_{u2}
}^{-1}\int_{0}^{1}\left( s-\frac{1}{2}\right) d\mathbf{B}_{\mathbf{u}1}(s).
\]
\end{theorem}
The limit of $t_{IV}$ given by Case 1 of Theorem 2 is of the same form as the
limit derived by \cite{vogelsang-nawaz-JTSA} for case of a single trend ratio
($n=1$). Under the null hypothesis (when $\overline{\mathbf{\Delta}
}=\mathbf{0}$\textbf{)}, the limits of $Wald_{IV}$ and $t_{IV}$ are identical
to limits obtained by \cite{bunzel-vogelsang} in deterministic trend
regression models. While the limiting random variables are nonstandard because
of the fixed-$b$ limit of $\widehat{\mathbf{\Omega}}_{\mathbf{\epsilon}}$, the
$\mathbf{P}_{b}(\widetilde{\mathbf{W}}_{\mathbf{\epsilon}}^{\ast}(r))$ term in
the denominator, critical values are available using formulas from
\cite{bunzel-vogelsang}. The critical values depend on the bandwidth through
$b=M/T$ and the kernel, $k(x)$.
When the trend slopes are very small or zero, the limiting distributions
change, become more complicated, and depend on nuisance parameters. This is to
be expected given the weak instrument problem that occurs when the $\beta
_{2}^{(i)}$ slopes are small and the fact that $\mathbf{\theta}$ is not
defined when slopes are zero. The finite sample simulations will illustrate
the extent to which inference breaks down as the trend slopes become very
small or zero.
\section{Product Approach for Testing Equal Trend Ratios}
When testing a simple hypothesis about a single trend slope ratio,
\cite{vogelsang-nawaz-JTSA} used Fieller's method (\cite{fieller1954}) to
construct confidence intervals that are robust to very small trends slopes
including the case of zero trend slopes. This approach is based on rewriting a
simple hypothesis about $\theta^{(i)}$ in terms of a linear restriction
involving $\beta_{1}^{(i)}$ and $\beta_{2}^{(i)}$. However, once a null
hypothesis involves a linear combination of at least two $\theta^{(i)}$,
Fieller's method cannot be applied.
One important empirical application is testing equality of trend slope ratios
for two pairs of series. See VMCS26 for tests of equal trend ratios between
observed and model generated temperature series. For the null hypothesis of
equal trend ratios, it is possible to develop a testing approach similar in
spirit to Fieller's method that gives potentially more robust inference when
slopes are very small. Suppose there are two pairs of series with trend ratios
$\theta^{(1)}$ and $\theta^{(2)}$ and the hypothesis of interest is
\[
H_{0}:\theta^{(1)}=\theta^{(2)},
\]
or equivalently
\begin{equation}
H_{0}:\theta^{(1)}-\theta^{(2)}=0. \label{h0equaltheta}
\end{equation}
Anticipating a local asymptotic calculation, suppose the alternative is
specified as local to zero
\[
H_{1L}:\theta^{(1)}-\theta^{(2)}=\overline{\mathbf{\Delta}}T^{-3/2+\kappa}.
\]
Rewriting the alternative hypothesis in terms of the trend slopes gives
\[
H_{1L}:\frac{\beta_{1}^{(1)}}{\beta_{2}^{(1)}}-\frac{\beta_{1}^{(2)}}
{\beta_{2}^{(2)}}=\overline{\mathbf{\Delta}}T^{-3/2+\kappa}.
\]
Multiplying both sides of $H_{1L}$ by $\beta_{2}^{(1)}\beta_{2}^{(2)}$ gives
\begin{equation}
H_{1L}:\beta_{2}^{(2)}\beta_{1}^{(1)}-\beta_{2}^{(1)}\beta_{1}^{(2)}=\beta
_{2}^{(1)}\beta_{2}^{(2)}\overline{\mathbf{\Delta}}T^{-3/2+\kappa}.
\label{h1prod}
\end{equation}
Testing (\ref{h0equaltheta}) is equivalent to testing
\begin{equation}
H_{0}:\beta_{2}^{(2)}\beta_{1}^{(1)}-\beta_{2}^{(1)}\beta_{1}^{(2)}=0.
\label{h0prod}
\end{equation}
The advantage of using (\ref{h0prod}) is that the trend slopes can be directly
estimated by OLS and ratios are avoided.
To develop a test statistic for testing (\ref{h0prod}) against (\ref{h1prod})
define
\[
g_{\beta}=\beta_{2}^{(2)}\beta_{1}^{(1)}-\beta_{2}^{(1)}\beta_{1}^{(2)}.
\]
The natural estimator of $g_{\beta}$ is given by
\[
g_{\widehat{\beta}}=\widehat{\beta}_{2}^{(2)}\widehat{\beta}_{1}
^{(1)}-\widehat{\beta}_{2}^{(1)}\widehat{\beta}_{1}^{(2)},
\]
where $\widehat{\beta}_{1}^{(i)}$ and $\widehat{\beta}_{2}^{(i)}$ are the OLS
estimators (\ref{beta1ols}) and (\ref{beta2ols}).
Because $g_{\widehat{\beta}}$ is a nonlinear function of the slope estimators,
its asymptotic variance depends on $\mathbf{\Omega}_{\mathbf{u}}$, the long
run variance of $\mathbf{U}_{t}=\left[ u_{1t}^{(1)},u_{1t}^{(2)},u_{2t}
^{(1)},u_{2t\mathbf{U}_{t}}^{(2)}\right] ^{\prime}$, and the vector
\[
\mathbf{R}_{\beta}=\left[ \beta_{2}^{(2)},-\beta_{2}^{(1)},-\beta_{1}
^{(2)},\beta_{1}^{(1)}\right] ,
\]
which can be derived using the delta method or, as in the appendix, directly.
The feasible version of $\mathbf{R}_{\beta}$ is given by
\[
\mathbf{R}_{\widehat{\beta}}=\left[ \widehat{\beta}_{2}^{(2)},-\widehat{\beta
}_{2}^{(1)},-\widehat{\beta}_{1}^{(2)},\widehat{\beta}_{1}^{(1)}\right] .
\]
Let $\widehat{\mathbf{U}}_{t}=\left[ \widehat{u}_{1t}^{(1)},\widehat{u}
_{1t}^{(2)},\widehat{u}_{2t}^{(1)},\widehat{u}_{2t}^{(2)}\right] ^{\prime}$
where $\widehat{u}_{1t}^{(i)},\widehat{u}_{2t}^{(i)}$ are the residuals from
(\ref{1.1}) and (\ref{1.2}) estimated by OLS. Define the long run variance
estimator of $\mathbf{\Omega}_{\mathbf{u}}$ as
\[
\widehat{\mathbf{\Omega}}_{\mathbf{u}}=\widehat{\mathbf{\Gamma}}_{\mathbf{u}
0}+
{\displaystyle\sum\limits_{j=1}^{T-1}}
k\left( \frac{j}{M}\right) (\widehat{\mathbf{\Gamma}}_{\mathbf{u}
j}+\widehat{\mathbf{\Gamma}}_{\mathbf{u}j}^{\prime}),\text{ \ \ \ }
\widehat{\mathbf{\Gamma}}_{\mathbf{u}j}=T^{-1}
{\displaystyle\sum\limits_{t=j+1}^{T}}
\widehat{\mathbf{U}}_{t}\widehat{\mathbf{U}}_{t-j}^{\prime}.
\]
Using
\[
\widehat{\lambda}_{g}^{2}=\mathbf{R}_{\widehat{\beta}}\widehat{\mathbf{\Omega
}}_{\mathbf{u}}\mathbf{R}_{\widehat{\beta}}^{\prime},
\]
a $t$-statistic for testing (\ref{h0prod}) can be constructed as
\[
t_{prod}=\frac{g_{\widehat{\beta}}}{\sqrt{\widehat{\lambda}_{g}^{2}\left(
{\textstyle\sum\nolimits_{t=1}^{T}}
(t-\overline{t})^{2}\right) ^{-1}}}.
\]
The following Theorem gives the limiting distribution of $t_{prod}$ under the
local alternative (\ref{h1prod}).
\begin{theorem}
Suppose that (\ref{fclt}) holds. Let $\overline{\beta}_{1}^{(i)}
,\overline{\beta}_{2}^{(i)}$ be fixed with respect to $T$. Let $\mathbf{D}
_{\overline{\mathbf{\beta}}_{2}}$and $\mathbf{D}_{\mathbf{B}_{u2}}$be defined
as in Theorem 1 for the case of $n=2$. Suppose $\theta^{(1)}-\theta
^{(2)}=\overline{\mathbf{\Delta}}T^{-3/2+\kappa}$. The following hold as
$T\rightarrow\infty$:
\[
\]
Case 1 (large to small slopes): For $\beta_{1}^{(i)}=T^{-\kappa}
\overline{\beta}_{1}^{(i)}$, $\beta_{2}^{(i)}=T^{-\kappa}\overline{\beta}
_{2}^{(i)}$ with $0\leq\kappa<\frac{3}{2}$,
\[
t_{prod}\Rightarrow\frac{z_{u}^{\ast}+\frac{\overline{\beta}_{2}
^{(1)}\overline{\beta}_{2}^{(2)}\overline{\mathbf{\Delta}}}{\Lambda_{u}^{\ast
}\sqrt{12}}}{\sqrt{P_{b}(\widetilde{w}_{u}^{\ast}(r))}},
\]
where $z_{u}^{\ast}=\sqrt{12}\int_{0}^{1}\left( s-\frac{1}{2}\right)
dw_{u}^{\ast}(s)$, $\widetilde{w}_{u}^{\ast}(r)=\widetilde{w}_{u}^{\ast
}(r)-rw_{u}^{\ast}(1)-12L(r)\int_{0}^{1}\left( s-\frac{1}{2}\right)
dw_{u}^{\ast}(s)$, $L(r)=\int_{0}^{r}\left( s-\frac{1}{2}\right) ds$,
$w_{u}^{\ast}(r)$ is a standard Wiener process, and $\Lambda_{u}^{\ast}
=\sqrt{\mathbf{R}_{\overline{\beta}}\mathbf{\Omega}_{\mathbf{u}}
\mathbf{R}_{\overline{\beta}}^{\prime}}$. Note that $z_{u}^{\ast}\sim N(0,1)$
and is independent of $\widetilde{w}_{u}^{\ast}(r)$.
\[
\]
Case 2 (very small slopes): For $\beta_{1}^{(i)}=T^{-3/2}\overline{\beta}
_{1}^{(i)},$ $\beta_{2}^{(i)}=T^{-3/2}\overline{\beta}_{2}^{(i)}$
($\kappa=\frac{3}{2}$),
\[
t_{prod}\Rightarrow\frac{\mathbf{R}_{\overline{\beta}}\Psi+\Psi_{2}^{(2)}
\Psi_{1}^{(1)}-\Psi_{2}^{(1)}\Psi_{1}^{(2)}+\overline{\mathbf{\Delta}
}\overline{\beta}_{2}^{(1)}\overline{\beta}_{2}^{(2)}}{\sqrt{12\left(
\mathbf{R}_{\overline{\beta}}+\left[ \Psi_{2}^{(2)},-\Psi_{2}^{(1)},-\Psi
_{1}^{(2)},\Psi_{1}^{(1)}\right] \right) \mathbf{P}_{b}
(\widetilde{\mathbf{B}}_{\mathbf{u}}(r))\left( \mathbf{R}_{\overline{\beta}
}+\left[ \Psi_{2}^{(2)},-\Psi_{2}^{(1)},-\Psi_{1}^{(2)},\Psi_{1}
^{(1)}\right] \right) ^{\prime}}},
\]
where $\mathbf{\Psi}=\left[ \Psi_{1}^{(1)},\Psi_{1}^{(2)},\Psi_{2}^{(1)}
,\Psi_{2}^{(2)}\right] ^{\prime}=12\int_{0}^{1}\left( s-\frac{1}{2}\right)
d\mathbf{B}_{\mathbf{u}}(s)$, $\widetilde{\mathbf{B}}_{\mathbf{u}
}(r)=\mathbf{B}_{\mathbf{u}}(r)-r\mathbf{B}_{\mathbf{u}}(1)-12L(r)\int_{0}
^{1}\left( s-\frac{1}{2}\right) d\mathbf{B}_{\mathbf{u}}(s).$
\[
\]
Case 3 (zero slopes): For $\beta_{1}^{(i)}=0$, $\beta_{2}^{(i)}=0$,
\[
t_{prod}\Rightarrow\frac{\Psi_{2}^{(2)}\Psi_{1}^{(1)}-\Psi_{2}^{(1)}\Psi
_{1}^{(2)}}{\sqrt{12\left[ \Psi_{2}^{(2)},-\Psi_{2}^{(1)},-\Psi_{1}
^{(2)},\Psi_{1}^{(1)}\right] \mathbf{P}_{b}(\widetilde{\mathbf{B}
}_{\mathbf{u}}(r))\left[ \Psi_{2}^{(2)},-\Psi_{2}^{(1)},-\Psi_{1}^{(2)}
,\Psi_{1}^{(1)}\right] ^{\prime}}}.
\]
\end{theorem}
When trend slopes are large to small, the asymptotic distribution of
$t_{prod}$ has the same form as $t_{IV}$. The two asymptotic distributions are
the same under the null ($\overline{\mathbf{\Delta}}=0$) but are different
under the local alternative because the variance parameters in the limit are
different. As is the case for $t_{IV}$, the asymptotic distribution of
$t_{prod}$ changes when trends slopes are very small or zero. Given the
complexity of the limits in Cases 2 and 3, it is not clear what Theorem 3
predicts about the finite sample behavior of $t_{prod}$ as trend slopes become
very small or zero. Finite simulations in the next section will be used to
explore the implications of very small or zero trend slopes.
\section{Finite Sample Null Rejection Probabilities and Power for Tests of
Equal Ratios}
This section provides simulation results to show some finite sample properties
of tests of equal trend slopes. The following DGP was used. The $y_{1t}^{(i)}$
and $y_{2t}^{(i)}$ variables where generated by models (\ref{1.1}) and
(\ref{1.2}) where the noise is given by
\begin{align*}
u_{1t}^{(1)} & =\phi_{1}^{(1)}u_{1t-1}^{(1)}+\varepsilon_{1t}^{(1)},\text{
}u_{2t}^{(1)}=\phi_{2}^{(1)}u_{2t-1}^{(1)}+\varepsilon_{2t}^{(1)},\\
u_{1t}^{(2)} & =\phi_{1}^{(2)}u_{1t-1}^{(2)}+\varepsilon_{1t}^{(2)},\text{
}u_{2t}^{(2)}=\phi_{2}^{(2)}u_{2t-1}^{(2)}+\varepsilon_{2t}^{(2)},
\end{align*}
\[
\left[
\begin{array}
[c]{c}
\varepsilon_{1t}^{(1)}\\
\varepsilon_{2t}^{(1)}\\
\varepsilon_{1t}^{(2)}\\
\varepsilon_{2t}^{(2)}
\end{array}
\right] \sim iidN\left( \mathbf{0},\left[
\begin{array}
[c]{cc}
\begin{array}
[c]{cc}
1 & \varphi\\
\varphi & 1
\end{array}
&
\begin{array}
[c]{cc}
0 & 0\\
0 & 0
\end{array}
\\%
\begin{array}
[c]{cc}
0 & 0\\
0 & 0
\end{array}
&
\begin{array}
[c]{cc}
1 & \varphi\\
\varphi & 1
\end{array}
\end{array}
\right] \right) ,
\]
\[
u_{10}^{(1)}=u_{20}^{(1)}=u_{10}^{(2)}=u_{20}^{(2)}=0.
\]
All the noise series are generated by AR(1) processes. Pairs of series are
uncorrelated with each other but there can be within-pair correlation
($\varphi\neq0$).
Given that $t_{IV}$ and $t_{prod}$ are exactly invariant to the values of the
intercept parameters, without loss of generality the intercept parameters are
set to zero: $\mu_{1}^{(1)}=0,\mu_{2}^{(1)}=0,\mu_{1}^{(2)}=0,\mu_{2}^{(2)}
=0$. Results are reported for various magnitudes of $\beta_{1}^{(i)}$ and
$\beta_{2}^{(i)}$ where $\theta^{(i)}=\beta_{1}^{(i)}/\beta_{2}^{(i)}=1$ under
the null hypothesis of equal ratios across the two pairs. Empirical null
rejections are given for $T=50,100,200$ with $10,000$ replications used in all
cases. Empirical power is given for $T=100$ where $\theta^{(2)}$ takes on
values different from $\theta^{(1)}=1$.
Table 1 reports null rejection probabilities for 5\% nominal level tests for
testing $H_{0}:\theta^{(1)}=\theta^{(2)}$ against the two-sided alternative
$H_{1}:\theta^{(1)}\neq\theta^{(2)}$. Results are reported for the values of
$\beta_{1}^{(i)}=\beta_{2}^{(i)}=10,2,.2,.05,0.25,.005,0$ giving $\theta
^{(i)}=1$ except for the case of $\beta_{1}^{(i)}=\beta_{2}^{(i)}=0$ where
$\theta^{(i)}$ is not defined. The variance estimators use the Daniell kernel.
Results for four bandwidth sample size ratios are provided:
$b=A91,0.25,0.5,1.0$, where $A91$ is the AR(1) plug-in data dependent
bandwidth proposed by \cite{andrews-91}. Empirical rejections are computed
using fixed-$b$ asymptotic critical values with the critical value function
\[
cv_{0.025}(b)=1.9659+4.0603b+11.6626b^{2}+34.8269b^{3}-13.9506b^{4}
+3.2669b^{5}\text{,}
\]
as given by \cite{bunzel-vogelsang} for the Daniell kernel for $k(x)$.
The top panels of Table 1 give results for the case of iid noise as a
benchmark. Here $\phi_{1}^{(1)}=\phi_{2}^{(1)}=\phi_{1}^{(2)}=\phi_{2}
^{(2)}=\varphi=0$. As long as the trend slopes are large, empirical rejections
are close to the nominal level of 0.05 for both $t_{IV}$ and $t_{prod}$ and
all values of $b$. As the trend slopes get smaller, both tests have rejections
below 0.05 (conservative tests). This happens more quickly with $t_{IV}$ than
$t_{prod}$ and more quickly with smaller values of $b$ (smaller bandwidths).
As $T$ increases, empirical rejections are close to the nominal level for
relatively smaller trend slopes. This makes sense given the asymptotic
results. For given values of $\beta_{1}^{(i)}$ and $\beta_{2}^{(i)}$, the
values of the local trend parameters, $\overline{\beta}_{1}^{(i)}$ and
$\overline{\beta}_{2}^{(i)}$, are larger for bigger values of $T$.
The bottom panels of Table 1 give results for serially correlated noise with
$\phi_{1}^{(1)}=0.3$, $\phi_{2}^{(1)}=0.7$, $\phi_{1}^{(2)}=0.5$, $\phi
_{2}^{(2)}=0.9$. Pairs of series have within-pairs correlation given by
$\varphi=0.5$. The first pair of noise series has modest serial correlation in
the numerator series, $u_{1t}^{(1)}$, and moderate serial correlation in the
denominator series, $u_{2t}^{(1)}$. The second pair of noise series has
stronger serial correlation than the first pair with the denominator series
again having stronger serial correlation than the numerator series. The
relative strengths of serial correlation within pairs are a feature of
temperature series where upper atmosphere temperature (numerators) tend to
have less serial correlation than surface temperatures (denominators). The
relative strength across pairs are a feature of observed versus model
generated temperatures where observed temperatures (first pair) tend to have
less serial correlation than model generated series (second pair).
As in the iid case, empirical rejections in the serial correlation case tend
to fall as trend slopes become smaller indicating that the tests remain
conservative when trend slope are small (or even zero). For larger trend
slopes, the positive autocorrelation can lead to over-rejections especially if
the data-dependent bandwidth is used. As $b$ increases, the tendency to
over-reject is mitigated. This is a well-known property (see
\cite{bunzel-vogelsang} and \cite{kvnewasym} among others). It is not
surprising that over-rejections are larger with the data-dependent bandwidth
because those bandwidths tend to give relatively small values of $b$.
Comparing $t_{IV}$ and $t_{prod}$, one can see that $t_{IV}$ tends to
over-reject less than $t_{prod}$ especially when using the data-dependent
bandwidth. In contrast, $t_{IV}$ tends to under-reject more than $t_{prod}$ as
trend slopes become smaller. The main takeaway is that $t_{IV}$ over-rejects
less than $t_{prod}$ when bandwidths are small and trend slopes are large, but
$t_{IV}$ is more conservative when trend slopes are small (regardless of
bandwidth). As a practical matter, using the slightly larger bandwidth of
$b=0.25$ gives non-trivial reductions in over-rejections relative to the $A91$
bandwidth rule.
Table 2 gives results for empirical power of the tests. Results are only given
for the case of serially correlated noise (patterns are similar for iid noise)
with $T=100$. In all cases $\theta^{(1)}=1$, and results are given for a grid
of values of $\theta^{(2)}$ above and below $1$. The range of the grid
increases as the trend slopes decrease to provide information about the power
curves. It is not surprising that in order to see power, the range of
$\theta^{(2)}$ needs to be increased as trend slopes decrease - this is
predicted by the slower rate of convergence of $\widehat{\mathbf{\theta}}$
(larger sampling variance) as trend slopes become smaller. Empirical power is
not size-adjusted to show actual power in practice. Null rejections are in
bold for the cases where $\theta^{(2)}=1$.
When trend slopes are large, power is similar for $t_{IV}$ and $t_{prod}$ for
the bandwidths $b=0.25,0.5,1.0$. Power is roughly symmetric around null value
of $\theta^{(1)}=1$. When the $A91$ data-dependent bandwidth is used,
$t_{prod}$ over-rejects more than $t_{IV}$ and has correspondingly higher
power. With smaller trend slopes some differences in power emerge between
$t_{IV}$ and $t_{prod}$ that are not solely due to differences in null
rejections. For example, when $\beta_{2}^{(1)}=\beta_{2}^{(2)}=0.2$, $t_{IV}$
and $t_{prod}$ have similar null rejections with $b=0.5,1.0$. In these cases
power of $t_{IV}$ is higher for $\theta^{(2)}<1$, whereas power of $t_{prod}$
is higher for $\theta^{(2)}>1$. Power for both tests is higher with smaller
values of $b$ --- another well-known feature of tests based on kernel variance
estimators that use fixed-$b$ critical values. As trend slopes become smaller,
power of both tests decreases and power can be low even for values of
$\theta^{(2)}$ very far from $1$. Again, this is not surprising because
smaller trend slopes have relatively less information about $\theta^{(1)}$ and
$\theta^{(2)}$ for given strength of the noise. Interestingly, with very small
trend slopes, $\beta_{2}^{(1)}=\beta_{2}^{(2)}=0.05,0.025$, power initially
increases as $\theta^{(2)}$ moves away from the null value of $1$ but can
begin to fall when $\theta^{(2)}$ is very far from $1$.
The main takeaways from the power results are: i) for large trend slopes,
$t_{IV}$ and $t_{prod}$ have similar power, ii) for medium to small trend
slopes power cannot be ranked, iii) power of both tests decreases as the
bandwidth increases and iv) there is low power in detecting differences
between trend ratios when trend slopes are small.
\section{Practical Recommendations}
The theory and simulations indicate that $t_{IV}$ and $t_{prod}$ perform
similarly in practice when i) trend slopes are not very small, and ii) small
bandwidths are avoided when there is nontrivial serial correlation in the
noise. Both tests can over-reject when there is positive serial correlation in
the noise although this problem is mitigated by larger sample sizes. For very
small trend slopes both tests become conservative with $t_{IV}$ becoming
conservative more quickly than $t_{prod}$. Power of the tests cannot be ranked
in this case. Because of the conservative nature of both tests when trend
slopes are small (or even zero), any rejections obtained for tests of equal
ratios are robust to very small trend slopes. The price paid for this
robustness is lower power, but one cannot expect trend ratios to be precisely
estimated when trend slopes are very small.
If a practitioner uses either test (or both) with the Daniell kernel and a
non-small bandwidth for the variance estimator, then rejections of equal
ratios can be viewed as relatively robust to stationary serial correlation in
the noise and very small trend slopes.
\section{Empirical Application}
VMCS26 used the methods developed in this paper to test equivalence between
trend ratios of observed temperature series and temperature series generated
by recent runs of climate models. They compared trend ratios of temperature
series at various atmospheric heights relative to trends in surface
temperatures for the tropics region of the earth. These trend slope ratios are
called amplification ratios because climate models predict amplified warming
in the lower to mid troposphere relative to the surface in the tropics. VMCS26
found that climate models tend to have more amplification that seen in
observed temperature series and that the differences are statistically
significant using the $t_{IV}$ and $t_{prod}$ test statistics. While VMCS26
compared amplification ratios between each of five sets of observed
temperatures with each of 39 sets of climate model temperature series, they
did not compare amplification ratios among the five sets of observed
temperature series. Comparisons across observed series is interesting given
the different methods by which the observed series are measured, constructed,
and aggregated.
VMCS26 provide details on the five sets of observed temperature series which
can be summarized as follows. Observed temperatures for various pressure
levels in the atmosphere are measured in two ways in the five sets of observed
temperatures. The first method uses station-based, balloon-borne radiosonde
records. Temperature data is collected at specific pressure levels, generally
up to 20 hectopascals (hPa) which is approximately 27 km above the earth's
surface. A smaller hPa value indicates a higher level about the surface. The
second method is known as reanalyses of weather data combined with climate
models to generate temperature series. The five observed data series include
three from balloons. Two are data series from the University of Vienna,
RAOBCORE v1.9 and RICH v1.9. The third is from the U.S. National Oceanic and
Atmospheric Administration, the RATPAC-A v2 data series. These sets of data
are labeled RAOB, RICH, and RATP. The two global reanalyses data sets are
taken from the European Centre for Medium-Range Forecasts Reanalyses (ERA5),
and the Japanese Reanalyses for Three Quarters of a Century (JRA3Q).
The data is aggregated on an annual basis and spans the years 1958 to 2024 (67
years) for the surface and the grid of hPa levels 850, 700, 500, 400, 300,
200, 150, 100, 70, 50, 30, 20. Data is not available for the RATP data set for
hPa 20. Empirical results are presented in three tables where in all cases
confidence intervals at the 95\% level are computed using the same variance
estimators and fixed-$b$ critical values used in the finite sample simulations
(Daniell kernel, A91 bandwidth, fixed-$b$ critical values).
Table 3 provides summary statistics in the form of OLS estimated linear trend
slope parameters using regression (\ref{1.1}) for each of the five sets of
temperatures across the hPa levels. Estimated trend slopes and confidence
intervals are scaled to be in units of degrees Celsius per \textit{decade.}
Nearly all of the estimated trend slopes are statistically significantly
different from zero. For each of the five sets of series, there is a clear
pattern of trend slopes across hPa levels. There is warming at the surface.
There is less warming at 850 hPa but more warming from hPa levels 700 to 150.
At hPa 100 up to 20, there is cooling with cooling increasing at higher altitudes.
Table 4 reports estimated trend ratios for each hPa level relative to the
surface. Confidence intervals are computed using Fieller's method following
\cite{vogelsang-nawaz-JTSA}. All five sets of observed temperatures show the
same pattern in estimated trend ratios across hPa levels. Amplification
(ratios greater than 1) occurs for hPa levels 700 to 200. For hPa level 150
two observed series show amplification whereas three do not. For hPa levels
100 to 20, there is cooling at these higher altitudes and ratios are negative
and increase in magnitude for higher altitudes (lower hPa values). While the
patterns across hPa levels are similar for the five sets of observed
temperature series, there can be noticeable variation across the five sets for
a given hPa level. This raises the question as to whether ratios are equal
across pairs of observed series for a given hPa level.
Table 5 reports, for each hPa level, pairwise differences between estimated
ratios, $\widehat{\Delta}_{\theta}$, and pairwise $g_{\widehat{\beta}}$
values. 95\% fixed-$b$ critical values are given below each estimate. A
$^{\ast}$ superscript on an estimated value indicates a rejection of the null
hypothesis of equal ratios (indicates the confidence interval does not contain
the value $0$). For reporting purposes, the values of $g_{\widehat{\beta}}$
and confidence intervals are scaled by $10^{4}$. Of course, this has no effect
on whether the null hypothesis of equal ratios can be rejected. Table 5 is
divided into four panels where each panel reports results for three hPa
levels. In general, there are many pairs where the differences in trend ratios
are not small in magnitude and are statistically significant. For example,
among the three balloon data sets, RICH, RAOB, RATP, there are rejections of
equal ratios for 8 to 11 of the hPa levels across the three respective pairs.
One case where there is relative alignment of trend ratios is between the two
Reanalyses data sets, ERA5 and JRA3Q, where rejections of equal ratios occurs
for only 3 of 12 hPa levels. Except between the pair of Reanalyses data sets,
the results in Table 5 suggest there are differences in trend ratios among
pairs of the observed data sets across hPa levels.
\section{Conclusion}
This paper develops estimation and inference methods for systems of pairs of
trend stationary time series where the parameters of interest are ratios of
trend slopes between pairs of time series. Inference focuses on null
hypotheses that can be written as linear restrictions across trend ratios.
Trend slopes are estimated by IV using time as an instrument and test
statistics are robust to i) stationary serial correlation in the fluctuations
around trend, and ii) correlation between and across pairs of time series. For
the empirically relevant special case of testing for equal trend slopes
between two pairs of time series, an alternative testing approach is developed
by restating the equal trend ratio restriction as a restriction involving
products of the underlying trend slopes which are estimated by OLS. Theory and
finite sample results suggest that both the IV and products approach work well
and have similar properties when trend slopes are not too small and serial
correlation is not too strong. When trend slopes are very small relative to
the variation around the trend function, the null limiting distribution of
both tests become nonpivotal and both tests tend to under-reject under the
null and have low power. Lower power is expected when trend slopes are very
small because trend ratios cannot be precisely estimated. That both tests
become conservative when trend slopes are very small gives the tests useful
robustness to very small trend slopes. Overall, the IV and product approaches
have complementary finite sample properties and using both is recommended in
practice for testing the hypothesis of equal ratios across two pairs of time series.