EconBase
← Back to paper

Estimating and Testing Kinks in Panel Data Models

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

153,180 characters

Estimating and Testing Kinks in Panel Data Models


\title{\textbf{Estimating  and Testing Kinks in Panel Data Models}
\Large{\emph{}}}

{\author{Yousef Kaddoura\thanks{\"Orebro University
 E-mail address: \texttt{[email removed]}.}\\
{\small \"Orebro University}\\ }}





\maketitle
\begin{abstract}
Many economic and financial relationships may change gradually rather than abruptly. We study panel data models in which the coefficient vector is continuous and piecewise linear in calendar time, with a finite number of unknown kink dates at which its slope changes. We propose a penalised least squares estimator that applies adaptive weighted group penalties to the second differences of the coefficient path, and develop asymptotic theory showing that it recovers both the number and the locations of the kinks with probability approaching one. To our knowledge, this is the first panel framework to estimate an unknown number of common kink dates in a time-varying coefficient path under fixed effects. We establish that endpoint slopes converge at the usual cubic regime-length rate and interior slopes at rates determined by their own and adjacent regime lengths. We also develop a coefficient-by-coefficient extension allowing individual regressors to kink at different dates. Monte Carlo evidence supports the good finite sample properties, and we illustrate the method through an application in macro-finance, specifically the relationship between debt and growth.
\end{abstract}



\textbf{Keywords:} Adaptive group lasso; Kink detection; Panel data; Structural change; Time-varying coefficients; Growth.





\textbf{JEL Classification:} C13; C23; C33.




\section{Introduction} \label{one}
Certain economic and financial relationships may not change overnight. When a
central bank shifts its policy stance, when a trade agreement phases
in, or when a tax reform gradually alters incentives, the response
of economic and financial agents unfolds over time. The coefficients governing
these relationships may not jump discontinuously, but they might
``bend.'' Yet the dominant paradigm in the structural change
literature focuses precisely on such jumps: abrupt, discrete breaks
in parameters that partition a sample into distinct regimes with
different levels \citep[see][among others]{Andrews1993, BaiPerron1998, BaiPerron2003,
HorvathHuskova2012, HorvathHuskovaRiceWang2017,
AntochHanousekHorvathHuskovaWang2019}. This paper argues that for
a broad and economically important class of phenomena, the right
model is not a break but a \emph{kink}, i.e.\ a change in the rate
at which a coefficient evolves.

Kink models have a growing presence in the threshold literature \citep[see, for example,][]{HidalgoLeeSeo2019, ZhongWanZhang2022, SunWanZhangZhong2024, ZhangXieXiao2025}.
\citet{Hansen2017} formalizes estimation and inference for the regression
kink model in time series, where the effect of a threshold variable $x$ on
the response is linear on each side of an unknown threshold $\gamma$ but
continuous at $\gamma$ itself, producing a ``kink'' in the regression
function at that point. \citet{ZhangWangJiang2017} extend this framework to
panel data. These papers model nonlinearity as a function
of \textbf{\emph{where agents are}}: the coefficient path bends when a
covariate crosses some threshold level $\gamma$, and crucially, an agent can
move back and forth across that threshold as the variable fluctuates. This is
distinct again from the regression kink design of \citet{CardLeePeiWeber2015},
where the kink is located in a known running variable at a policy threshold
and the change in slope identifies a causal effect. There the kink is a
feature of the cross-sectional design.

The present paper takes a different and complementary perspective. We model
nonlinearity as a function of \textbf{\emph{when agents are}}. The coefficient
vector is modeled as piecewise linear in time, with multiple unknown kink dates
at which the rate of change shifts permanently. This distinction matters
economically. A central bank may not react because
inflation crossed some level last quarter, but it may do so because of a persistent
shift in its policy framework. \citet{ClaridaGaliGertler2000} document
  a shift around the Volcker appointment in 1979, but such a
transition need not be instantaneous  as
credibility may need to be established. Similarly, \citet{Trefler2004} shows that the
employment and productivity effects of the Canada--US Free Trade Agreement
unfolded over several years after ratification.
Returning to the growth literature,  \citet{Hansen2017} studies the relationship between
public debt and growth studied by \citet{ReinhartRogoff2010}. The question
is whether the effect of debt on growth changes once the debt-to-GDP ratio
crosses some threshold.  In our empirical application we revisit a similar question and find that the bend is dateable and that the relationship was already evolving before that date. This complements his analysis and
emphasizes that when agents are is an important complement
to where agents are in obtaining a fuller picture of the
underlying economic relationship.

Another related literature allows coefficients to evolve smoothly over time. In panel data, however, this literature remains relatively sparse; see \citet{LiChenGao2011}, \citet{SuWangJin2019}, \citet{Chen2019}, \citet{DongGaoPeng2021}, and \citet{SuJin2012}. Besides the when agents are point, a natural question, then, is why one should prefer the kink model for this task over its two closest neighbours: the level break model on one side and the fully nonparametric time-varying parameter (TVP) model on the other. The break model, as in \citet{BaiPerron1998}, is well suited to changes that occur overnight, but it rules out, by assumption, the possibility that the relationship was already evolving before the break date, constraining the within-regime path to be flat and discarding the information about the {speed} at which structural change unfolds. The TVP model imposes no such restriction, since its goal is to recover the function itself rather than to identify a small number of structural change dates. The cost, however, is that it delivers neither a dateable event nor a within-regime slope with a clear economic interpretation. Our model lies between these two extremes. It identifies a dateable kink and estimates how fast the relationship was evolving on each side of that date.  The within-regime slope measures the speed of structural change in a way that neither alternative can provide, making the kink model the natural choice whenever both the timing and the pace of a shifting economic relationship are objects of interest.

The estimator suggested in this paper proceeds in two steps. We
first run an unpenalised regression of the first-period differenced
model, which eliminates the fixed effects while leaving the
time-varying structure of the coefficients \citep[see][for a discussion in related settings]{KaddouraWesterlund2023, Kaddoura2025}, to construct
adaptive weights in the spirit of \citet{Zou2006} that are small
where the coefficient path already appears to bend and large where
it is locally linear. We then minimise a penalised least squares
objective that penalises the second differences of the coefficient
path using a weighted group-$\ell_2$ norm \citep[see][for the group
lasso and its adaptive extension]{YuanLin2006, Zou2006}. Since a
nonzero second difference at period $t$ is equivalent to a kink at $t$,
the penalty directly targets slope changes while enforcing continuity
throughout, distinguishing our approach from level break estimators
that allow jumps. The adaptive weights recover the oracle property in the
sense of \citet{Zou2006}, where the estimated kink dates coincide with the
true ones with probability approaching one, recovering both the
correct number of kinks and their exact locations. Conditional on
correct recovery, a post-kink ordinary least squares regression on
the piecewise linear basis estimates the initial level and
within-regime slopes, delivering asymptotic normality.

The estimation strategy proposed in this paper is related to several
strands of the penalisation literature, though it differs from each
in important ways. At a conceptual level, our use of second differences is reminiscent of the Hodrick--Prescott filter \citep{HodrickPrescott1997}. Our method is not an HP filter, however, since we work in a panel regression setting with observed covariates and fixed effects, and use an adaptive group penalty on second differences to induce sparse kinks. This panel structure, where first-period differencing removes the incidental fixed effect before the penalty is applied, is what separates our setting from the time series filtering and trend filtering literature \citep{KimEtAl2009, TibshiraniRyan2014} and the generalised lasso
of \citet{TibshiraniTaylor2011}. The closest to our setting is
\citet{QianSuET2016} and \citet{QianSu2016} who apply an adaptive
group fused Lasso to detect level breaks by penalising \emph{first}
differences of the coefficient path. We penalise second differences instead,
which induces piecewise linearity rather than jumps and so targets kinks rather
than breaks. Because a slope keeps shifting the
coefficient path in every later regime, it is identified over the whole
subsequent span rather than within a single regime alone.

The contributions of this paper are therefore threefold. First, we are the first to
introduce the kink panel model as a new class of structural change
specifications in which the coefficient path is continuous and
piecewise linear in time, filling the gap between the fully flexible
time-varying parameter literature and the level break literature. To the best of our knowledge, this is the first panel-data structural-change framework in which common unknown kink dates are estimated in a time-varying coefficient path.
Second, we establish selection consistency, exact recovery of the kink dates, and oracle asymptotic normality for the post-kink estimator. Unlike the level-break model of \citet{QianSu2016}, in which each regime coefficient is identified from observations within that regime, continuity in our kink model links adjacent segments and changes the convergence rates of the slopes. Endpoint slopes retain the classical within-regime trend rates, whereas interior slopes borrow strength from neighboring regimes through their shared boundary levels. To the best of our knowledge, this is the first explicit characterization of such adjacent-regime slope rates in a panel model with multiple unknown kinks.   Third, we introduce a
coefficient-by-coefficient extension, in the spirit of
\citet{Kaddoura2025}, in which each regressor is permitted to kink
at its own set of dates independently of the others, accommodating
richer forms of partial structural change in which some coefficients
bend while others remain globally linear. Together, these
contributions provide applied researchers with a toolkit for dating
and estimating gradual, continuous shifts in economic relationships.

The remainder of the paper is organized as follows. Section~\ref{sec:model}
sets out the kink panel model and its key representation. Section~\ref{sec:estimator}
introduces the penalised estimator and the post-kink estimation step.
Section~\ref{asymptotics} develops the asymptotic theory, including
selection consistency, recovery of the kink dates, asymptotic normality
of the post-kink estimator, and the coefficient-by-coefficient
extension. Section~\ref{sec:MC} reports the Monte Carlo evidence, and Section~\ref{sec:debtgrowth}
presents the empirical applications.

\section{Model}\label{sec:model}

Consider a scalar  variable $y_{i,t}\in\mathbb{R}$ and a vector of covariates $\*x_{i,t}:=[x_{i,t,1},\ldots,x_{i,t,p}]'\in\mathbb{R}^{p\times 1}$, observed for cross-sectional units $i=1,\ldots,N$ over time periods $t=1,\ldots,T$, where $p$ is fixed. Consider the following data-generating process (DGP):
\begin{align}
    y_{i,t} = \mu_i + \*x_{i,t}'\+\theta_t + \varepsilon_{i,t},
    \label{eq:DGP}
\end{align}
where $\+\theta_t\in\mathbb{R}^{p\times 1}$ is a conformable vector of unknown coefficients that may vary across time periods in a manner made precise below, $\mu_i\in\mathbb{R}$ is a cross-sectional fixed effect, and $\varepsilon_{i,t}\in\mathbb{R}$ is an idiosyncratic error term. Throughout this paper ``$a:=b$'' indicates that $a$ is defined by $b$.

The object of interest is to estimate $\+\theta_t$ and its  kink behaviour. To allow for kinks, we let $\+\theta_t$ evolve over $K+1$ distinct regimes. Let $1=:T_0<T_1<\cdots<T_K<T<T_{K+1}:=T+1$ and define the $\ell$th regime set as
\(
    \mathcal{K}_\ell := \{T_{\ell-1},\ldots,T_\ell-1\},
    \) for \( \ell=1,\ldots,K+1,
     \)
where $K\in\{0,\ldots,T-2\}$. For each regime $\ell$, let $\+\kappa_\ell\in\mathbb{R}^{p\times 1}$ denote the within-regime slope (in~$t$). We model $\+\theta_t$ as piecewise linear in $t$ via
\begin{align}
    \+\theta_t = \+\theta_{T_{\ell-1}} + (t-T_{\ell-1})\+\kappa_\ell
    \label{eq:Kink}
\end{align}
for $ t\in\mathcal{K}_\ell$ and $\ \ell=1,\ldots,K+1$.
Under~\eqref{eq:Kink}, $\+\theta_{T_{\ell-1}}$ is the level of the coefficient path at the start of regime~$\ell$, and $\+\kappa_\ell$ governs the linear evolution of $\+\theta_t$ within that regime. Kinks occur at periods $t\in\mathcal{T}_K:=\{T_1,\ldots,T_K\}$, where the slope changes across regimes. Since regime~$\ell$ ends at time $T_\ell-1$ and regime~$\ell+1$ starts at time $T_\ell$, we impose continuity at kink locations (no jumps) via the recursion
\begin{align}
    \+\theta_{T_\ell} = \+\theta_{T_{\ell-1}} + (T_\ell-T_{\ell-1})\+\kappa_\ell
    \label{eq:continuity}
\end{align}
for $\ell = 1,\hdots, K$. In particular,~\eqref{eq:continuity} implies $\+\theta_{T_\ell}-\+\theta_{T_\ell-1}=\+\kappa_\ell$,\footnote{
To see this, note that $T_\ell-1\in\mathcal{K}_\ell$, so~\eqref{eq:Kink} gives $\+\theta_{T_\ell-1}=\+\theta_{T_{\ell-1}}+\bigl((T_\ell-1)-T_{\ell-1}\bigr)\+\kappa_\ell$, while~\eqref{eq:continuity} gives $\+\theta_{T_\ell}=\+\theta_{T_{\ell-1}}+(T_\ell-T_{\ell-1})\+\kappa_\ell$. Subtracting yields $\+\theta_{T_\ell}-\+\theta_{T_\ell-1}=\+\kappa_\ell$.}
so the coefficient path does not jump when moving from $T_\ell-1$ to $T_\ell$.



In fact, how we model the parameters is closely related to the literature in which local linear methods are used to approximate smoothly evolving coefficient functions (see, for example, \citealt{FanZhang1999}; \citealt{Cai2007}). The key difference is that, in our setting, piecewise linearity is imposed as a global structural approximation with a set of unknown kink dates, rather than serving as a local approximation to an otherwise smooth coefficient path.\footnote{In the special case where $K=1$, $p=1$,  $x_{i,t,1}=1$,
our model becomes close to the deterministic panel trend-break
model of \citet{Kim2011}.}

\begin{remark}[Unified representation]\label{rem:unified}
Equations~\eqref{eq:Kink}--\eqref{eq:continuity} can be condensed into a single representation. Iterating~\eqref{eq:continuity} yields
\(
    \+\theta_{T_{\ell-1}}
    =
    \+\theta_1
    +
    \sum_{j=1}^{\ell-1}(T_j-T_{j-1})\+\kappa_j \) for \(
     \ell=1,\ldots,K+1,\)
where the sum is empty (hence zero) when $\ell=1$. Substituting this into~\eqref{eq:Kink} gives, for $t\in\mathcal{K}_\ell$,
\begin{align}
    \+\theta_t
    =
    \+\theta_1
    +
    \sum_{j=1}^{\ell-1}(T_j-T_{j-1})\+\kappa_j
    +
    (t-T_{\ell-1})\+\kappa_\ell
    \label{eq:Kinkalt}
\end{align}
for $ \ell=1,\ldots,K+1$. Hence, the entire path $\{\+\theta_t\}_{t=1}^T$ is determined by the initial level $\+\theta_1$, the within-regime slopes $\+\kappa_1,\ldots,\+\kappa_{K+1}$, and the kink dates $\mathcal{T}_K$.
\end{remark}




An equivalent way of expressing the piecewise-linear restriction~\eqref{eq:Kink}--\eqref{eq:continuity} is through first differences. Define $\Delta\+\theta_t:=\+\theta_t-\+\theta_{t-1}$ for $t=2,\ldots,T$ and consider the following proposition.

\begin{proposition}[Piecewise-constant first differences]\label{prop:FD}
Under~\eqref{eq:Kink}--\eqref{eq:continuity}, the first differences are piecewise constant:
\begin{align}
    \Delta\+\theta_t = \+\kappa_\ell
    \label{eq:FDconst}
\end{align}
for $ t\in\{T_{\ell-1}+1,\ldots,\min(T_\ell,\, T)\}$ and $\ \ell=1,\ldots,K+1$.
Conversely,~\eqref{eq:FDconst} implies~\eqref{eq:Kink} and~\eqref{eq:continuity}.
\end{proposition}

\begin{proof}
\emph{Forward direction.} Fix a regime $\ell\in\{1,\ldots,K+1\}$. For $t\in\{T_{\ell-1}+1,\ldots,T_\ell-1\}$, both $t$ and $t-1$ belong to
$\mathcal{K}_\ell$, so~\eqref{eq:Kink} applies to both and
\(
    \Delta\+\theta_t
    = \bigl[\+\theta_{T_{\ell-1}}+(t-T_{\ell-1})\+\kappa_\ell\bigr]
    -\bigl[\+\theta_{T_{\ell-1}}+(t-1-T_{\ell-1})\+\kappa_\ell\bigr]
    = \+\kappa_\ell.
\)
For $\ell\leq K$ and $t=T_\ell$, continuity~\eqref{eq:continuity} combined
with~\eqref{eq:Kink} at $t=T_\ell-1$ gives
$\+\theta_{T_\ell}-\+\theta_{T_\ell-1}=\+\kappa_\ell$, so
$\Delta\+\theta_{T_\ell}=\+\kappa_\ell$ as well. For the last regime
$\ell=K+1$, no boundary $t=T_{K+1}=T+1$ exists since $\Delta\+\theta_t$
is defined only for $t=2,\ldots,T$. This establishes~\eqref{eq:FDconst}.
\emph{Converse direction.} Assume~\eqref{eq:FDconst}. Fix $\ell$ and any
$t\in\mathcal{K}_\ell$. For every $s\in\{T_{\ell-1}+1,\ldots,t\}$, we have
$s\leq t\leq T_\ell-1\leq\min(T_\ell,T)$, with equality in the last bound
only in the final regime $\ell=K+1$, where $T_{K+1}=T+1$ and
$\min(T_\ell,T)=T$. In every case $s\leq\min(T_\ell,T)$, so~\eqref{eq:FDconst}
gives $\Delta\+\theta_s=\+\kappa_\ell$. Telescoping yields
\(
    \+\theta_t
    = \+\theta_{T_{\ell-1}}+\sum_{s=T_{\ell-1}+1}^{t}\Delta\+\theta_s
    = \+\theta_{T_{\ell-1}}+(t-T_{\ell-1})\+\kappa_\ell,
\)
which is~\eqref{eq:Kink}. For $\ell\leq K$, taking $t=T_\ell-1$ in this
identity gives
$\+\theta_{T_\ell-1}=\+\theta_{T_{\ell-1}}+(T_\ell-1-T_{\ell-1})\+\kappa_\ell$,
and combining with $\Delta\+\theta_{T_\ell}=\+\kappa_\ell$
from~\eqref{eq:FDconst} gives
$\+\theta_{T_\ell}=\+\theta_{T_{\ell-1}}+(T_\ell-T_{\ell-1})\+\kappa_\ell$,
which is~\eqref{eq:continuity}.
\end{proof}

An immediate consequence of Proposition~\ref{prop:FD} is that a kink at $T_\ell$ corresponds to a change in slope:
\begin{align}
    \+\kappa_{\ell+1}\neq\+\kappa_\ell
    \Longleftrightarrow
    \|\Delta\+\theta_{T_\ell+1}-\Delta\+\theta_{T_\ell}\|
    = \|\+\kappa_{\ell+1}-\+\kappa_\ell\| > 0.
    \label{eq:kinkdetect}
\end{align}
Hence, estimating kink locations reduces to detecting the times at which the \emph{second differences} $\Delta^2\+\theta_t:=\Delta\+\theta_{t+1}-\Delta\+\theta_t$ are nonzero. This motivates a sparsity assumption on $\{\Delta^2\+\theta_t\}_{t=2}^{T-1}$: most second differences are zero (corresponding to the linear segments), and the nonzero ones identify the kink dates.
Take Figure \ref{fig:kink} as an illustration. The upper panel shows a coefficient path $\+\theta_t$ across three regimes $\mathcal{K}_1,\mathcal{K}_2,\mathcal{K}_3$, while the lower panel shows the corresponding piecewise constant slope $\+\kappa_\ell$, with jumps at the two kink dates.


\bigskip
\begin{figure}[h!]
    \centering    \includegraphics[width=0.8\linewidth]{kink_figure_7.pdf}
    \caption{\footnotesize{Example of the piecewise-linear coefficient path. The upper panel shows $\+\theta_t$ across three regimes $\mathcal{K}_1,\mathcal{K}_2,\mathcal{K}_3$, with kinks at $T_1$ and $T_2$. The lower panel shows the corresponding piecewise-constant first differences $\Delta\+\theta_t$.}}
    \label{fig:kink}
\end{figure}




\bigskip









\begin{remark}[Second-derivative interpretation]\label{rem:secondderiv}
The second difference has a natural continuous time counterpart. For a twice differentiable function $b(\cdot)$, the standard finite-difference approximation gives $b''(t)\approx\bigl[b(t+1)-2b(t)+b(t-1)\bigr]$. In our discrete-time setting, $\Delta^2\+\theta_t=\+\theta_{t+1}-2\+\theta_t+\+\theta_{t-1}$ is the exact analog. Hence, penalising $\sum_{t=2}^{T-1}\|\Delta^2\+\theta_t\|$ can be interpreted as penalising the ``curvature'' of the coefficient path: the penalty is zero when $\+\theta_t$ is globally linear and nonzero only at kink points.
\end{remark}



It is also worth asking what happens if the truth contains a level break rather
than a kink. In fact, the kink model contains the break
model as a class of discrete-time observed sequences. Suppose the coefficient equals \(\+\theta_A\) through \(T^*-1\) and
\(\+\theta_A+\+\alpha\) from \(T^*\) onward, similar to  \citet{BaiPerron1998}.
Although continuity rules out a literal jump in the interpolated path, it does not rule out the same sequence at the observed integer dates. Suppose further that  \(K=2\) with
\(T_1=T^*-1\) and \(T_2=T^*\), \(\+\theta_1=\+\theta_A\),
\(\+\kappa_1=\+\kappa_3=\*0_{p\times1}\) and \(\+\kappa_2=\+\alpha\),
equations~\eqref{eq:Kink} and~\eqref{eq:continuity} reproduce the break path
exactly, with the jump carried by the middle slope. Thus, the observed level break is represented by two adjacent kinks, with slopes \(\*0_{p\times1}\), \(\+\alpha\), and \(\*0_{p\times1}\).
The reason is that the coefficient is only ever observed at integer dates.
Between \(T^*-1\) and \(T^*\) there are no observations, so the data cannot distinguish a literal jump from a steep but continuous movement across that interval. The kink model therefore connects the two observed coefficient levels by a straight line. A change occurring over several sampling periods is gradual, whereas one occurring over a single period is abrupt at the observed sampling frequency.


\begin{comment}
One might wonder what happens if the truth contains a level break
rather than a kink. Suppose the coefficient equals $\+\theta_A$
through $T^*-1$ and $\+\theta_A+\+\delta$ from $T^*$ onward, as in
\citet{BaiPerron1998}. Extended to a continuous time axis as a step
function, this path jumps at $T^*$, and continuity
\eqref{eq:continuity} rules out that interpolation. It does not rule
out the underlying sequence. Take $K=2$ with kink dates
$T_1=T^*-1$ and $T_2=T^*$, so that the middle regime has unit
length, and set $\+\theta_1=\+\theta_A$ with
$\+\kappa_1=\+\kappa_3=\*0_{p\times1}$ and $\+\kappa_2=\+\delta$.
Because the outer slopes are zero, only the middle slope moves the
path, and the unified representation~\eqref{eq:Kinkalt} gives
\begin{align*}
    \+\theta_t
    =\+\theta_A
    +\begin{cases}
    \*0_{p\times1},& t\le T^*-2,\\[2pt]
    (t-T_1)\+\delta=\*0_{p\times1},& t=T^*-1,\\[2pt]
    (T_2-T_1)\+\delta=\+\delta,& t\ge T^*,
    \end{cases}
\end{align*}
since that slope has been active for zero periods at its only
sampled date $T^*-1$ and for one full period thereafter. The two
sequences coincide at every observed date. What separates the two specifications is the interpolation between
sampled dates. The step function applies the entire change at
$T^*$, while the kink accumulates it linearly over $[T^*-1,T^*]$.
The panel is observed only at the integer dates $1,\ldots,T$, so no
data bear on that interval. A level break is therefore a kink
configuration with a regime of unit length, whose slope is the
magnitude of the break.
\end{comment}

\section{Estimator}\label{sec:estimator}


By~\eqref{eq:Kink}, the behaviour of the coefficient path $\+\Theta_T:=[\+\theta_1',\ldots,\+\theta_T']'\in\mathbb{R}^{Tp\times 1}$ is governed by the within-regime slopes $\*K_K:=[\+\kappa_1',\ldots,\+\kappa_{K+1}']'\in\mathbb{R}^{(K+1)p\times 1}$, the initial level $\+\theta_1$, and the kink dates $\mathcal{T}_K$. Conversely, once $\+\theta_1$, $\*K_K$, and $\mathcal{T}_K$ are known, the entire path $\+\Theta_T$ can be recovered via~\eqref{eq:Kinkalt}. We use superscript ``$0$'' to denote true values: $K^0$ is the true number of kinks, $\mathcal{T}^0_{K^0}:=\{T_1^0,\ldots,T_{K^0}^0\}$ the true kink dates, $\*K^0_{K^0}:=[\+\kappa_1^{0\prime},\ldots,\+\kappa_{K^0+1}^{0\prime}]'\in\mathbb{R}^{(K^0+1)p\times 1}$ the true slopes, and $\+\Theta^0_T:=[\+\theta_1^{0\prime},\ldots,\+\theta_T^{0\prime}]'\in\mathbb{R}^{Tp\times 1}$ the true coefficient path.

Two transformations of the data are used to handle the fixed effects
$\mu_i$, one per estimation stage. The selection stage works with
deviations from the initial observation. For each $i$ and $t=2,\ldots,T$,
define $\check{a}_{i,t}:=a_{i,t}-a_{i,1}$, and in particular
$\check{y}_{i,t}:=y_{i,t}-y_{i,1}$ and
$\check{\varepsilon}_{i,t}:=\varepsilon_{i,t}-\varepsilon_{i,1}$. The
post-kink stage instead removes $\mu_i$ by the within transformation. For
each $i$ and $t=1,\ldots,T$, define
$\ddot{a}_{i,t}:=a_{i,t}-T^{-1}\sum_{s=1}^{T}a_{i,s}$, and in particular
$\ddot{y}_{i,t}$ and $\ddot{\varepsilon}_{i,t}$.



\subsection{Penalised estimator}\label{sec:initial}

As a preliminary step, we compute an unpenalised  OLS estimator from the first-period-differenced model:
\begin{align}
    \{\dot{\+\theta}_t\}_{t=1}^T
    := \underset{\{\+\theta_t\}_{t=1}^T}{\arg\min}\;
    \frac{1}{NT}\sum_{i=1}^N\sum_{t=2}^{T}
    \bigl( \check{y}_{i,t} - \*x_{i,t}'\+\theta_t + \*x_{i,1}'\+\theta_1\bigr)^2.
    \label{eq:initial}
\end{align}
Since~\eqref{eq:initial} imposes no piecewise-linear structure, the resulting estimator $\dot{\+\theta}_t$ is $\sqrt{N}$-consistent for $\+\theta_t^0$ under the assumptions stated in Section~\ref{sec:assumptions}, but does not exploit the kink structure.

The initial estimator is used to construct adaptive weights that steer the penalty in the penalised objective function defined below. Specifically, define for all $t=2,\hdots, T-1$
\(
    \dot{\omega}_t
    := \|\Delta\dot{\+\theta}_{t+1}-\Delta\dot{\+\theta}_t\|^{-\zeta_1},
\)
where $\zeta_1>0$ is a user-specified constant (typically $\zeta_1=2$).   This adaptive weighting is essential for the oracle property established later.\footnote{For numerical stability, one may replace
$\|\Delta\dot{\+\theta}_{t+1}-\Delta\dot{\+\theta}_t\|^{-\zeta_1}$
with
$\bigl(\|\Delta\dot{\+\theta}_{t+1}-\Delta\dot{\+\theta}_t\|+\delta_t\bigr)^{-\zeta_1}$,
where $\delta_t>0$ is a small perturbation of order $O(N^{-1/2})$ satisfying
$\delta_t\to 0$ as $N\to\infty$. This prevents near-division-by-zero in
finite samples while leaving the asymptotic properties unchanged.}
Now, given these weights, we propose minimising the following objective function:
\begin{align}
    \mathcal{L}_{\vartheta_1}(\+\Theta_T)
    &:= \frac{1}{NT}\sum_{i=1}^N\sum_{t=2}^{T}
    \bigl( \check{y}_{i,t} - \*x_{i,t}'\+\theta_t + \*x_{i,1}'\+\theta_1\bigr)^2
    +  {\vartheta_1}\sum_{t=2}^{T-1}\dot{\omega}_t\,
    \|\Delta\+\theta_{t+1}-\Delta\+\theta_t\|,
    \label{eq:objective}
\end{align}
where $\vartheta_1=\vartheta_1(N,T)>0$ is a tuning parameter that can be chosen by minimizing an information criterion (see \citealt{QianSu2016} for details). We do this explicitly Sections \ref{sec:MC} and \ref{sec:debtgrowth}. Three features of~\eqref{eq:objective}  are worth commenting on.

\medskip
\noindent\textbf{($i$)~\emph{Loss function.}}\;
The first term is a least-squares loss computed on the first-period-differenced data. Taking first-period differences of~\eqref{eq:DGP} gives $ \check{y}_{i,t}=\*x_{i,t}'\+\theta_t-\*x_{i,1}'\+\theta_1+\check{\varepsilon}_{i,t}$, so the fixed effect $\mu_i$ cancels. The initial level $\+\theta_1$ enters through the term $\*x_{i,1}'\+\theta_1$.

\bigskip

\noindent\textbf{($ii$)~\emph{Penalty}.}\;
The second term penalises the second differences $\Delta^2\+\theta_t=\Delta\+\theta_{t+1}-\Delta\+\theta_t$ using the group-$\ell_2$ norm $\|\cdot\|$. By~\eqref{eq:kinkdetect}, $\|\Delta^2\+\theta_t\|>0$ if and only if there is a kink at~$t$. The group-$\ell_2$ structure encourages the \emph{entire} $p\times 1$ vector $\Delta^2\+\theta_t$ to be set to zero simultaneously, thereby inducing piecewise linearity in the coefficient path. The sum runs over $t=2,\ldots,T-1$ because $\Delta^2\+\theta_t=\+\theta_{t+1}-2\+\theta_t+\+\theta_{t-1}$ requires $t-1\geq 1$ and $t+1\leq T$.

\bigskip

\noindent\textbf{($iii$)~\emph{Adaptive weights}.}\;
The weights $\dot{\omega}_t$  make the penalty adaptive. At true kink dates, the penalty is weak (small $\dot{\omega}_t$), so the nonzero second differences survive; at non-kink dates, the penalty is strong (large $\dot{\omega}_t$), so $\Delta^2\+\theta_t$ is shrunk to $\*0_{p\times 1}$ exactly. This is the mechanism by which the estimator simultaneously detects kink locations and estimates the coefficient path.

\medskip

The penalised estimator of $\+\Theta_T^0$ is then given by
\begin{align}
    \widehat{\+\Theta}_T
    :=
    \begin{bmatrix}
        \widehat{\+\theta}_1 \\ \vdots \\ \widehat{\+\theta}_T
    \end{bmatrix}
    =
    \underset{\+\Theta_T}{\arg\min}\;\mathcal{L}_{\vartheta_1}(\+\Theta_T).
    \label{eq:penest}
\end{align}
From the penalised estimator $\widehat{\+\Theta}_T$, we extract the estimated kink dates and regime sets as follows.\footnote{Demeaning would also eliminate $\mu_i$ and is linear in
$\+\Theta_T$, so it could serve at this stage too. However, under the within
transformation the regressor attaching to $\+\theta_s$ at observation
$(i,t)$ is $\*x_{i,t}\*1\{s=t\}-T^{-1}\*x_{i,s}$, so every date enters
every equation which leads to a matrix of
order $Tp\times Tp$. Differencing against $t=1$ instead  reduces the  dimension to
$2p\times2p$.}

\begin{definition}[Estimated kinks and regimes]\label{def:estkinks}
For a given $\widehat{\+\Theta}_T$, define the set of estimated kink locations as
\begin{align}
    \widehat{\mathcal{T}}_{\widehat{K}}
    := \left\{t\in\{2,\ldots,T-1\}:\|\Delta\widehat{\+\theta}_{t+1}-\Delta\widehat{\+\theta}_t\|>0\right\}
    = \left\{\widehat{T}_1,\ldots,\widehat{T}_{\widehat{K}}\right\},
    \label{eq:estkinks}
\end{align}
where $\widehat{T}_1<\cdots<\widehat{T}_{\widehat{K}}$ are the ordered elements and $\widehat{K}:=|\widehat{\mathcal{T}}_{\widehat{K}}|$ is the estimated number of kinks. By convention, set $\widehat{T}_0:=1$ and $\widehat{T}_{\widehat{K}+1}:=T+1$. These kink dates partition $\{1,\ldots,T\}$ into $\widehat{K}+1$ estimated regime sets:
\(
    \widehat{\mathcal{K}}_\ell
    :=  \left\{\widehat{T}_{\ell-1},\ldots,\widehat{T}_\ell-1\right\}
\) for $\ell=1,\ldots,\widehat{K}+1.$
If $\|\Delta\widehat{\+\theta}_{t+1}-\Delta\widehat{\+\theta}_t\|=0$ for all $t=2,\ldots,T-1$, then $\widehat{K}=0$ and $\widehat{\mathcal{T}}_0=\varnothing$, corresponding to a single linear regime spanning all periods.
\end{definition}


\subsection{Post kink estimator}\label{sec:postkink}
Recall from Definition~\ref{def:estkinks} that
$\widehat{\mathcal{T}}_{\widehat{K}}=\{\widehat{T}_1,\ldots,\widehat{T}_{\widehat{K}}\}$, with
$\widehat{T}_0:=1$ and $\widehat{T}_{\widehat{K}+1}:=T+1$, and that the
induced regime sets are
$\widehat{\mathcal{K}}_\ell:=\{\widehat{T}_{\ell-1},\ldots,
\widehat{T}_\ell-1\}$ for $\ell=1,\ldots,\widehat{K}+1$. Holding $\widehat{\mathcal{T}}_{\widehat{K}}$
fixed, the unified representation in Remark~\ref{rem:unified} gives for $ t\in\widehat{\mathcal{K}}_\ell$
\begin{align}
    \+\theta_t
    =
    \+\theta_1
    +\sum_{j=1}^{\ell-1}\bigl(\widehat{T}_j-\widehat{T}_{j-1}\bigr)\+\kappa_j
    +\bigl(t-\widehat{T}_{\ell-1}\bigr)\+\kappa_\ell,
    \label{eq:pk_rep}
\end{align}
where $\+\theta_1$ is the initial level and $\+\kappa_\ell$ the slope in
regime $\widehat{\mathcal{K}}_\ell$. The continuity
conditions~\eqref{eq:continuity} are already embedded
in~\eqref{eq:pk_rep}, so the entire $Tp$-dimensional path $\+\Theta_T$ is
determined by the $(\widehat{K}+2)p$-dimensional  parameter
vector
\(
    \+\psi_{\widehat{K}}
    := [\+\theta_1', \*K_{\widehat{K}}']' =
    \bigl[\+\theta_1',\,\+\kappa_1',\ldots,\+\kappa_{\widehat{K}+1}'\bigr]'.
\)
To express~\eqref{eq:pk_rep} as a function of $\+\psi_{\widehat{K}}$,
define for each $t\in\widehat{\mathcal{K}}_\ell$ and each
$j=1,\ldots,\widehat{K}+1$ the scalar \emph{basis loading}
\begin{align}
    d_{j,t}\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)
    :=
    \begin{cases}
        \widehat{T}_j-\widehat{T}_{j-1} & j < \ell, \\[4pt]
        t-\widehat{T}_{\ell-1}           & j = \ell, \\[4pt]
        0                                & j > \ell,
    \end{cases}
    \label{eq:djt}
\end{align}
where $\ell$ is the unique regime with $t\in\widehat{\mathcal{K}}_\ell$.
The loading measures how long the slope $\+\kappa_j$ has been active by
period $t$, i.e. the full length $\widehat{T}_j-\widehat{T}_{j-1}$ of
$\widehat{\mathcal{K}}_j$ once that regime has ended, the time elapsed
since the current regime began while $t$ lies within it, and zero before
it starts. With this notation~\eqref{eq:pk_rep} takes the compact form
\begin{align}
    \+\theta_t
    =
    \+\theta_1
    +\sum_{j=1}^{\widehat{K}+1}d_{j,t}\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)\,\+\kappa_j
    \label{eq:pk_loading}
\end{align}
for each $t=1,\hdots T$ and  is linear in $\+\psi_{\widehat{K}}$. Collecting the loadings into
the $(\widehat{K}+2)p\times 1$ regressor
\(
    \+\phi_{i,t}\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)
    :=
    \bigl[
        \*x_{i,t}',\,
        d_{1,t}\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)\,\*x_{i,t}',\ldots,
        d_{\widehat{K}+1,t}\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)\,\*x_{i,t}'
    \bigr]'
\)
gives $\*x_{i,t}'\+\theta_t=\+\phi_{i,t}\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)'\+\psi_{\widehat{K}}$
for every $t$, so that the DGP~\eqref{eq:DGP} reads for each $i=1,\ldots,N$ and
$t=1,\ldots,T,$
\begin{align}
     y_{i,t}
    =
    \mu_i
    +\+\phi_{i,t}\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)'\+\psi_{\widehat{K}}
    +\varepsilon_{i,t},
    \label{eq:pk_expanded}
\end{align}
a standard linear panel regression in which the coefficient vector
$\+\psi_{\widehat{K}}$ is common to all $i$ and $t$ and the time
variation of the path is carried entirely by the regressor.

It remains to remove $\mu_i$, and here the choice of transformation
matters in a way it did not for selection. The penalised problem of
Section~\ref{sec:initial} differences against the initial period, which
is enough to recover $\widehat{\mathcal{T}}_{\widehat{K}}$ but subtracts the single equation $t=1$ from
every other, returning the  error $\varepsilon_{i,1}$ in all
$T-1$ rows, where it accumulates with the time index. The post-kink
stage subtracts the unit mean instead, which spreads the transformation
evenly across periods and leaves a transformed error whose covariance
remains bounded in $T$. Averaging the
DGP~\eqref{eq:DGP} over $t$ for fixed $i$ and subtracting gives
\begin{align}
    \ddot y_{i,t}
    =
    \*x_{i,t}'\+\theta_t
    -\dfrac{1}{T}\sum_{s=1}^{T}\*x_{i,s}'\+\theta_s
    +\ddot\varepsilon_{i,t},
    \label{eq:pk_demeanraw}
\end{align}
and the subtracted term carries a different coefficient $\+\theta_s$ at
every date, so the equation for period $t$ involves the entire path
$\+\theta_1,\ldots,\+\theta_T$ rather than $\+\theta_t$ alone.  However, every coefficient in the sum is the same linear
function of $\+\psi_{\widehat{K}}$. Substituting
$\*x_{i,s}'\+\theta_s=\+\phi_{i,s}\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)'\+\psi_{\widehat{K}}$ at
every date $s$, the common vector factors out of the sum,
\begin{align}
    \dfrac1T\sum_{s=1}^{T}\*x_{i,s}'\+\theta_s
   & =
    \dfrac1T\sum_{s=1}^{T}\+\phi_{i,s}\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)'\+\psi_{\widehat{K}}
    \notag \\ & =
    \Bigl(\frac1T\sum_{s=1}^{T}\+\phi_{i,s}\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)\Bigr)'
    \+\psi_{\widehat{K}}
     \notag \\ &=
    \bar{\+\phi}_i\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)'\+\psi_{\widehat{K}},
    \label{eq:pk_factor}
\end{align}
so the path collapses to the single structural vector
and~\eqref{eq:pk_demeanraw} becomes
\begin{align}
    \ddot y_{i,t}
    =
    \ddot{\+\phi}_{i,t}\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)'\+\psi_{\widehat{K}}
    +\ddot\varepsilon_{i,t}
    \label{eq:pk_linear}
\end{align}
 for all $i=1,\ldots,N$ and $t=1,\ldots,T,$ a regression on the same parameter vector as~\eqref{eq:pk_expanded},
with no coefficients of other periods entering. The transformation acts
on the regressor alone.

In matrix form, stack the observations as
$\*y:=[y_{1,1},\ldots,y_{1,T},\ldots,y_{N,T}]'\in\mathbb{R}^{NT\times1}$,
let $\+\Phi\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)\in\mathbb{R}^{NT\times(\widehat{K}+2)p}$ collect
the rows $\+\phi_{i,t}\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)'$ in the same order, and let
$\+\varepsilon$ be conformable. The transformation
in~\eqref{eq:pk_linear} is then the centering matrix
\(
    \*C
    :=\*I_N\otimes\*C_T,
    \) with \(
    \*C_T:=\*I_T-\frac1T\+{\iota}_T\+{\iota}_T',
\)
which is symmetric and idempotent, demeans each unit's $T$ observations,
and satisfies $\*C(\*I_N\otimes\+{\iota}_T)=\*0$, so that the fixed effects are
annihilated. The post-kink estimator of $\+\psi_{\widehat{K}}$ is the OLS
solution in~\eqref{eq:pk_linear},
\begin{align}
\widetilde{\+\psi}_{\widehat{K}}
&:=
\underset{\+\psi_{\widehat{K}}\in\mathbb{R}^{(\widehat{K}+2)p}}{\arg\min}
\;\frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T}
\bigl(\ddot y_{i,t}
-\ddot{\+\phi}_{i,t}\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)'\+\psi_{\widehat{K}}\bigr)^2
\notag\\
&=
\underset{\+\psi_{\widehat{K}}\in\mathbb{R}^{(\widehat{K}+2)p}}{\arg\min}
\;\frac{1}{NT}
\bigl(\*y-\+\Phi\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)\+\psi_{\widehat{K}}\bigr)'\*C
\bigl(\*y-\+\Phi\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)\+\psi_{\widehat{K}}\bigr)
\notag\\
&=
\bigl[\+\Phi\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)'\*C\,\+\Phi\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)\bigr]^{-1}
\+\Phi\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)'\*C\,\*y,
    \label{eq:postest}
\end{align}
where the second equality holds by symmetry and idempotency of $\*C$, and the third
holding under Assumption~\ref{ass:rank} below. Partitioning
$\widetilde{\+\psi}_{\widehat K}=[\widetilde{\+\theta}_1',\,
\widetilde{\+\kappa}_1',\ldots,\widetilde{\+\kappa}_{\widehat K+1}']'$
and defining the $p\times(\widehat K+2)p$ loading matrix
\(
    \*L_t\bigl(\widehat{\mathcal{T}}_{\widehat{K}}\bigr)
    :=\bigl[\*I_p,\;
    d_{1,t}\bigl(\widehat{\mathcal{T}}_{\widehat{K}}\bigr)\*I_p,\ldots,
    d_{\widehat K+1,t}\bigl(\widehat{\mathcal{T}}_{\widehat{K}}\bigr)\*I_p\bigr],
\)
the post-kink coefficient path for each $t=1,\ldots,T$  follows from~\eqref{eq:pk_loading} as
\begin{align}
    \widetilde{\+\theta}_t
    :=\*L_t\bigl(\widehat{\mathcal{T}}_{\widehat{K}}\bigr)\widetilde{\+\psi}_{\widehat K}
    =\widetilde{\+\theta}_1
    +\sum_{j=1}^{\widehat K+1}
    d_{j,t}\bigl(\widehat{\mathcal{T}}_{\widehat{K}}\bigr)\,\widetilde{\+\kappa}_j
    \label{eq:pk_path}
\end{align}


\begin{remark}[The regressor matrix]
\label{rem:structure}
Sorting rows by regime, the raw design is block lower triangular,
\begin{align*}
\+\Phi\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)
\;=\;
\begin{bmatrix}
\*D_{1,\theta}   & \*D_{1,1}   & \*0         & \cdots & \*0 \\[4pt]
\*D_{2,\theta}   & \*D_{2,1}   & \*D_{2,2}   & \cdots & \*0 \\[4pt]
\vdots           & \vdots      & \vdots      & \ddots & \vdots \\[4pt]
\*D_{\widehat{K}+1,\theta} &
\*D_{\widehat{K}+1,1} &
\*D_{\widehat{K}+1,2} &
\cdots &
\*D_{\widehat{K}+1,\widehat{K}+1}
\end{bmatrix},
\end{align*}
with row block $\ell$ collecting the observations $(i,t)$,
$t\in\widehat{\mathcal{K}}_\ell$, and
\(
    \*D_{\ell,\theta}
    := \bigl[\*x_{i,t}'\bigr]_{i,\,t\in\widehat{\mathcal{K}}_\ell},
\)
\(
    \*D_{\ell,j}
    := \bigl[d_{j,t}\big(\widehat{\mathcal{T}}_{\widehat{K}}\big)\,\*x_{i,t}'\bigr]_{i,\,
    t\in\widehat{\mathcal{K}}_\ell}
\)
for $j=1,\ldots,\ell$. The blocks above the diagonal vanish because
slopes from regimes not yet begun contribute nothing at $t$. The
diagonal block carries the loading $t-\widehat{T}_{\ell-1}$, whose
variation in $t$ separates $\+\kappa_\ell$ from the level $\+\theta_1$,
and the blocks below it carry the frozen loading
$d_{j,t}=\widehat{T}_j-\widehat{T}_{j-1}$, through which an earlier
slope keeps shifting the path in every later regime.
\end{remark}

\section{Asymptotic results}\label{asymptotics}

We first introduce some notation. If $\*A$ is a matrix,
$\mu_{\min}(\*A)$ and $\mu_{\max}(\*A)$ denote its smallest and largest
singular values (its smallest and largest eigenvalues when $\*A$ is
symmetric), $\mathrm{tr}\,\*A$ its trace,
$\|\*A\|:=\sqrt{\mathrm{tr}\,\*A'\*A}$ its Frobenius norm, and
$\|\*A\|_{op}:=\mu_{\max}(\*A)$ its operator norm, so that
$\|\*A\|_{op}\le\|\*A\|$. For two symmetric matrices $\*A$ and $\*B$, we
write $\*A\preceq\*B$ if $\*B-\*A$ is positive semidefinite. All limits
are taken as $N,T\to\infty$ jointly unless stated otherwise, and we
abbreviate $\lim_{N,T\to\infty}$ to $\lim$ where no confusion arises.
For a sequence of random variables or vectors $\{Z_N\}$, we write
$Z_N=O_p(a_N)$ if $Z_N/a_N$ is bounded in probability, i.e.\ for every
$\varepsilon>0$ there exists $M<\infty$ such that
$\sup_N\mathbb{P}(|Z_N/a_N|>M)<\varepsilon$, and $Z_N=o_p(a_N)$ if
$Z_N/a_N\to_p0$. For two positive sequences $\{a_N\}$ and $\{b_N\}$, we
write $a_N\asymp b_N$ if there exist finite constants $0<c\le C<\infty$
such that $c\le a_N/b_N\le C$ for all $N$ sufficiently large. For
positive scalars $a$ and $b$, we write $a\wedge b:=\min\{a,b\}$. The
terms w.p.1, w.p.a.1, $\to_d$, $\to_p$, $\mathcal{N}(\cdot,\cdot)$, and
$\*0_{r\times1}$ signify with probability one, with probability
approaching one, convergence in distribution, convergence in
probability, a normal distribution, and an $r\times1$ vector of zeros,
respectively.


\subsection{Assumptions for selection}\label{sec:assumptions}

Let $I_\ell^0:=T_\ell^0-T_{\ell-1}^0$ denote the length of regime
$\ell$ for $\ell=1,\ldots,K^0+1$, and let
$I_{\min}^0:=\min_{1\le\ell\le K^0+1}I_\ell^0$ be the length of the
shortest regime, with the conventions $I_0^0:=0$ and
$I_{K^0+2}^0:=0$, which will be convenient at the sample endpoints.
When $K^0\ge1$, the relevant signal for selection is
\(
    J_{\min}
    :=\min_{1\le\ell\le K^0}
    \bigl\|\+\kappa^0_{\ell+1}-\+\kappa^0_\ell\bigr\|,
\)
the smallest change in slope at a true kink date or, equivalently, the smallest nonzero second difference of
the true coefficient path. This is the quantity that the adaptive
penalty must distinguish from the estimation error in the unrestricted
path. Because selection works with deviations from the initial
observation, the transformed error is
$\check\varepsilon_{i,t}=\varepsilon_{i,t}-\varepsilon_{i,1}$ for
$t=2,\ldots,T$, and the following moment conditions are stated in terms
of it.

\begin{assumption}[Selection: errors and regressors]\label{ass:errors}
\leavevmode
\begin{enumerate}[label=(\alph*)]
    \item We have that $\mathbb{E}(\varepsilon_{i,t}\*x_{i,s})
    =\*0_{p\times1}$ for all $i$, $t$ and $s$;
    \item We have that for any $N$ and $T$,
\(
    \sup_{2\le t\le T}\frac1N\sum_{i=1}^N\sum_{j=1}^N
    \bigl|\mathbb{E}(\check{\varepsilon}_{i,t}\*x_{i,t}'
    \*x_{j,t}\check{\varepsilon}_{j,t})\bigr|\le C
\)
and
\(
    \sup_{2\le t\le T}\frac1N\sum_{i=1}^N\sum_{j=1}^N
    \bigl|\mathbb{E}(\check{\varepsilon}_{i,t}\*x_{i,1}'
    \*x_{j,1}\check{\varepsilon}_{j,t})\bigr|\le C
\)
for some finite constant $C$;
  \item We have that for any $N$ and $T$,
    \(
        \frac1N\sum_{i=1}^N\sum_{j=1}^N
        \mathbb{E}\!\left(\*x_{i,1}'\*x_{j,1}
        \Bigl(\sum_{t=2}^T\check{\varepsilon}_{i,t}\Bigr)
        \!\Bigl(\sum_{s=2}^T\check{\varepsilon}_{j,s}\Bigr)\right)
        =O(T^2).
    \)
\end{enumerate}
\end{assumption}

Assumption~\ref{ass:errors} controls the score of the
first-period-differenced least squares criterion. Part (a)
imposes strict exogeneity, and parts (b) and (c) permit at most weak
cross-sectional dependence together with heteroskedasticity of
unrestricted form (see \citealt{KaddouraWesterlund2023};
\citealt{Kaddoura2025}). Part (c) is specific to first-period
differencing. The reference error $\varepsilon_{i,1}$ is repeated in
all $T-1$ transformed equations of unit $i$, so its cumulative score
contribution may grow with $T$. The strict
exogeneity in part (a) can be relaxed by replacing the penalised least
squares objective with a GMM-type analog as in \citet{QianSu2016}.
The preliminary estimator of \eqref{eq:initial} must also identify
$\+\theta_1^0$ and each $\+\theta_t^0$ well enough for the adaptive
weights to be informative. For $t=2,\ldots,T$, let
$\*z_{i,t}:=[\*x_{i,1}',(\*x_{i,t}-\*x_{i,1})']'$ and
$\*G_t:=N^{-1}\sum_{i=1}^N\*z_{i,t}\*z_{i,t}'$, the joint second moment
matrix of the initial-period regressor and its deviation at date $t$.

\begin{assumption}[Selection: identification and rank]\label{ass:rank}
There exist constants $0<\underline c\le\overline c<\infty$ such that,
w.p.1,
\(
    \underline c
    \le\inf_{2\le t\le T}\mu_{\min}(\*G_t)
    \le\sup_{2\le t\le T}\mu_{\max}(\*G_t)
    \le\overline c.
\)
\end{assumption}

The lower bound in Assumption~\ref{ass:rank} ensures that $\+\theta_1$
and $\+\theta_t$ are jointly identified in every differenced
cross-section, which is what the unrestricted regression
\eqref{eq:initial} estimates, while the upper bound prevents the design
from becoming arbitrarily ill scaled. The condition concerns the
selection stage only. The tuning parameter has two tasks. It must be small enough at the true
kink dates that a genuine change in slope survives the shrinkage, and
large enough at the remaining dates that a second difference generated
only by sampling noise is set exactly to zero.

\begin{assumption}[Selection: signal strength and tuning parameter]
\label{ass:signal}
\leavevmode
\begin{enumerate}[label=(\alph*)]
    \item $\sqrt{NTK^0}\,\vartheta_1 J_{\min}^{-\zeta_1}\to_p
    c_1\in[0,\infty)$ and $\sqrt{N/K^0}\,J_{\min}\to\infty$ as
    $N,T\to\infty$;
    \item $\dfrac{N^{(\zeta_1+1)/2}\,\vartheta_1}{T^{1+\zeta_1/2}}
    \to_p\infty$ as $N,T\to\infty$.
\end{enumerate}
\end{assumption}

Assumption~\ref{ass:signal} is similar in spirit to the corresponding
condition in \citet{QianSu2016}, to which we refer for further
discussion. The first clause of part (a) and part (b) squeeze
$\vartheta_1$ from both sides, small enough that the penalty vanishes
at the true kinks and large enough that it diverges at the non-kinks.
The power $T^{1+\zeta_1/2}$ in part (b) exceeds the corresponding
requirement in \citet{QianSu2016} because a slope loads on
$d_{j,t}=t-T_{j-1}^0$ with $\sum_t d_{j,t}^2$ of order up to $T^3$,
rather than the order $T$ of the indicator loading in the flat break
case; at fixed $T$, part (b) reduces to their Assumption A.2(iii). The
second clause of part (a) keeps the adaptive weights bounded at every
true kink, and permits $J_{\min}$ to be fixed or to shrink at any rate
slower than $\sqrt{K^0/N}$. No separate detection condition is needed:
combining parts (a) and (b) yields $\sqrt{N/T}\,J_{\min}\to\infty$
whenever a tuning parameter satisfying both parts exists, so any kink
small enough to escape detection is also too small for the tuning
window to accommodate.\footnote{By part (a), w.p.a.1
$\vartheta_1\le(c_1+1)J_{\min}^{\zeta_1}(NTK^0)^{-1/2}$, which is
compatible with part (b) only if
$J_{\min}^{\zeta_1}N^{\zeta_1/2}T^{-(3+\zeta_1)/2}(K^0)^{-1/2}
\to\infty$. Multiplying by $T^{3/2}\sqrt{K^0}\ge1$ gives the claim.}
Parts (a) and (b) are compatible whenever
$(\sqrt N\,J_{\min})^{\zeta_1}\gg T^{(3+\zeta_1)/2}\sqrt{K^0}$. At
$\zeta_1=2$ with $K^0$ fixed and $J_{\min}$ bounded away from zero,
this reduces to $N/T^{5/2}\to\infty$. The constant
$\zeta_1>0$ enters through the adaptive weights and is set to
$\zeta_1=2$ in practice (see, e.g., \citealt{QianSu2016};
\citealt{Kaddoura2025}).\footnote{When $K^0=0$, the coefficient path is
globally linear and $J_{\min}$ is not needed; part (a) is replaced by
the single requirement $\vartheta_1\to_p0$.}

\subsection{Selection consistency}
\label{sec:selectionasymptotics}

The selection argument proceeds in three steps. We first control the
penalised estimator of the unrestricted coefficient path, which bounds
the error in its second differences. We then show that the adaptive
penalty sets every second difference outside the true kink set exactly
to zero. The minimum-signal condition finally guarantees that no true
kink is lost, and exact recovery follows.

\begin{theorem}[Consistency of the penalised estimator]
\label{thm:consistency}
Suppose that Assumptions~\ref{ass:errors}, \ref{ass:rank}
and~\ref{ass:signal}(a) hold. Then, as $N,T\to\infty$,
\begin{enumerate}[label=\textnormal{(\alph*)}]
\item $\widehat{\+\theta}_1-\+\theta_1^0=O_p\!\left(N^{-1/2}\right)$;
\item $\dfrac1T\displaystyle\sum_{t=1}^T
      \bigl\|\widehat{\+\theta}_t-\+\theta_t^0\bigr\|^2
      =O_p\!\left(N^{-1}\right)$;
\item $\displaystyle\max_{1\le t\le T}
      \bigl\|\widehat{\+\theta}_t-\+\theta_t^0\bigr\|
      =O_p\!\left(\sqrt{T/N}\right)$.
\end{enumerate}
\end{theorem}

Theorem~\ref{thm:consistency} establishes three rates for the
penalised estimator. Part (a) gives the $\sqrt N$ rate for the initial
level, part (b) the $N^{-1}$ rate for the path on average over $t$, and
part (c) the uniform rate $\sqrt{T/N}$. Part (c) bounds the largest
deviation along the whole path and should not be read as a pointwise
$\sqrt N$ rate at each date. The reason is the second difference
penalty: a slope shift at one date moves the path in every later
period, so the penalty loads cumulatively and the dates cannot be
separated before the kink structure is known. These are rates for the
unrestricted $Tp$-dimensional selection problem; the faster rates
available after refitting the selected model are the subject of
Section~\ref{sec:postselection}.\footnote{These rates can be improved upon.
Under cross-sectional independence and sub-Gaussian errors, a maximal
inequality should give $O_p(\sqrt{\log T/N})$ instead, weakening the
requirement at $\zeta_1=2$ from $N/T^{5/2}\to\infty$ to
$N/(T^{3/2}\log T)\to\infty$.} The theorem says that the estimated
path is steered towards the truth, but it does not by itself establish
that the estimated kink dates coincide with the true ones, which is the
content of the next result.

\begin{theorem}[Elimination of false kinks]
\label{thm:sign}
Suppose that Assumptions~\ref{ass:errors}, \ref{ass:rank}
and~\ref{ass:signal} hold. Then, as $N,T\to\infty$,
\begin{align*}
    \mathbb{P}\!\left(\left\{
      \left\|\Delta\widehat{\+\theta}_{t+1}
        -\Delta\widehat{\+\theta}_t\right\|=0
        \;\;\forall\;t\in\mathcal{T}^{0c}_{K^0}
   \right\}\right)\to1,
\end{align*}
where
$\mathcal{T}^{0c}_{K^0}:=\{2,\ldots,T-1\}\setminus\mathcal{T}^0_{K^0}$.
\end{theorem}

Theorem~\ref{thm:sign} establishes one half of the recovery argument,
namely that $\widehat{\mathcal T}_{\widehat K}\subseteq
\mathcal T^0_{K^0}$ w.p.a.1. The other half follows from
Theorem~\ref{thm:consistency} and the triangle inequality. At a true
kink date $t\in\mathcal T^0_{K^0}$,
\(
    \bigl\|\Delta^2\widehat{\+\theta}_t\bigr\|
    \ge\bigl\|\Delta^2\+\theta_t^0\bigr\|
    -\bigl\|\Delta^2\bigl(\widehat{\+\theta}_t-\+\theta_t^0\bigr)\bigr\|
    \ge J_{\min}
    -4\max_{1\le s\le T}
    \bigl\|\widehat{\+\theta}_s-\+\theta_s^0\bigr\|.
\)
The maximum is $O_p(\sqrt{T/N})$ by Theorem~\ref{thm:consistency}(c),
and $\sqrt{N/T}\,J_{\min}\to\infty$ under Assumption~\ref{ass:signal},
so the right side is strictly positive
w.p.a.1 and every true kink date survives. Combining the two halves
gives the following corollary.

\begin{corollary}[Exact recovery of the kink structure]
\label{cor:kinks}
Suppose that the assumptions of Theorem~\ref{thm:sign} hold. Then, as
$N,T\to\infty$,
\begin{enumerate}
    \item $\mathbb{P}\!\left(\left\{\widehat K=K^0\right\}\right)\to1$;
    \item $\mathbb{P}\!\left(\left\{\widehat T_1=T_1^0,\ldots,
    \widehat T_{K^0}=T^0_{K^0}
    \;\Big|\;\widehat K=K^0\right\}\right)\to1$.
\end{enumerate}
\end{corollary}

Corollary~\ref{cor:kinks} is the oracle property of the estimator. It
recovers both the correct number of kinks and their exact locations
w.p.a.1. Consequently, for all $N$ and $T$ large enough, the post-kink
regression~\eqref{eq:postest} is run on the correct design matrix
$\+\Phi(\mathcal{T}^0_{K^0})$, and we may now study its distribution in the next subsection.

\subsection{Post-selection asymptotics}
\label{sec:postselection}

Now that exact recovery has been established, we turn to the post-kink
estimator of Section~\ref{sec:postkink}. On the event
$\{\widehat{\mathcal T}_{\widehat K}=\mathcal T^0_{K^0}\}$, whose
probability tends to one by Corollary~\ref{cor:kinks}, the feasible
estimator coincides with the oracle estimator that treats the kink
dates as known. The distribution theory can therefore be developed for the
oracle regression, which is ordinary least squares in a
$(K^0+2)p$-dimensional panel regression, and transferred to the
feasible estimator at the end.

The estimator is computed in the cumulative coordinates
\(\+\psi_{K^0}^0
=[\+\theta_1^{0\prime},\+\kappa_1^{0\prime},\ldots,
\+\kappa_{K^0+1}^{0\prime}]'\),
but these are not the coordinates in which the rates are read. For
\(t\in\mathcal K_\ell^0\), the coefficient path satisfies
\(\+\theta_t^0
=\+\theta_{T_{\ell-1}^0}^0+(t-T_{\ell-1}^0)\+\kappa_\ell^0\),
so regime \(\ell\) is directly informative about its starting level
\(\+\theta_{T_{\ell-1}^0}^0
=\+\theta_1^0+\sum_{j=1}^{\ell-1}I_j^0\+\kappa_j^0\)
and its own slope, but not about the individual terms of the sum,
which are separated only by the early regimes themselves. The
precision of a cumulative coordinate therefore depends on the full
configuration of regime lengths. We therefore reparameterise the
model in terms of the boundary levels, which are exactly what the regimes identify.
Each is informative only through its two adjacent regimes, so each
carries its own rate, and the within-regime slopes are recovered as
scaled differences of consecutive boundary levels.

To that end, we consider the following reparameterisation. For
$j=0,\ldots,K^0+1$, define the boundary levels
\begin{align}
    \+\eta_j^0
    :=\+\theta_1^0+\sum_{m=1}^{j}I_m^0\+\kappa_m^0,
    \label{eq:knotdef}
\end{align}
where the sum is empty when $j=0$, so that $\+\eta_0^0=\+\theta_1^0$.
Each interior $\+\eta_j^0$ is the level of the coefficient path at the
$j$th kink date. Indeed, $T_0^0=1$, so the continuity
recursion~\eqref{eq:continuity} at $\ell=1$ gives
\(
    \+\theta_{T_1^0}^0
    =\+\theta_1^0+\bigl(T_1^0-T_0^0\bigr)\+\kappa_1^0
    =\+\theta_1^0+I_1^0\+\kappa_1^0
    =\+\eta_1^0,\)
and if $\+\theta_{T_{j-1}^0}^0=\+\eta_{j-1}^0$ for some $j\le K^0$,
then~\eqref{eq:continuity} at $\ell=j$ gives
\begin{align}
    \+\theta_{T_j^0}^0
    =\+\theta_{T_{j-1}^0}^0+I_j^0\+\kappa_j^0
    =\+\eta_{j-1}^0+I_j^0\+\kappa_j^0
    =\+\eta_j^0,
    \label{eq:etainduction}
\end{align}
so that $\+\eta_j^0=\+\theta_{T_j^0}^0$ for $j=1,\ldots,K^0$ by
induction. The final coordinate is the endpoint obtained by extending
the last linear segment one period beyond the sample. To see this, note
that $T\in\mathcal K_{K^0+1}^0$, so~\eqref{eq:Kink} and
$I_{K^0+1}^0=T_{K^0+1}^0-T_{K^0}^0=T+1-T_{K^0}^0$ give
\begin{align}
    \+\eta_{K^0+1}^0-\+\theta_T^0
    &=\bigl[\+\eta_{K^0}^0+I_{K^0+1}^0\+\kappa_{K^0+1}^0\bigr]
    -\bigl[\+\eta_{K^0}^0
    +\bigl(T-T_{K^0}^0\bigr)\+\kappa_{K^0+1}^0\bigr]
    \notag\\
    &=\+\kappa_{K^0+1}^0.
    \label{eq:etaendpoint}
\end{align}
It is not a new parameter, since $T_{K^0}^0<T$ implies
$I_{K^0+1}^0\ge2$, so the last slope, and with it the endpoint, is
identified from the observed part of the final regime. Including it
makes every regime an interval between two boundary levels.

Stacking $\+\eta^0
:=[\+\eta_0^{0\prime},\ldots,\+\eta_{K^0+1}^{0\prime}]'$,
equation~\eqref{eq:knotdef} reads $\+\eta^0=\*M_{K^0}\+\psi_{K^0}^0$
with
\begin{align}
    \*M_{K^0}
    :=\begin{bmatrix}
    \*I_p & \*0 & \*0 & \cdots & \*0\\
    \*I_p & I_1^0\*I_p & \*0 & \cdots & \*0\\
    \*I_p & I_1^0\*I_p & I_2^0\*I_p & \cdots & \*0\\
    \vdots & \vdots & \vdots & \ddots & \vdots\\
    \*I_p & I_1^0\*I_p & I_2^0\*I_p & \cdots & I_{K^0+1}^0\*I_p
    \end{bmatrix},
    \label{eq:Mmatrix}
\end{align}
which is block lower triangular with diagonal blocks
$\*I_p,I_1^0\*I_p,\ldots,I_{K^0+1}^0\*I_p$ and hence nonsingular for
every $N$ and $T$, every regime length being positive. Its inverse is
the successive difference operator. Subtracting consecutive terms
in~\eqref{eq:knotdef},
\begin{align}
    \+\eta_j^0-\+\eta_{j-1}^0
    =\Bigl(\+\theta_1^0+\sum_{m=1}^{j}I_m^0\+\kappa_m^0\Bigr)
    -\Bigl(\+\theta_1^0+\sum_{m=1}^{j-1}I_m^0\+\kappa_m^0\Bigr)
    =I_j^0\+\kappa_j^0,
    \label{eq:etadiff}
\end{align}
so that $\+\theta_1^0=\+\eta_0^0$ and
\begin{align}
    \+\kappa_j^0=\dfrac{\+\eta_j^0-\+\eta_{j-1}^0}{I_j^0}  =  \dfrac{
        \+\theta_{T_j^0}^0-\+\theta_{T_{j-1}^0}^0
    }{
        T_j^0-T_{j-1}^0
    }
    \label{eq:slopefromknots}
\end{align}
for $j=1,\ldots,K^0+1$, or,  In matrix
form, we have $\+\psi_{K^0}^0=\*M_{K^0}^{-1}\+\eta^0$, where
\begin{align}
    \*M_{K^0}^{-1}
    =\begin{bmatrix}
    \*I_p & \*0 & \*0 & \cdots & \*0\\
    -(I_1^0)^{-1}\*I_p & (I_1^0)^{-1}\*I_p & \*0 & \cdots & \*0\\
    \*0 & -(I_2^0)^{-1}\*I_p & (I_2^0)^{-1}\*I_p & \cdots & \*0\\
    \vdots & \vdots & \ddots & \ddots & \vdots\\
    \*0 & \*0 & \cdots & -(I_{K^0+1}^0)^{-1}\*I_p
    & (I_{K^0+1}^0)^{-1}\*I_p
    \end{bmatrix}.
    \label{eq:Minverse}
\end{align}
Equation~\eqref{eq:slopefromknots} is the map by which the slopes are
reported once the boundary levels have been estimated, and it is exact
and linear, so no approximation is involved at any point. The
boundary levels are therefore an exact nonsingular reparameterisation
of the post-kink model, changing the coordinates of the theory but
neither the estimator, its fitted values, nor the column space of the
design. It shows that the rate of \(\widetilde{\+\kappa}_j\) depends locally on the rates of the two boundary-level estimators enclosing regime \(j\), rather than on the full span of subsequent regimes.

The representation of the path in the new coordinates is an
interpolation. Fix a regime $\ell$ and a date
$t\in\mathcal K_\ell^0$, and let
$u_{\ell,t}:=(t-T_{\ell-1}^0)/I_\ell^0$ denote the fraction of regime
$\ell$ elapsed by date $t$, so that $u_{\ell,t}\in[0,1)$ at observed
dates. Substituting $\+\theta_{T_{\ell-1}^0}^0=\+\eta_{\ell-1}^0$
and~\eqref{eq:slopefromknots} into~\eqref{eq:Kink} gives
\begin{align}
    \+\theta_t^0
    =(1-u_{\ell,t})\+\eta_{\ell-1}^0+u_{\ell,t}\+\eta_\ell^0,
    \label{eq:knotrep}
\end{align}
so that every date is a convex combination of the two boundary levels
enclosing its regime. The identity is exact, being nothing but the
model's continuity and piecewise linearity, and it makes the source of
the convergence rates transparent, where the information about $\+\theta_t^0$
is the information about the two boundary levels adjacent to date $t$.

The implied loading of date $t$ on $\+\eta_j^0$ is
\begin{align}
    h_{j,t}
    :=
    \begin{cases}
    \displaystyle
    u_{j,t}
    =
    \frac{t-T_{j-1}^{0}}
         {T_{j}^{0}-T_{j-1}^{0}},
    & t\in\mathcal{K}_{j}^{0},
    \\[8pt]
    \displaystyle
    1-u_{j+1,t}
    =
    \frac{T_{j+1}^{0}-t}
         {T_{j+1}^{0}-T_{j}^{0}},
    & t\in\mathcal{K}_{j+1}^{0},
    \\[8pt]
    0,
    & \text{otherwise},
    \end{cases}
    \label{eq:hdef}
\end{align}
for $j=0,\ldots,K^0+1$, with the conventions
$\mathcal K_0^0=\mathcal K_{K^0+2}^0:=\varnothing$. The loading is
supported only on the two regimes adjacent to boundary $j$. It rises
linearly from zero over regime $j$, falls back linearly over regime
$j+1$, and vanishes elsewhere. At any date exactly two loadings are
nonzero and they sum to one, so \eqref{eq:knotrep} reads
$\+\theta_t^0=\sum_{j=0}^{K^0+1}h_{j,t}\+\eta_j^0$.

\begin{remark}[Linear B-spline representation]\label{rem:bspline}
A useful and perhaps surprising implication of this representation is that
the coefficient path admits an exact B-spline form. In particular, the path
$t\mapsto\+\theta_t^0$ is a spline of degree one with knots at the kink
dates, and $h_{0,t},\ldots,h_{K^0+1,t}$ are precisely its first-degree
B-spline basis functions, restricted to the observed dates
\citep[]{deBoor2001}. The associated B-spline coefficients are the
boundary levels and, because first-degree B-splines interpolate at the
knots, they coincide with the values of the coefficient path at those dates.
Unlike in nonparametric spline regression, however, the spline is not an
approximation as it is the exact data-generating coefficient path with finitely
many unknown knots and therefore introduces no approximation bias.
\end{remark}

Collecting the loadings into the regressor
$\+\phi_{i,t}^\eta:=[h_{0,t}\*x_{i,t}',\ldots,h_{K^0+1,t}\*x_{i,t}']'$,
 DGP~\eqref{eq:DGP} becomes
$y_{i,t}=\mu_i+\+\phi_{i,t}^{\eta\prime}\+\eta^0+\varepsilon_{i,t}$, a
linear panel regression in which every observation loads on at most two
parameter blocks. As noted earlier, this representation is equivalent to a varying-coefficient regression in which the time variation in the coefficient on $\*x_{i,t}$ is captured by a B-spline basis \citep{HastieTibshirani1993,HuangWuZhou2002}. This connection is noteworthy in the sense that although the estimator is derived from a structural kink model rather than from a conventional spline specification, the post-kink regression admits an exact first-degree B-spline representation.
Now, let $\+\Phi^\eta(\mathcal T^0_{K^0})$ stack the rows
$\+\phi_{i,t}^{\eta\prime}$ conformably with~\eqref{eq:pk_expanded}.
The two designs are related through the partial sums of the
loadings. Fix $t\in\mathcal K_\ell^0$. By~\eqref{eq:hdef}, the only potentially
nonzero loadings at date $t$ are $h_{\ell-1,t}=1-u_{\ell,t}$ and
$h_{\ell,t}=u_{\ell,t}$, so, for $m=0,\ldots,K^0+1$,
\begin{align}
    \sum_{j=m}^{K^0+1}h_{j,t}
    =\mathbf 1\{m<\ell\}+u_{\ell,t}\,\mathbf 1\{m=\ell\},
    \label{eq:partialsums}
\end{align}
and multiplying by $I_m^0$ and using
$I_\ell^0u_{\ell,t}=t-T_{\ell-1}^0$,
\begin{align}
    I_m^0\sum_{j=m}^{K^0+1}h_{j,t}
    &=I_m^0\,\mathbf 1\{m<\ell\}
    +\bigl(t-T_{\ell-1}^0\bigr)\mathbf 1\{m=\ell\}
    \notag \\
    &=d_{m,t}\bigl(\mathcal T^0_{K^0}\bigr),
    \label{eq:partialsumd}
\end{align}
where the second equality follows from~\eqref{eq:djt} for
$m=1,\ldots,K^0+1$ and from the convention of setting
$d_{0,t}(\mathcal T_{K^0}^0):=0$ for $m=0$. Now consider the product
$\+\Phi^\eta(\mathcal T^0_{K^0})\*M_{K^0}$. By~\eqref{eq:Mmatrix},
block column $0$ of $\*M_{K^0}$ carries $\*I_p$ in every block row,
and block column $m\ge1$ carries $I_m^0\*I_p$ in block rows
$j\ge m$ and $\*0_{p\times p}$ above, so the row belonging to
observation $(i,t)$ has $m$th block entry
\begin{align}
    \bigl[\+\phi_{i,t}^{\eta\prime}\*M_{K^0}\bigr]_m
    &=I_m^0\,\Bigl(\sum_{j=m}^{K^0+1}h_{j,t}\Bigr)\*x_{i,t}'\,
    \mathbf 1\{m\ge1\}
    +\Bigl(\sum_{j=0}^{K^0+1}h_{j,t}\Bigr)\*x_{i,t}'\,
    \mathbf 1\{m=0\}
    \notag \\
    &=d_{m,t}\bigl(\mathcal T^0_{K^0}\bigr)\*x_{i,t}'\,
    \mathbf 1\{m\ge1\}
    +\*x_{i,t}'\,\mathbf 1\{m=0\}
    \notag \\
    &=\bigl[\+\phi_{i,t}
    \bigl(\mathcal T^0_{K^0}\bigr)'\bigr]_m,
    \label{eq:rowident}
\end{align}
where the second equality holds by~\eqref{eq:partialsumd} and by the case $m=0$
of~\eqref{eq:partialsums}, which is the partition of unity
$\sum_{j=0}^{K^0+1}h_{j,t}=1$, and the third by the definition of
$\+\phi_{i,t}(\mathcal T^0_{K^0})$. Since~\eqref{eq:rowident} holds
at every observation, stacking the rows gives
\begin{align}
    \+\Phi(\mathcal T^0_{K^0})
    =\+\Phi^\eta(\mathcal T^0_{K^0})\,\*M_{K^0},
    \label{eq:designequiv}
\end{align}
so the two designs share the same column space and identical fitted
values, and the OLS estimators satisfy
$\widetilde{\+\eta}=\*M_{K^0}\widetilde{\+\psi}$ exactly in finite
samples.\footnote{The identity~\eqref{eq:designequiv} is what permits the
distribution theory to be developed on
$\+\Phi^\eta(\mathcal T^0_{K^0})$ while the estimator is computed
in the $\+\psi$ coordinates of~\eqref{eq:postest}. The two designs
have identical fitted values and residuals, so
$\widetilde{\+\eta}=\*M_{K^0}\widetilde{\+\psi}$ holds exactly in
finite samples, covariance matrices transform by the fixed
nonsingular map $\*M_{K^0}$.}



\begin{remark}[Identification as a degree-one B-spline basis]
To identify the basis precisely, let
\(\+\tau:=(T_0^0,T_0^0,T_1^0,\ldots,T_{K^0}^0,
T_{K^0+1}^0,T_{K^0+1}^0)\)
be the open knot vector.  Starting from the degree-zero interval indicators
\(B_{j,0}(t):=\mathbf{1}\{\tau_j\leq t<\tau_{j+1}\}\), the Cox--de Boor
recursion gives
\(\displaystyle
B_{j,1}(t)
=
\frac{t-\tau_j}{\tau_{j+1}-\tau_j}B_{j,0}(t)
+
\frac{\tau_{j+2}-t}{\tau_{j+2}-\tau_{j+1}}B_{j+1,0}(t)
\),
where a term with zero denominator is defined as zero. On the interval
to the left of \(T_j^0\), this recursion yields \(u_{j,t}\), while on
the interval to its right it yields \(1-u_{j+1,t}\). Hence
\(h_{j,t}=B_{j,1}(t)\) for \(j=0,\ldots,K^0+1\).
\end{remark}










Different boundary levels are estimated at different rates, and the
right scale for $\+\eta_j^0$ is the norm of its own loading sequence,
\(
    r_j^2:=\sum_{t=1}^{T}h_{j,t}^2
\)
for $j=0,\ldots,K^0+1$. Dividing the $j$th block of the design by
$r_j$ produces a time loading of unit length, so $r_j$ removes the
deterministic magnitude of the basis before the regressors and the
within transformation are brought in. The norms are available in
closed form. By~\eqref{eq:hdef}, $h_{j,t}$ is nonzero only on
$\mathcal K_j^0\cup\mathcal K_{j+1}^0$, rising linearly over regime
$j$ and falling linearly over regime $j+1$, so the sum splits into
two ramp contributions,
\(
    r_j^2
    =\sum_{t\in\mathcal K_j^0}u_{j,t}^2
    +\sum_{t\in\mathcal K_{j+1}^0}(1-u_{j+1,t})^2.
\)
Substituting $q=t-T_{j-1}^0$ in the first sum and $q=t-T_j^0$ in the
second, so that $u_{j,t}=q/I_j^0$ and $u_{j+1,t}=q/I_{j+1}^0$ with
$q$ running over the regime, gives $r_j^2=A(I_j^0)+B(I_{j+1}^0)$,
where for an integer $L\ge1$,
\(
    A(L):=\sum_{q=0}^{L-1}(q/L)^2=(L-1)(2L-1)/(6L)
\)
and
\(
    B(L):=\sum_{q=0}^{L-1}(1-q/L)^2=(L+1)(2L+1)/(6L),
\)
with $A(0):=B(0):=0$ at the endpoints, so that $r_0^2=B(I_1^0)$ and
$r_{K^0+1}^2=A(I_{K^0+1}^0)$. A ramp over a regime of length $L$
thus contributes $L/3$ per regime up to a bounded term, and
$r_j^2\asymp I_j^0+I_{j+1}^0$, so $r_j^2$ counts the periods in the
two regimes adjacent to boundary $j$, which are the only ones
informative about $\+\eta_j^0$. The following lemma records the
exact bounds.




\begin{lemma}[The rate of convergence]\label{lem:basis}
Let $\*H^0:=[h_{j,t}]_{1\le t\le T,\,0\le j\le K^0+1}$ and
$\*D^0:=\mathrm{diag}(r_0,\ldots,r_{K^0+1})$. Then, uniformly over all
admissible kink configurations,
\(
    \frac18\bigl(I_j^0+I_{j+1}^0\bigr)
    \le r_j^2
    \le I_j^0+I_{j+1}^0
\)
for $j=0,\ldots,K^0+1$, and
\(
    \frac12\*I_{K^0+2}
    \preceq(\*D^0)^{-1}\*H^{0\prime}\*H^0(\*D^0)^{-1}
    \preceq\frac32\*I_{K^0+2}.
\)
\end{lemma}
The normalised column by column is then  well
conditioned, whatever the number of kinks and however the sample is
divided across regimes. Non-adjacent loadings have disjoint support,
while adjacent ones overlap on a single regime, and the overlap
carries less than half of the mass of either column. This is the
first-degree case of the stability property of B-splines, whose
condition number is bounded independently of the knot sequence
\citep{deBoor1972,SchererShadrin1999}. In particular, since the
 bounds in Lemma \ref{lem:basis} imply
\(
    \frac23 r_j^{-2}
    \le
    \bigl[(\*H^{0\prime}\*H^0)^{-1}\bigr]_{jj}
    \le
    2 r_j^{-2},
\)
the diagonal of the inverse  satisfies
\([(\*H^{0\prime}\*H^0)^{-1}]_{jj}\asymp r_j^{-2}\),
which is precisely what gives us the rate of
\(\widetilde{\+\eta}_j\) below. Further, collect the scales in
$\mathbb D_{K^0+1}:=\mathrm{diag}(r_0,\ldots,r_{K^0+1})\otimes\*I_p$
and let $\*C$ be the centering matrix.
Evaluating the design at the true kink dates, define the normalised  matrix
\(
    \widehat{\*Q}_N
    :=\mathbb D_{K^0+1}^{-1}\,\frac1N\,
    \+\Phi^\eta(\mathcal T^0_{K^0})'\*C\,
    \+\Phi^\eta(\mathcal T^0_{K^0})\,\mathbb D_{K^0+1}^{-1}
\)
and
\(
    \+\Psi_{NT}
    :=\frac1{\sqrt N}\+\Phi^\eta(\mathcal T^0_{K^0})'\*C\+\varepsilon.
\)
 They are
linked to the estimator by the exact factorisation
\(
    \+\Phi^\eta(\mathcal T^0_{K^0})'\*C\,
    \+\Phi^\eta(\mathcal T^0_{K^0})
    =N\,\mathbb D_{K^0+1}\widehat{\*Q}_N\mathbb D_{K^0+1}.
\)


\begin{assumption}[Post-selection: design and rank]\label{ass:postrank}
We have $\|\widehat{\*Q}_N-\*Q_0\|_{op}=o_p(1)$, where $\*Q_0$ is
finite and positive definite with
$0<c\le\mu_{\min}(\*Q_0)\le\mu_{\max}(\*Q_0)\le C<\infty$ for some
constants $c$ and $C$.
\end{assumption}

The lower eigenvalue bound excludes the possibility that the centering
or the regressors asymptotically eliminate a normalised direction of
the design, and the upper bound prevents any normalised direction from
accumulating information at an order larger than its scale. By
Lemma~\ref{lem:basis}, the normalised time basis is uniformly well
conditioned, so positive definiteness of $\*Q_0$ is a restriction on
the regressors and not a restriction on the kink configuration.



\begin{assumption}[Post-selection: score and central limit theorem]
\label{ass:postclt}
Let $\{\*H_{NT}\}$ be any  sequence of
$l\times p(K^0+2)$ selection matrices with $l$ fixed and
$\limsup_{N,T\to\infty}\|\*H_{NT}\|_{op}<\infty$.
\begin{enumerate}[label=(\alph*)]
    \item The limit
    \(
        \+\Sigma_0
        :=\lim\mathrm{Var}\!\left(
        \mathbb D_{K^0+1}^{-1}\+\Psi_{NT}\right)
    \)
    exists and is finite and positive definite;
    \item whenever
    \(
        \+\Upsilon
        :=\lim\*H_{NT}\*Q_0^{-1}\+\Sigma_0\*Q_0^{-1}\*H_{NT}'
    \)
    exists and is positive definite, we have
    \(
        \*H_{NT}\*Q_0^{-1}\mathbb D_{K^0+1}^{-1}\+\Psi_{NT}
        \to_d\mathcal N\!\left(\*0_{l\times1},\+\Upsilon\right).
    \)
\end{enumerate}
\end{assumption}

Assumption~\ref{ass:postclt} is a high-level central limit condition
for fixed-dimensional contrasts of the normalised score, in the spirit
of \citet{QianSu2016}. It holds under arbitrary within-unit serial
correlation and heteroskedasticity with bounded spectrum, and weak
cross-sectional dependence can be maintained directly. The matrices
$\*H_{NT}$ allow the same result to deliver the individual boundary
levels, the slopes, and the path values below, and $\+\Upsilon$ is the
usual sandwich variance of the corresponding OLS contrast. We stress
again that Assumptions~\ref{ass:postrank} and~\ref{ass:postclt}
supplement, and do not replace,
Assumptions~\ref{ass:errors}--\ref{ass:signal}.

Let $\widetilde{\+\eta}^{\,o}$ denote the OLS estimator of $\+\eta^0$
from the centered boundary-level regression evaluated at the true kink
dates, so that $\widetilde{\+\eta}^{\,o}=\*M_{K^0}\widetilde{\+\psi}^{\,o}$
with $\widetilde{\+\psi}^{\,o}$ the oracle version
of~\eqref{eq:postest}. Since $\*C$ is symmetric and idempotent,
\begin{align}
    \widetilde{\+\eta}^{\,o}-\+\eta^0
    =\bigl[\+\Phi^\eta(\mathcal T^0_{K^0})'\*C
    \+\Phi^\eta(\mathcal T^0_{K^0})\bigr]^{-1}
    \+\Phi^\eta(\mathcal T^0_{K^0})'\*C\+\varepsilon,
    \label{eq:oracleerr}
\end{align}
and inserting the factorisation~\eqref{eq:gramfact} together with
$\+\Phi^\eta(\mathcal T^0_{K^0})'\*C\+\varepsilon=\sqrt N\+\Psi_{NT}$
yields the exact identity
\(
    \sqrt N\,\mathbb D_{K^0+1}
    \bigl(\widetilde{\+\eta}^{\,o}-\+\eta^0\bigr)
    =\widehat{\*Q}_N^{-1}\mathbb D_{K^0+1}^{-1}\+\Psi_{NT}.
\)
Because $\widehat{\*Q}_N^{-1}=\*Q_0^{-1}+o_p(1)$ in operator norm under
Assumption~\ref{ass:postrank}, combining the identity  with
Assumption~\ref{ass:postclt} and Slutsky's theorem gives the oracle
limit.

\begin{theorem}[Oracle post-kink asymptotic normality]
\label{thm:oracle}
Suppose that $K^0$ is fixed and that Assumptions~\ref{ass:postrank} and~\ref{ass:postclt} hold.
Then, as $N,T\to\infty$,
\begin{align*}
    \sqrt N\,\*H_{NT}\,\mathbb D_{K^0+1}
    \bigl(\widetilde{\+\eta}^{\,o}-\+\eta^0\bigr)
    \to_d
    \mathcal N\!\left(\*0_{l\times1},\+\Upsilon\right).
\end{align*}
\end{theorem}

The feasible estimator uses $\widehat{\mathcal T}_{\widehat K}$. Let
$\widetilde{\+\eta}_{\widehat K}
:=\*M_{\widehat K}\widetilde{\+\psi}_{\widehat K}$ be its boundary-level
representation. On the event
$\{\widehat{\mathcal T}_{\widehat K}=\mathcal T^0_{K^0}\}$ the feasible
and oracle regressions share the design, the response, and the OLS
formula, so $\widetilde{\+\eta}_{\widehat K}=\widetilde{\+\eta}^{\,o}$
exactly, and Corollary~\ref{cor:kinks} shows that the probability of
this event tends to one. The feasible estimator therefore inherits the
oracle limit.

\begin{theorem}[Feasible post-kink asymptotic normality]
\label{thm:normality}
Suppose that $K^0$ is fixed and that the assumptions of Corollary~\ref{cor:kinks} and
Assumptions~\ref{ass:postrank}--\ref{ass:postclt} hold. Then, as
$N,T\to\infty$,
\begin{enumerate}[label=(\roman*)]
\item $    \sqrt N\,\*H_{NT}\,\mathbb D_{K^0+1}
    \bigl(\widetilde{\+\eta}_{\widehat K}-\+\eta^0\bigr)
    \to_d
    \mathcal N\!\left(\*0_{l\times1},\+\Upsilon\right);$
    \item  $\widetilde{\+\eta}_j-\+\eta_j^0
=O_p\!\left(\left[N\left(I_j^0+I_{j+1}^0\right)\right]^{-1/2}\right)$ for each for each $j=0,\hdots, K^0+1$.
\end{enumerate}
\end{theorem}

The selection assumptions enter Theorem~\ref{thm:normality} only
through exact recovery. Conditional on the correct kink set, the limit
is the ordinary OLS distribution of the centered post-kink regression.
The theorem contains the rates of every object of interest. Let
$\*P_j:=\*e_{j+1}'\otimes\*I_p$ select the $j$th boundary block, so
that $\*P_j\mathbb D_{K^0+1}=r_j\*P_j$. Choosing $\*H_{NT}=\*P_j$,
whose limit variance is the $(j+1)$th diagonal $p\times p$ block of
$\*Q_0^{-1}\+\Sigma_0\*Q_0^{-1}$ and hence positive definite, gives a
nondegenerate limit for
$\sqrt N\,r_j(\widetilde{\+\eta}_j-\+\eta_j^0)$, and
Lemma~\ref{lem:basis} gives us the following rate of convergence
\begin{align}
    \widetilde{\+\eta}_j-\+\eta_j^0
    =O_p\!\left(\left[N\left(I_j^0+I_{j+1}^0\right)\right]^{-1/2}\right)
    \label{eq:etarate}
\end{align}
for $j=0,\ldots,K^0+1$. This says that a boundary level is estimated at the rate of
the combined length of the two regimes that inform it.
















Consider next the within-regime slopes. By~\eqref{eq:slopefromknots},
the estimation error of a slope is the successive difference of two
boundary-level errors divided by the regime length, and this identity
already contains the rate. For the limiting distribution, define the
slope scale
\begin{align}
    s_j:=I_j^0\left(r_{j-1}^{-2}+r_j^{-2}\right)^{-1/2}
    \label{eq:sj}
\end{align}
and the contrast
$\*H_{j,NT}^{\kappa}
:=s_j(I_j^0)^{-1}\bigl(r_j^{-1}\*P_j-r_{j-1}^{-1}\*P_{j-1}\bigr)$,
which produces $\sqrt N\,s_j(\widetilde{\+\kappa}_j-\+\kappa_j^0)$
exactly when applied to the left side of Theorem~\ref{thm:normality}
and satisfies $\*H_{j,NT}^{\kappa}\*H_{j,NT}^{\kappa\prime}=\*I_p$
because the two selectors act on disjoint blocks. Let
\(
    \+\Upsilon_j^{\kappa}
    :=\lim\*H_{j,NT}^{\kappa}\*Q_0^{-1}\+\Sigma_0\*Q_0^{-1}
    \*H_{j,NT}^{\kappa\prime},
\)
whenever the limit exists.

\begin{corollary}[Distribution and convergence rates of the slopes]
\label{cor:kappa}
Under the assumptions of Theorem~\ref{thm:normality}, for each
$j=1,\ldots,K^0+1$, suppose that $\+\Upsilon_j^{\kappa}$ exists and
is positive definite. Then, as $N,T\to\infty$,
\begin{enumerate}[label=(\roman*)]
    \item
    \(
    \sqrt N\,s_j
    \bigl(\widetilde{\+\kappa}_j-\+\kappa_j^0\bigr)
    \to_d
    \mathcal N\!\left(\*0_{p\times1},\+\Upsilon_j^{\kappa}\right);
    \)
    \item
    \(
    \widetilde{\+\kappa}_j-\+\kappa_j^0
    =O_p\!\left(
    \left[N\left(I_j^0\right)^2
    \left\{\left(I_{j-1}^0+I_j^0\right)\wedge\left(I_j^0+I_{j+1}^0\right)\right\}
    \right]^{-1/2}
    \right).
    \)
\end{enumerate}
\end{corollary}

The order of $s_j^2$ follows from Lemma~\ref{lem:basis} because
$s_j^2/(I_j^0)^2$ is the harmonic mean of $r_{j-1}^2$ and $r_j^2$,
which lies within a factor of two of their minimum.\footnote{To see this, let \(H(a,b):=2ab/(a+b)\) denote the harmonic mean of two positive scalars \(a\) and \(b\). With \(a=r_{j-1}^2\) and \(b=r_j^2\), the definition of \(s_j\) in~\eqref{eq:sj} gives \(s_j^2/(I_j^0)^2 = (r_{j-1}^{-2}+r_j^{-2})^{-1} = r_{j-1}^2 r_j^2/(r_{j-1}^2+r_j^2) = \frac12 H(r_{j-1}^2,r_j^2)\). The  inequality \(\frac12(a\wedge b)\le ab/(a+b)\le a\wedge b\) bounds this quantity between half the minimum and the minimum itself, so \(s_j^2/(I_j^0)^2\) and \(r_{j-1}^2\wedge r_j^2\) differ by at most a factor of two.} The rate has a simple reading. By~\eqref{eq:slopefromknots}, a slope
is the difference of the two boundary levels enclosing its regime,
divided by the regime length $I_j^0$, so its precision is
$(I_j^0)^{-1}$ times the precision of the worse of its two
endpoints. The first boundary level has
no regime to its left, so it is identified from regime~$1$ alone at
rate $(NI_1^0)^{-1/2}$, and dividing by $I_1^0$ gives
$[N(I_1^0)^3]^{-1/2}$ for the first slope and  the same argument applied at the right end gives
$[N(I_{K^0+1}^0)^3]^{-1/2}$ for the last. These are exactly the rates of a trend regression run on the
regime alone. An interior slope inherits the precision of its two endpoints. Each
boundary level enclosing regime $j$ converges at the
rate $[N(I_j^0+I_{j+1}^0)]^{-1/2}$ of its own pair of adjacent regimes, the
slope is their difference divided by $I_j^0$, and so it converges at
the worse of the two endpoint rates divided by $I_j^0$, which is
Corollary~\ref{cor:kappa}.



\begin{remark}[Equivalent representations of the slope rates]\label{rem:rateinterpretation}  Since \(\kappa_j^0=(\eta_j^0-\eta_{j-1}^0)/I_j^0\), and since \(r_{j-1}^2\asymp I_{j-1}^0+I_j^0\) and \(r_j^2\asymp I_j^0+I_{j+1}^0\), the exact slope scale satisfies \(s_j^2=(I_j^0)^2(r_{j-1}^{-2}+r_j^{-2})^{-1}\asymp (I_j^0)^2\frac{(I_{j-1}^0+I_j^0)(I_j^0+I_{j+1}^0)}{I_{j-1}^0+2I_j^0+I_{j+1}^0}\). Moreover, \((I_{j-1}^0+I_j^0)\wedge(I_j^0+I_{j+1}^0)=I_j^0+(I_{j-1}^0\wedge I_{j+1}^0)\), so an interior slope has rate \(O_p([N(I_j^0)^2\{I_j^0+(I_{j-1}^0\wedge I_{j+1}^0)\}]^{-1/2})\). At the endpoints, the conventions \(I_0^0=I_{K^0+2}^0=0\) imply \((I_0^0+I_1^0)\wedge(I_1^0+I_2^0)=I_1^0\) and \((I_{K^0}^0+I_{K^0+1}^0)\wedge(I_{K^0+1}^0+I_{K^0+2}^0)=I_{K^0+1}^0\), yielding the usual rates \(O_p([N(I_1^0)^3]^{-1/2})\) and \(O_p([N(I_{K^0+1}^0)^3]^{-1/2})\) for the first and last slopes, respectively. \end{remark}

\begin{remark}[Balanced regimes]\label{rem:balanced}
If $I_\ell^0\asymp T$ for all $\ell$, then
$r_j^2\asymp T$ and $s_j^2\asymp T^3$ for every $j$, and the slope rate
reduces to $(NT^3)^{-1/2}$, that is $T^{-3/2}$ in the time dimension.
The rate in Corollary~\ref{cor:kappa} makes precise which
departures from balance are costly: a short first or last regime is
estimated at the cube of its own length, while a short interior regime
borrows strength from its neighbours through the shared boundary
levels.
\end{remark}

\begin{remark}[Connection to classical trend regression]
\label{rem:trend}
The rate in Corollary~\ref{cor:kappa} has a natural connection to OLS
estimation of a slope in a deterministic trend regression. In the
special case $K^0=0$, there is a single regime with $I_1^0=T$, so
$s_1^2\asymp T^3$ and the rate reduces to $O_p((NT^3)^{-1/2})$. This is
the panel analog of the classical $T^{-3/2}$ rate for OLS in
$y_t=\alpha+\beta t+\varepsilon_t$, where
$\mathrm{Var}(\hat\beta)\propto
\bigl(\sum_{t=1}^T(t-\overline t)^2\bigr)^{-1}\asymp T^{-3}$ with
$\overline t:=(T+1)/2$; see \citet[][Ch.~16]{Hamilton1994}.
\end{remark}

We finally consider the coefficient path. Fix a regime index $\ell$ and
a date $t\in\mathcal K_\ell^0$. By~\eqref{eq:knotrep}, the path error
is the same convex combination of the two adjacent boundary-level
errors, and the natural scale is
\(
    w_t
    :=\bigl[(1-u_{\ell,t})^2r_{\ell-1}^{-2}
    +u_{\ell,t}^2r_\ell^{-2}\bigr]^{-1/2},
\)
with contrast
$\*H_{t,NT}^{\theta}
:=w_t\bigl[(1-u_{\ell,t})r_{\ell-1}^{-1}\*P_{\ell-1}
+u_{\ell,t}r_\ell^{-1}\*P_\ell\bigr]$, again satisfying
$\*H_{t,NT}^{\theta}\*H_{t,NT}^{\theta\prime}=\*I_p$. Let
\(
    \+\Upsilon_t^{\theta}
    :=\lim\*H_{t,NT}^{\theta}\*Q_0^{-1}\+\Sigma_0\*Q_0^{-1}
    \*H_{t,NT}^{\theta\prime},
\)
whenever the limit exists.















\begin{corollary}[Pointwise distribution of the coefficient path]
\label{cor:thetat}
Suppose that the assumptions of Theorem~\ref{thm:normality} hold and
fix a regime index $\ell\in\{1,\ldots,K^0+1\}$. Then, as
$N,T\to\infty$,
\begin{enumerate}[label=(\roman*)]
    \item for any deterministic sequence of dates
    $t\in\mathcal K_\ell^0$ such that $\+\Upsilon_t^{\theta}$ exists
    and is positive definite,
    \(
    \sqrt N\,w_t
    \bigl(\widetilde{\+\theta}_t-\+\theta_t^0\bigr)
    \to_d
    \mathcal N\!\left(\*0_{p\times1},\+\Upsilon_t^{\theta}\right);
    \)
    \item uniformly over $t\in\mathcal K_\ell^0$,
    \(
    \widetilde{\+\theta}_t-\+\theta_t^0
    =O_p\!\left(
    \left[N\left\{\left(I_{\ell-1}^0+I_\ell^0\right)
    \wedge\left(I_\ell^0+I_{\ell+1}^0\right)\right\}\right]^{-1/2}
    \right).
    \)
\end{enumerate}
\end{corollary}

The second claim follows from the convex-combination representation
\(
\widetilde{\+\theta}_t-\+\theta_t^0
=(1-u_{\ell,t})(\widetilde{\+\eta}_{\ell-1}-\+\eta_{\ell-1}^0)
+u_{\ell,t}(\widetilde{\+\eta}_{\ell}-\+\eta_{\ell}^0),
\)
which implies, uniformly over \(t\in\mathcal K_\ell^0\),
\(
\|\widetilde{\+\theta}_t-\+\theta_t^0\|
\le
\max\{\|\widetilde{\+\eta}_{\ell-1}-\+\eta_{\ell-1}^0\|,
\|\widetilde{\+\eta}_{\ell}-\+\eta_{\ell}^0\|\}.
\)
Moreover, \((1-u_{\ell,t})^2+u_{\ell,t}^2\le1\) implies
\(w_t^2\ge r_{\ell-1}^2\wedge r_\ell^2\), showing that the pointwise rate is never slower than the stated uniform rate. The weight $w_t$ traces how precision moves
along the path. At the kink dates the rate is that of the corresponding
boundary level, and in between the two adjacent loadings share the
load. Precision at a date is therefore a local property, governed by
the lengths of the regimes adjacent to the nearby boundaries, and not
by the amount of sample preceding or following the date.

\begin{remark}[Relation to spline regression asymptotics]
\label{rem:splinerates}
The local character of the rates is familiar from least squares spline
regression, where the pointwise variance of the fit is governed by the
sample size times the local knot spacing, so that precision at a point
depends on the mesh near that point alone
\citep{AgarwalStudden1980,ZhouShenWolfe1998,Huang2003}.
\end{remark}

\begin{remark}[Initial level]\label{rem:initlevel}
The initial level is the case $t=1$, $\ell=1$, and $u_{1,1}=0$ of
Corollary~\ref{cor:thetat}, in which $\*H_{1,NT}^{\theta}=\*P_0$ and
$w_1=r_0$ with $r_0^2\asymp I_1^0$. Hence
$\widetilde{\+\theta}_1-\+\theta_1^0=O_p((NI_1^0)^{-1/2})$, which
reduces to the familiar $(NT)^{-1/2}$ rate under balanced regimes. The
initial level is identified from the first regime alone, whose
observations are the only ones in which its loading is active.
\end{remark}

\begin{remark}[Fixed $T$]\label{rem:fixedT}
The results above extend to the fixed-$T$ case; see
\citet{KaddouraWesterlund2023} for a related discussion. When $T\ge3$
is fixed as $N\to\infty$, the normalisation $\mathbb D_{K^0+1}$ and the
transformation $\*M_{K^0}$ reduce to constant matrices, the selection
matrices may be taken fixed, and Assumption~\ref{ass:signal} simplifies
to $\sqrt N\vartheta_1J_{\min}^{-\zeta_1}\to_pc_1\in[0,\infty)$,
$\sqrt N\,J_{\min}\to\infty$, and
$N^{(\zeta_1+1)/2}\vartheta_1\to_p\infty$. All theorems continue to
hold, with every post-kink parameter converging at the usual
$N^{-1/2}$ rate.
\end{remark}


\begin{comment}
The path result above is pointwise in $t$, applying to each date
sequence separately. The following theorem establishes the
corresponding uniform rates over $t=1,\ldots,T$ for the coefficient
path and over $j=1,\ldots,K^0+1$ for the slopes, both for fixed $K^0$
and when $K^0$ grows.

\begin{theorem}[Uniform convergence rates of the post-kink estimators]
\label{thm:uniformtheta}
Suppose that the conclusions of Corollary~\ref{cor:kinks} hold, and
define $m_{\min}^0:=\min_{0\le j\le K^0+1}(I_j^0+I_{j+1}^0)$.
\begin{enumerate}[label=(\arabic*)]
\item If $K^0$ is fixed and
Assumptions~\ref{ass:postrank}--\ref{ass:postclt} hold, then
\begin{align*}
    \max_{1\le t\le T}
    \bigl\|\widetilde{\+\theta}_t-\+\theta_t^0\bigr\|
    =O_p\!\left(\bigl(Nm_{\min}^0\bigr)^{-1/2}\right).
\end{align*}
\item If $K^0$ is allowed to grow with $N$ and $T$, assume instead that
there exists a sequence of positive definite matrices
$\{\*Q_0^{(K^0)}\}$ such that
$\|\widehat{\*Q}_N-\*Q_0^{(K^0)}\|_{op}=o_p(1)$ with
$0<c\le\mu_{\min}(\*Q_0^{(K^0)})\le\mu_{\max}(\*Q_0^{(K^0)})\le
C<\infty$, and that
$\mathbb E\|\mathbb D_{K^0+1}^{-1}\+\Psi_{NT}\|^2=O(K^0+1)$. Then:
    \begin{enumerate}[label=(\alph*)]
        \item
        \(
        \displaystyle\max_{1\le j\le K^0+1}
        \bigl\|\widetilde{\+\kappa}_j-\+\kappa_j^0\bigr\|
        =O_p\!\left(
        \sqrt{\frac{K^0+1}{N(I_{\min}^0)^3}}
        \right);
        \)
        \item
        \(
        \displaystyle\max_{1\le t\le T}
        \bigl\|\widetilde{\+\theta}_t-\+\theta_t^0\bigr\|
        =O_p\!\left(
        \sqrt{\frac{K^0+1}{Nm_{\min}^0}}
        \right)
        =O_p\!\left(
        \sqrt{\frac{K^0+1}{NI_{\min}^0}}
        \right).
        \)
    \end{enumerate}
\item Under the conditions of part~(2):
    \begin{enumerate}[label=(\alph*)]
        \item If $(K^0+1)/\{N(I_{\min}^0)^3\}\to0$, then
        \(
        \max_{1\le j\le K^0+1}
        \|\widetilde{\+\kappa}_j-\+\kappa_j^0\|=o_p(1);
        \)
        \item If $(K^0+1)/(NI_{\min}^0)\to0$, then
        \(
        \max_{1\le t\le T}
        \|\widetilde{\+\theta}_t-\+\theta_t^0\|=o_p(1).
        \)
    \end{enumerate}
\end{enumerate}
\end{theorem}

Theorem~\ref{thm:uniformtheta} strengthens the pointwise results.
Part~(1) gives the fixed-$K^0$ benchmark for the coefficient path: the
largest deviation is governed by the least informed boundary level,
whose combined adjacent span is $m_{\min}^0$, of order $T$ under
balanced regimes. Part~(2) allows $K^0$ to grow, which is the regime in
which the model approaches nonparametric spline estimation with an
increasing number of knots. The Gram condition keeps the normalised
regression uniformly invertible, which by Lemma~\ref{lem:basis} is a
restriction on the regressors and not on the configuration, and the
score condition is the growing-dimension analog of a bounded score,
stacking $K^0+2$ blocks each normalised to a second moment of order
one; it holds, for example, whenever units are independent with error
covariances of uniformly bounded spectrum. The factor $\sqrt{K^0+1}$ in
the path bound reflects that each date loads on only the two boundary
levels adjacent to it, the local-support property of the B-spline
basis.
\end{comment}

Theorem~\ref{thm:normality} implies that inference on $\+\eta^0$
requires consistent estimation of
\(
    \mathrm{Avar}(\widetilde{\+\eta}_{\widehat K})
    :=N^{-1}\mathbb D_{K^0+1}^{-1}\*Q_0^{-1}
    \+\Sigma_0\*Q_0^{-1}\mathbb D_{K^0+1}^{-1},
\)
and covariance matrices and Wald contrasts transform by the same fixed
nonsingular map $\*M_{K^0}$, so the estimator may be computed directly
in the original $\+\psi$ coordinates of Section~\ref{sec:postkink}. Let
$\*y_i:=[y_{i,1},\ldots,y_{i,T}]'$ and let
$\widehat{\*u}_i
:=\*y_i-\+\Phi_i\bigl(\widehat{\mathcal T}_{\widehat K}\bigr)
\widetilde{\+\psi}_{\widehat K}$ denote the post-kink residual vector
of unit $i$, so that $\*C_T\widehat{\*u}_i$ collects the fitted
residuals of~\eqref{eq:pk_linear}. A natural cluster-robust variance
estimator is
\(
    \widehat{\+\Omega}
    :=\frac1N\widehat{\*A}^{-1}\widehat{\*G}\,\widehat{\*A}^{-1},
\)
where
\(
    \widehat{\*A}
    :=\frac1N\+\Phi\bigl(\widehat{\mathcal T}_{\widehat K}\bigr)'\*C
    \+\Phi\bigl(\widehat{\mathcal T}_{\widehat K}\bigr)
\)
and
\(
    \widehat{\*G}
    :=\frac1N\sum_{i=1}^N\hat{\*g}_i\hat{\*g}_i'
\)
with
\(
    \hat{\*g}_i
    :=\+\Phi_i\bigl(\widehat{\mathcal T}_{\widehat K}\bigr)'
    \*C_T\widehat{\*u}_i.
\)
For a linear hypothesis $H_0:\*R\+\psi^0_{K^0}=\*r$, with $\*R$ of full
row rank, the Wald statistic is
\(
    \mathcal{W}
    :=\bigl(\*R\widetilde{\+\psi}_{\widehat K}-\*r\bigr)'
    \bigl(\*R\widehat{\+\Omega}\*R'\bigr)^{-1}
    \bigl(\*R\widetilde{\+\psi}_{\widehat K}-\*r\bigr).
\)
For the path, since $\widetilde{\+\theta}_t
=\*L_t(\widehat{\mathcal T}_{\widehat K})
\widetilde{\+\psi}_{\widehat K}$ by~\eqref{eq:pk_path}, an asymptotic
pointwise $95\%$ confidence interval for the $m$th component of
$\+\theta_t^0$ is
\(
    \widetilde\theta_{m,t}\pm1.96\,
    \bigl[\widehat{\+\Omega}_{\theta,t}\bigr]_{mm}^{1/2},
\)
where
\(
    \widehat{\+\Omega}_{\theta,t}
    :=\*L_t(\widehat{\mathcal T}_{\widehat K})\widehat{\+\Omega}
    \*L_t(\widehat{\mathcal T}_{\widehat K})'.
\)\footnote{On the recovery event of Corollary~\ref{cor:kinks},
$\widehat{\+\Omega}$ coincides with the oracle cluster-robust
estimator at the true design, so consistency follows by the usual
argument under cross-sectional independence of
$\{(\*x_i,\+\varepsilon_i)\}$ together with standard moment and
cluster-leverage conditions, and transfers to the feasible estimator
because that event has probability tending to one. Under weak
cross-sectional dependence a cross-sectionally robust estimator would
be needed.}

\subsection{Component-wise kinks}\label{sec:component}

The baseline estimator imposes a common set of kink dates on the
entire coefficient vector. This is appropriate when the relationship
shifts jointly across regressors, but restrictive when only some
coefficients change slope at a given date. Following the
coefficient-by-coefficient logic of \citet{Kaddoura2025}, we now let
each coefficient carry its own kink structure. The analysis mirrors
the baseline case. Selection first recovers the coordinate-specific
kink sets, and post-selection estimation then runs the
within-transformed regression on the selected structure, analysed one
coordinate at a time in boundary-level coordinates.

For each $j=1,\ldots,p$, let coefficient $j$ have $K_j$ kinks at dates
$\mathcal T_{j,K_j}:=\{T_{j,1},\ldots,T_{j,K_j}\}$, where
$1=:T_{j,0}<T_{j,1}<\cdots<T_{j,K_j}<T<T_{j,K_j+1}:=T+1$, with regimes
$\mathcal K_{j,\ell}:=\{T_{j,\ell-1},\ldots,T_{j,\ell}-1\}$ for
$\ell=1,\ldots,K_j+1$. Within regime $\ell$ the coefficient evolves as
\begin{align}
    \theta_{j,t}
    =\theta_{j,T_{j,\ell-1}}+(t-T_{j,\ell-1})\kappa_{j,\ell}
    \label{eq:cwpath}
\end{align}
for $t\in\mathcal K_{j,\ell}$, where $\kappa_{j,\ell}\in\mathbb R$ is
the within-regime slope, and continuity at the coordinate-specific
kink dates is imposed by
$\theta_{j,T_{j,\ell}}
=\theta_{j,T_{j,\ell-1}}+(T_{j,\ell}-T_{j,\ell-1})\kappa_{j,\ell}$ for
$\ell=1,\ldots,K_j$. Equivalently, the scalar second difference
$\Delta^2\theta_{j,t}:=\theta_{j,t+1}-2\theta_{j,t}+\theta_{j,t-1}$ is
nonzero exactly when coefficient $j$ changes slope at date $t$.
To ease the exposition, let \(a^0\) and \(\widehat a\) denote the true
value and estimator, respectively, of any quantity \(a\). In particular,
for each \(j=1,\ldots,p\), let
\(\mathcal T_{j,K_j^0}^0=\{T_{j,1}^0,\ldots,T_{j,K_j^0}^0\}\) and
\(\widehat{\mathcal T}_{j,\widehat K_j}
=\{\widehat T_{j,1},\ldots,\widehat T_{j,\widehat K_j}\}\)
denote the true and estimated coordinate-specific kink sets, with
\(T_{j,0}^0=\widehat T_{j,0}:=1\) and
\(T_{j,K_j^0+1}^0=\widehat T_{j,\widehat K_j+1}:=T+1\).



The preliminary estimator of~\eqref{eq:initial} yields the adaptive
weights
$\dot\omega^{\mathrm{cw}}_{j,t}
:=|\Delta^2\dot\theta_{j,t}|^{-\zeta_2}$ for $j=1,\ldots,p$ and
$t=2,\ldots,T-1$, where $\zeta_2>0$, and the component-wise estimator
minimises
\begin{align}
    \mathcal L^{\mathrm{cw}}_{\vartheta_2}(\+\Theta_T)
    :=\frac1{NT}\sum_{i=1}^N\sum_{t=2}^{T}
    \bigl(\check y_{i,t}-\*x_{i,t}'\+\theta_t
    +\*x_{i,1}'\+\theta_1\bigr)^2
    +\vartheta_2\sum_{j=1}^{p}\sum_{t=2}^{T-1}
    \dot\omega^{\mathrm{cw}}_{j,t}\,
    \bigl|\Delta^2\theta_{j,t}\bigr|,
    \label{eq:objective_componentwise}
\end{align}
where $\vartheta_2>0$ is a tuning parameter. In contrast to the
group-$\ell_2$ penalty of~\eqref{eq:objective}, the penalty acts on
every scalar second difference separately, so coefficient $j$ may kink
at a date at which the remaining coefficients stay locally linear.
The relevant count of active constraints is no longer $K^0$ but the
total number of scalar kinks, $M^0:=\sum_{j=1}^pK_j^0$, and, when
$M^0\ge1$, the relevant signal is
$J^{\mathrm{cw}}_{\min}
:=\min_{j,\ell}|\kappa^0_{j,\ell+1}-\kappa^0_{j,\ell}|$, the minimum
taken over $j=1,\ldots,p$ and $\ell=1,\ldots,K_j^0$.

\begin{proposition}[Component-wise recovery of the kink structure]
\label{thm:cbc}
Suppose that Assumptions~\ref{ass:errors}, \ref{ass:rank}
and~\ref{ass:signal} hold, the latter with $J_{\min}$, $\vartheta_1$,
$\zeta_1$, and $K^0$ replaced by $J^{\mathrm{cw}}_{\min}$,
$\vartheta_2$, $\zeta_2$, and $M^0$, respectively, and with part (a)
replaced by $\vartheta_2\to_p0$ when $M^0=0$. Then, as $N,T\to\infty$,
\begin{enumerate}[label=(\roman*)]
    \item
    \(
    \mathbb{P}\!\left(\left\{
    \left|\Delta^2\widehat\theta^{\mathrm{cw}}_{j,t}\right|=0
    \;\;\forall\;t\in\mathcal{T}_{j,K_j^0}^{0c},
    \;\forall\;j=1,\ldots,p
    \right\}\right)\to1;
    \)
    \item
    \(
    \mathbb{P}\!\left(\left\{
    \widehat K_j=K_j^0
    \;\;\forall\;j=1,\ldots,p
    \right\}\right)\to1;
    \)
    \item
    \(
    \mathbb{P}\!\left(\left\{
    \widehat T_{j,\ell}=T_{j,\ell}^0,
    \;\ell=1,\ldots,K_j^0,
    \;\forall\;j
    \;\Big|\;
    \widehat K_j=K_j^0\;\forall\;j
   \right\}\right)\to1.
    \)
\end{enumerate}
\end{proposition}

Proposition~\ref{thm:cbc} is the component-wise analog of
Theorem~\ref{thm:sign} and Corollary~\ref{cor:kinks}. Selection
operates separately for each coefficient, so the estimator recovers
not only whether a kink is present, but also which coefficient kinks
and at which date.

Post-selection estimation again runs on the within transform and is
analysed in boundary-level coordinates, one coordinate at a time. For
coordinate $j$, the true kink dates induce regimes of lengths
$I_{j,\ell}^0:=T_{j,\ell}^0-T_{j,\ell-1}^0$, with the conventions
$I_{j,0}^0:=0$, $I_{j,K_j^0+2}^0:=0$, and
$\mathcal K_{j,0}^0=\mathcal K_{j,K_j^0+2}^0:=\varnothing$, and
boundary levels
\begin{align}
    \eta_{j,m}^0
    :=\theta_{j,1}^0
    +\sum_{\ell=1}^{m}I_{j,\ell}^0\kappa_{j,\ell}^0
    \label{eq:cwknotdef}
\end{align}
for $m=0,\ldots,K_j^0+1$, so that
$\kappa_{j,\ell}^0=(\eta_{j,\ell}^0-\eta_{j,\ell-1}^0)/I_{j,\ell}^0$.
With $u_{j,\ell,t}:=(t-T_{j,\ell-1}^0)/I_{j,\ell}^0$ the fraction of
regime $\ell$ of coordinate $j$ elapsed by date $t$, the loading of
date $t$ on $\eta_{j,m}^0$ is
\begin{align}
    h_{j,m,t}
    :=
    \begin{cases}
        u_{j,m,t}, & t\in\mathcal K_{j,m}^0,\\[2pt]
        1-u_{j,m+1,t}, & t\in\mathcal K_{j,m+1}^0,\\[2pt]
        0, & \text{otherwise},
    \end{cases}
    \label{eq:cwhdef}
\end{align}
the per-coordinate analog of~\eqref{eq:hdef}, and the interpolation
identity~\eqref{eq:knotrep} holds coordinate by coordinate,
$\theta_{j,t}^0
=(1-u_{j,\ell,t})\eta_{j,\ell-1}^0+u_{j,\ell,t}\eta_{j,\ell}^0$ for
$t\in\mathcal K_{j,\ell}^0$. Each coefficient is therefore a
free-knot linear spline on its own knot sequence, and the post-kink
model is a varying-coefficient regression in which the $p$
coordinates carry separate first-degree B-spline bases. Setting
$r_{j,m}^2:=\sum_{t=1}^Th_{j,m,t}^2$, Lemmas~\ref{lem:rjorder}
and~\ref{lem:basis} apply to each coordinate with the same absolute
constants, so that $r_{j,m}^2\asymp I_{j,m}^0+I_{j,m+1}^0$.

Substituting the interpolation identity into the DGP~\eqref{eq:DGP},
\begin{align}
    \*x_{i,t}'\+\theta_t^0
    =\sum_{j=1}^p\sum_{m=0}^{K_j^0+1}
    h_{j,m,t}\,x_{i,t,j}\,\eta_{j,m}^0,
    \label{eq:cwdesign}
\end{align}
which is linear in the boundary levels with known coefficients once
the kink dates are given. To write~\eqref{eq:cwdesign} as a
regression, stack the boundary levels coordinate by coordinate, the
$K_1^0+2$ boundary levels of coordinate $1$ first, then those of
coordinate $2$, and so on,
\(
    \+\eta^{0,\mathrm{cw}}
    :=[\eta_{1,0}^0,\ldots,\eta_{1,K_1^0+1}^0,\ldots,
    \eta_{p,0}^0,\ldots,\eta_{p,K_p^0+1}^0]'
    \in\mathbb R^{q^0\times1},
\)
where $q^0:=\sum_{j=1}^p(K_j^0+2)$ is fixed since every $K_j^0$ is
fixed, and define the conformable regressor
\begin{align}
    \+\phi^{\eta,\mathrm{cw}}_{i,t}
    :=\bigl[h_{1,0,t}\,x_{i,t,1},\ldots,h_{1,K_1^0+1,t}\,x_{i,t,1},
    \ldots,
    h_{p,0,t}\,x_{i,t,p},\ldots,h_{p,K_p^0+1,t}\,x_{i,t,p}\bigr]'
    \in\mathbb R^{q^0\times1},
    \label{eq:cwphidef}
\end{align}
so that~\eqref{eq:cwdesign} reads
$\*x_{i,t}'\+\theta_t^0
=\+\phi^{\eta,\mathrm{cw}\prime}_{i,t}\+\eta^{0,\mathrm{cw}}$ and the
DGP becomes
$y_{i,t}=\mu_i+\+\phi^{\eta,\mathrm{cw}\prime}_{i,t}
\+\eta^{0,\mathrm{cw}}+\varepsilon_{i,t}$. Let
$\+\Phi^{\eta,\mathrm{cw}}\in\mathbb R^{NT\times q^0}$ stack the rows
$\+\phi^{\eta,\mathrm{cw}\prime}_{i,t}$ and
$\+\varepsilon\in\mathbb R^{NT\times1}$ the errors
$\varepsilon_{i,t}$  conformably with $\*y$
in~\eqref{eq:pk_expanded}.

Collecting the rates in
\(
    \mathbb D^{\mathrm{cw}}
    :=\mathrm{diag}\bigl(r_{1,0},\ldots,r_{1,K_1^0+1},\ldots,
    r_{p,0},\ldots,r_{p,K_p^0+1}\bigr)
    \in\mathbb R^{q^0\times q^0},
\)
which is nonsingular because every $r_{j,m}$ is positive by
Lemma~\ref{lem:rjorder} applied per coordinate, the last kink date of
every coordinate lying strictly before $T$. Further, define
\(
    \widehat{\*Q}_N^{\mathrm{cw}}
    :=\bigl(\mathbb D^{\mathrm{cw}}\bigr)^{-1}\,\frac1N\,
    \+\Phi^{\eta,\mathrm{cw}\prime}\*C\,\+\Phi^{\eta,\mathrm{cw}}\,
    \bigl(\mathbb D^{\mathrm{cw}}\bigr)^{-1}
\)
and
\(
    \+\Psi_{NT}^{\mathrm{cw}}
    :=\frac1{\sqrt N}\,
    \+\Phi^{\eta,\mathrm{cw}\prime}\*C\,\+\varepsilon,
\)
with $\*C=\*I_N\otimes\*C_T$ the centering matrix.

The structure of these objects is read off from the columns of the
design. The column of $\+\Phi^{\eta,\mathrm{cw}}$ belonging to the
boundary level $\+\eta_{j,m}^0$ has $(i,t)$ entry
\(
    \chi_{i,t;j,m}:=h_{j,m,t}\,x_{i,t,j},
\)
the $j$th regressor switched on with weight $h_{j,m,t}$ over the two
regimes of coordinate $j$ adjacent to boundary $m$, and
premultiplication by $\*C$ demeans each column unit by unit,
\(
    \ddot\chi_{i,t;j,m}
    :=\chi_{i,t;j,m}-\frac{1}{T}\sum_{s=1}^{T}\chi_{i,s;j,m}.
\)
Hence the $\bigl((j,m),(j',m')\bigr)$ entry of
$N\mathbb D^{\mathrm{cw}}\widehat{\*Q}_N^{\mathrm{cw}}
\mathbb D^{\mathrm{cw}}
=\+\Phi^{\eta,\mathrm{cw}\prime}\*C\,\+\Phi^{\eta,\mathrm{cw}}$ is,
using $\*C'\*C=\*C$,
\begin{align}
    \left[
    N\mathbb D^{\mathrm{cw}}\widehat{\*Q}_N^{\mathrm{cw}}
    \mathbb D^{\mathrm{cw}}
    \right]_{(j,m),(j',m')}
    &=\sum_{i=1}^{N}\sum_{t=1}^{T}
    \ddot\chi_{i,t;j,m}\,\ddot\chi_{i,t;j',m'}
    \notag\\
    &=\sum_{i=1}^{N}\sum_{t=1}^{T}
    h_{j,m,t}h_{j',m',t}\,x_{i,t,j}x_{i,t,j'}
    \notag\\
    &\qquad
    -\frac1T\sum_{i=1}^{N}
    \left(\sum_{s=1}^{T}\chi_{i,s;j,m}\right)
    \left(\sum_{r=1}^{T}\chi_{i,r;j',m'}\right),
    \label{eq:cwgramentry}
\end{align}
the second equality by expanding the demeaned products. The first
term is the one that carries the structure, since it is nonzero only
at dates where both hat functions are active. Within a coordinate
($j'=j$), the hat functions are built on the same kink set,
$h_{j,m,t}$ is supported on
$\mathcal K_{j,m}^0\cup\mathcal K_{j,m+1}^0$ and vanishes at
$t=T_{j,m-1}^0$ and $t=T_{j,m+1}^0$, so for $|m-m'|\ge2$ the supports
of $h_{j,m,\cdot}$ and $h_{j,m',\cdot}$ are disjoint and
\(
    \sum_{i=1}^{N}\sum_{t=1}^{T}
    h_{j,m,t}h_{j,m',t}\,x_{i,t,j}^2=0,
    \) for \( |m-m'|\ge2.
\)
The leading term of each diagonal block is therefore tridiagonal in
$m$, the deterministic hat overlaps of Lemma~\ref{lem:basis} weighted
by $x_{i,t,j}^2$. Across coordinates ($j'\neq j$), the hat functions
are built on different kink sets, so no zero pattern is available,
and the leading term is
\(    \sum_{i=1}^{N}\;\sum_{t\in\mathcal S_{j,m;j',m'}}
    h_{j,m,t}h_{j',m',t}\,x_{i,t,j}x_{i,t,j'},
    \) where
    \( \mathcal S_{j,m;j',m'}
    :=\bigl\{t:h_{j,m,t}h_{j',m',t}>0\bigr\}
\)
a hat-weighted sample cross moment of $x_{i,t,j}$ and $x_{i,t,j'}$
over the window on which both hats are active. In both cases the
demeaning term in~\eqref{eq:cwgramentry} is a within-unit mean
correction that does not respect the banding, and its control is part
of the rank condition below. In sum,
$\widehat{\*Q}_N^{\mathrm{cw}}$ couples boundary levels of the same
coordinate only when they are adjacent, and boundary levels of
different coordinates only through the comovement of their regressors
on overlapping regimes. The post-kink estimator
$\widetilde{\+\eta}^{\mathrm{cw}}$ is the OLS estimator of
$\+\eta^{0,\mathrm{cw}}$ from the centered regression evaluated at
the estimated kink sets,
\begin{align}
\widetilde{\+\eta}^{\mathrm{cw}}
    :=\bigl[\+\Phi^{\eta,\mathrm{cw}\prime}\*C\,
    \+\Phi^{\eta,\mathrm{cw}}\bigr]^{-1}
    \+\Phi^{\eta,\mathrm{cw}\prime}\*C\,\*y,
    \label{eq:cwols}
\end{align}
with $\+\Phi^{\eta,\mathrm{cw}}$ built from
$\widehat{\mathcal T}_{j,\widehat K_j}$, $j=1,\ldots,p$, the inverse
existing w.p.a.1 under Assumption~\ref{ass:cwrank}(a). On the
recovery event of Proposition~\ref{thm:cbc} it coincides with the
oracle estimator evaluated at the true coordinate-specific dates.
\begin{assumption}[Component-wise post-selection design and central
limit theorem]\label{ass:cwrank}
Every $K_j^0$, $j=1,\ldots,p$, is fixed, and
$q^0=\sum_{j=1}^p(K_j^0+2)$.
\begin{enumerate}[label=(\alph*)]
    \item We have
    $\|\widehat{\*Q}_N^{\mathrm{cw}}-\*Q_0^{\mathrm{cw}}\|_{op}
    =o_p(1)$, where $\*Q_0^{\mathrm{cw}}$ is finite and positive
    definite with
    $0<c\le\mu_{\min}(\*Q_0^{\mathrm{cw}})
    \le\mu_{\max}(\*Q_0^{\mathrm{cw}})\le C<\infty$;
    \item the limit
    \(
        \+\Sigma_0^{\mathrm{cw}}
        :=\lim\operatorname{Var}\!\left(
        (\mathbb D^{\mathrm{cw}})^{-1}
        \+\Psi_{NT}^{\mathrm{cw}}\right)
    \)
    exists and is finite and positive definite, and for any
    deterministic sequence of $l\times q^0$ matrices
    $\{\*H_{NT}^{\mathrm{cw}}\}$ with $l$ fixed and
    $\limsup_{N,T\to\infty}\|\*H_{NT}^{\mathrm{cw}}\|_{op}<\infty$,
    whenever
    \(
        \+\Upsilon^{\mathrm{cw}}
        :=\lim\*H_{NT}^{\mathrm{cw}}(\*Q_0^{\mathrm{cw}})^{-1}
        \+\Sigma_0^{\mathrm{cw}}(\*Q_0^{\mathrm{cw}})^{-1}
        \*H_{NT}^{\mathrm{cw}\prime}
    \)
    exists and is positive definite, we have
    \(
        \*H_{NT}^{\mathrm{cw}}(\*Q_0^{\mathrm{cw}})^{-1}
        (\mathbb D^{\mathrm{cw}})^{-1}
        \+\Psi_{NT}^{\mathrm{cw}}
        \to_d\mathcal N(\*0_{l\times1},\+\Upsilon^{\mathrm{cw}}).
    \)
\end{enumerate}
\end{assumption}

Assumption~\ref{ass:cwrank} is the component-wise analog of
Assumptions~\ref{ass:postrank} and~\ref{ass:postclt}, stated on the
true design so that every object has the fixed dimension $q^0$ and
the scaling matrix is deterministic. The division of labour is the
same as in the baseline case, with one addition. Within a coordinate,
Lemma~\ref{lem:basis} controls the time basis, so the lower
eigenvalue bound in part (a) is not a hidden restriction on any
coordinate's kink configuration. Across coordinates, however, the
loadings of different coefficients may overlap arbitrarily in time,
and their conditioning is governed by the regressors, so the positive
definiteness of $\*Q_0^{\mathrm{cw}}$ is where the cross-coordinate
regressor design enters.


\begin{proposition}[Asymptotic normality of the component-wise
post-kink estimator]\label{prop:cwnorm}
Suppose that the assumptions of Proposition~\ref{thm:cbc} and
Assumption~\ref{ass:cwrank} hold, and define
\(
    s_{j,\ell}
    :=I_{j,\ell}^0
    \bigl(r_{j,\ell-1}^{-2}+r_{j,\ell}^{-2}\bigr)^{-1/2}
\)
for $j=1,\ldots,p$ and $\ell=1,\ldots,K_j^0+1$. Then, as
$N,T\to\infty$,
\begin{enumerate}[label=(\roman*)]
    \item for any deterministic sequence
    $\{\*H_{NT}^{\mathrm{cw}}\}$ as in
    Assumption~\ref{ass:cwrank}(b),
    \(
    \sqrt N\,\*H_{NT}^{\mathrm{cw}}\,\mathbb D^{\mathrm{cw}}
    \bigl(\widetilde{\+\eta}^{\mathrm{cw}}
    -\+\eta^{0,\mathrm{cw}}\bigr)
    \to_d
    \mathcal N\bigl(\*0_{l\times1},\+\Upsilon^{\mathrm{cw}}\bigr);
    \)
    \item
    \(
    \widetilde\kappa^{\mathrm{cw}}_{j,\ell}-\kappa^0_{j,\ell}
    =O_p\!\left(\bigl[Ns_{j,\ell}^2\bigr]^{-1/2}\right)
    =O_p\!\left(\Bigl[N\bigl(I_{j,\ell}^0\bigr)^2
    \bigl\{(I_{j,\ell-1}^0+I_{j,\ell}^0)
    \wedge(I_{j,\ell}^0+I_{j,\ell+1}^0)\bigr\}\Bigr]^{-1/2}\right)
    \)
    for $j=1,\ldots,p$ and $\ell=1,\ldots,K_j^0+1$;
    \item
    \(
    \widetilde\theta^{\mathrm{cw}}_{j,t}-\theta^0_{j,t}
    =O_p\!\left(\Bigl[N
    \bigl\{(I_{j,\ell-1}^0+I_{j,\ell}^0)
    \wedge(I_{j,\ell}^0+I_{j,\ell+1}^0)\bigr\}\Bigr]^{-1/2}\right)
    \)
    uniformly over $t\in\mathcal K^0_{j,\ell}$, for
    $j=1,\ldots,p$ and $\ell=1,\ldots,K_j^0+1$, and in particular
    $\widetilde\theta^{\mathrm{cw}}_{j,1}-\theta^0_{j,1}
    =O_p\bigl((NI_{j,1}^0)^{-1/2}\bigr)$.
\end{enumerate}
\end{proposition}

The component-wise rates therefore inherit the adjacent-regime
structure of the baseline model. Each coordinate-specific slope is the
contrast of the two boundary levels enclosing its regime, and each
boundary level is informed by the two regimes of that coordinate
adjacent to it, so the precision of coefficient $j$ is governed
entirely by its own regime configuration, independently of where the
remaining coefficients kink.




\section{Monte Carlo simulations}\label{sec:MC}

\subsection{Design}\label{sec:MCdesign}

We study the finite sample behaviour of the estimator in a restricted
version of the DGP in~\eqref{eq:DGP}. The regressors are generated as
\begin{align*}
    \*x_{i,t}=0.2\,\left(\*1_{p\times1}\cdot\xi_i\right)+\*z_{i,t},
\end{align*}
where $\xi_i\sim\mathcal N(0,1)$ and
$\*z_{i,t}\sim\mathcal N(\*0_{p\times1},\*I_p)$ are independent across
$i$ and $t$, so that $0.2\,\xi_i$ induces a common component in the
regressors within a unit. Following \citet{bai2002determining}, the
error carries serial dependence, cross-sectional dependence, and
heteroskedasticity through
\begin{align*}
    \varepsilon_{i,t}
    =\varpi_1\varepsilon_{i,t-1}+\varsigma_{i,t}
    +\varpi_2\sum_{r=1}^{R}
    \bigl(\varsigma_{i-r,t}+\varsigma_{i+r,t}\bigr),
\end{align*}
with $\varepsilon_{i,0}=0$ and $\varsigma_{i,t}\sim\mathcal N(0,1)$,
where $\varpi_1$ governs serial dependence, $\varpi_2$ cross-sectional
dependence, and $R$ its range, with baseline
$(\varpi_1,\varpi_2,R)=(0.2,0.1,5)$. The fixed effects are
$\mu_i\sim\mathcal N(0,1)$, independent of all else. The initial level
is $\+\theta_1^0=[1,-1]'$ when $p=2$ and $\+\theta_1^0=[1,-1,1]'$ when
$p=3$, and the path is built from the within-regime slopes
through~\eqref{eq:Kinkalt}. The sample sizes cover all nine
combinations of $N\in\{50,100,200\}$ and $T\in\{20,40,80\}$, and all
results are based on $1000$ replications.

The eight designs fall into three groups. The first group holds the
error dependence at its baseline and varies the kink configuration,
so that the number of kinks and the way identification is spread
across regimes are the only things that change. The second group
perturbs that configuration in the two directions the theory
identifies as costly, namely unbalanced regime lengths and stronger
error dependence. The third group breaks the common-kink restriction
altogether and lets each coordinate carry its own kink structure.
DGPs~1--3 and~5--7 have $p=2$ and use the group estimator of
Section~\ref{sec:initial}; DGP~4 has $p=2$ and DGP~8 has $p=3$, and
both use the component-wise estimator of Section~\ref{sec:component}.

\bigskip
\noindent\textbf{DGP~1 ($K^0=0$).}\;
A single regime with $\+\kappa_1^0=[0,0]'$, so the path is flat. The
estimator should return $\widehat K=0$, and the design measures
over-selection when there is nothing to find.

\bigskip
\noindent\textbf{DGP~2 ($K^0=1$).}\;
A kink at $T_1^0=\lfloor T/2\rfloor$ with $\+\kappa_1^0=[0.3,-0.2]'$
and $\+\kappa_2^0=[-0.2,0.3]'$. Each coordinate rises and then
reverses, so the kink is a genuine change of direction. Both regimes
are of length $I^0_\ell\asymp T/2$, so both slopes are endpoint
slopes in the sense of Corollary~\ref{cor:kappa}.

\bigskip
\noindent\textbf{DGP~3 ($K^0=2$).}\;
Kinks at $\lfloor T/3\rfloor$ and $\lfloor2T/3\rfloor$ with slopes
$[0,0]'$, $[0.5,-0.4]'$, and $[0,0]'$. The path is flat in the outer
regimes and linear in the middle one, so all identification of the two
kinks comes from the middle regime. This is the most demanding
selection design among those with balanced regimes.

\bigskip
\noindent\textbf{DGP~4 (coordinate-specific kinks).}\;
Coordinate one kinks once at $\lfloor T/3\rfloor$ with regime slopes
$(0.4,-0.3)$ and coordinate two kinks once at $\lfloor2T/3\rfloor$
with regime slopes $(-0.3,0.4)$, so that $K_1^0=K_2^0=1$ but the kink
dates differ. No common kink set describes the path, and the
component-wise estimator should recover the two coordinate-specific
kink sets.

\bigskip
\noindent\textbf{DGP~5 (uneven regimes, $K^0=2$).}\;
Kinks at $\lfloor T/5\rfloor$ and $\lfloor 4T/5\rfloor$ with slopes
$[0,0]'$, $[0.5,-0.4]'$, and $[0,0]'$, so the middle regime is long
and the two outer regimes short. The design isolates the role of the
regime configuration in Corollary~\ref{cor:kappa}: $\+\kappa_2^0$ is
an interior slope flanked by a long own regime, while $\+\kappa_1^0$
and $\+\kappa_3^0$ are endpoint slopes identified from a short own
regime alone.

\bigskip
\noindent\textbf{DGP~6 (DGP~2 under stronger dependence).}\;
The kink configuration of DGP~2, with the error dependence raised to
$(\varpi_1,\varpi_2)=(0.5,0.2)$. The design measures the sensitivity
of selection and estimation to serial and cross-sectional dependence
in the errors when identification is otherwise strong.

\bigskip
\noindent\textbf{DGP~7 (DGP~3 under stronger dependence).}\;
The kink configuration of DGP~3, the most demanding selection design,
under the same stronger dependence $(\varpi_1,\varpi_2)=(0.5,0.2)$.
This is the joint stress of weak identification and strong error
dependence.

\bigskip
\noindent\textbf{DGP~8 (heterogeneous kink counts, $p=3$).}\;
Three coordinates with different numbers of kinks: coordinate one is
flat ($K_1^0=0$), coordinate two kinks once at $\lfloor T/2\rfloor$
with slopes $(0.4,-0.3)$, and coordinate three kinks twice at
$\lfloor T/3\rfloor$ and $\lfloor2T/3\rfloor$ with slopes
$(0,0.5,0)$. No common kink set describes the path, and the
component-wise estimator must recover a different number of kinks in
each coordinate, including none in the first.

\subsection{Implementation}\label{sec:MCimpl}

The tuning parameter $\vartheta_1$ is chosen by minimising
\begin{align}
    \mathrm{IC}(\vartheta_1)
    :=\log\Biggl\{\frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^T
    \Bigl(\ddot y_{i,t}
    -\ddot{\+\phi}_{i,t}
    \bigl(\widehat{\mathcal T}_{\widehat K_{\vartheta_1}}\bigr)'
    \widetilde{\+\psi}_{\widehat K_{\vartheta_1}}\Bigr)^2\Biggr\}
    +\varrho_{N,T}\,p\,\bigl(\widehat K_{\vartheta_1}+1\bigr),
    \label{eq:IC}
\end{align}
where $\varrho_{N,T}:=\log(NT)/NT$ is a BIC-type rate in the spirit
of \citet{QianSu2016}. The residuals are within transformed and
post-kinks. The convex objective~\eqref{eq:objective} is
solved along the grid $\vartheta_1=c\,(NT)^{-1/2}$ with $c$ on a
logarithmic grid. The kink set is built by stepwise selection on the
criterion, adding the candidate whose inclusion lowers it most and
dropping any date whose removal lowers it further, and each selected
date is then moved within a window of three periods to the position
minimising the post-kink sum of squared residuals, with a final
backward pass under the criterion. The adaptive weights use the
perturbed form
$\dot\omega_t=(\|\Delta\dot{\+\theta}_{t+1}
-\Delta\dot{\+\theta}_t\|+\delta_t)^{-\zeta_1}$ with
$\delta_t=N^{-1/2}$ and $\zeta_1=2$.

DGPs~4 and~8 use the separable
analogue~\eqref{eq:objective_componentwise} with $\zeta_2=2$. The
candidate pairs are the active scalar second differences accumulated
along the $\vartheta_2$ grid, and the coordinate-specific kink sets
are built by the same stepwise procedure on~\eqref{eq:IC}, with
$p(\widehat K+1)$ replaced by $\sum_{j=1}^p(\widehat K_j+1)$, adding
or dropping the pair, in whichever coordinate, that improves the
criterion most, followed by the same local refinement and final
backward pass applied to the selected pairs.

\subsection{Results}\label{sec:MCresults}

Selection is measured by three statistics, each averaged over the
replications. Correct selection is the frequency, in percent, with
which the estimated kink dates coincide exactly with the true ones,
which for DGP~1 is the frequency of $\widehat K=0$. The kink-count
bias $\mathbb E[\widehat K-K^0]$ separates missed kinks from spurious
ones. Misclassification is the percentage of dates assigned to the
wrong regime, the regime label of date $t$ implied by a kink set
$\mathcal T$ being
$r_t(\mathcal T):=\sum_{T_k\in\mathcal T}\*1\{T_k\le t\}$.
Estimation is measured by the bias and root mean squared error of the
within-regime slopes, regime by regime and pooled across coordinates,
conditional on exact recovery of the kink set, and of the coefficient
path, averaged over dates and coordinates, unconditional.
Table~\ref{tab:mcsel} reports selection for the common-kink designs,
Tables~\ref{tab:mcselcw} and~\ref{tab:mcselcw8} the component-wise
results, and Tables~\ref{tab:mcest}, \ref{tab:mcestrob},
\ref{tab:mccw} and~\ref{tab:mccw8} the estimation results.

\begin{center}
\textbf{[Tables 1--7 about here]}
\end{center}

Estimation improves in both dimensions and at the predicted rates.
Every RMSE column falls monotonically in $N$ and in $T$. Moreover, selection improves in $N$. In DGP~1 there are no kinks and the
estimator finds none, with correct selection never below $99.3$
percent. DGP~2, with one well-separated kink, runs from $97.8$ to
$99.8$ percent. DGP~3 is harder, since the outer regimes are flat and
both kinks are identified from the middle regime alone, and there
correct selection rises from $80.0$ percent at $N=50$ to $98.7$
percent at $N=200$ when $T=20$, and from $73.0$ to $96.2$ percent in
DGP~7 when $T=80$. Comparing DGP~6 with DGP~2 and DGP~7 with DGP~3
shows that error dependence costs a few percentage points.

The span matters less. Correct selection is roughly flat in $T$,
moving from $80.0$ to $83.9$ to $82.0$ percent in DGP~3 at $N=50$ and
from $70.9$ to $79.8$ to $73.0$ percent in DGP~7. What rises with $T$
is over-selection: in DGP~7 at $N=50$ the kink-count bias runs from
$0.05$ to $0.28$ and misclassification from $3.8$ to $12.5$ percent,
so the errors are spurious kinks rather than missed ones. With
$J_{\min}$ fixed, a longer span widens the $\sqrt{T/N}$ band of
Theorem~\ref{thm:consistency}(c) within which noise survives the
shrinkage, while also lengthening the regimes from which the true
slope changes are estimated; the two roughly offset in the
exact-recovery rate. The effect is confined to $N=50$, the kink-count
bias being flat in $T$ at $N=100$ and $N=200$ in every design. Precision depends on the configuration, not only the sample size.
DGP~5 has a long middle regime and short outer ones, so
$\+\kappa_2^0$ is an interior slope with two well-informed boundary
levels while $\+\kappa_1^0$ and $\+\kappa_3^0$ are endpoint slopes
converging at the cube of their own short regime. In all nine cells
the interior slope is estimated more precisely than the endpoint
slopes. Against DGP~3, which differs only in where the kinks sit, the
middle slope improves from $0.019$ to $0.010$ at $(N,T)=(50,20)$
while the first deteriorates from $0.030$ to $0.051$.

The component-wise estimator behaves the same way coordinate by
coordinate. In DGP~4 both coordinates are recovered with comparable
accuracy, from about $90$ percent at $N=50$ to about $97$ percent at
$N=200$. In DGP~8 the three coordinates carry $0$, $1$ and $2$ kinks
and each performs according to its own configuration. The globally
linear first coordinate is correctly identified as such in $95.0$ to
$98.3$ percent of replications despite its neighbours kinking, the
second behaves like DGP~2 and reaches $97.7$ percent, and the third
has the configuration of DGP~3 and runs from $61.2$ to $93.4$
percent. This is what Proposition~\ref{prop:cwnorm} asserts. The cross-section is therefore the dimension that buys correct
dating, and a researcher with a short panel and a long span should
expect over-selection rather than missed kinks. That error is mild
for the fitted path. Even in DGP~7 the path RMSE falls with $T$ at
every $N$, from $0.100$ to $0.051$ at $N=50$. This implies that a spurious kink adds a basis function to a
design that already spans the truth, so the fitted values are barely
affected.




\section{Dating the Debt--Growth Relationship}\label{sec:debtgrowth}

The debt--growth literature asks whether the association changes once
public debt crosses a particular level. \citet{ReinhartRogoff2010} report
markedly lower growth above 90 percent of GDP.
\citet{HerndonAshPollin2014} replicate that exercise and show the finding
turns on data construction and weighting. The work that followed finds the
estimated nonlinearity fragile. It weakens once endogeneity and
cross-country heterogeneity are taken seriously, and it does not survive as
a universal cutoff \citep{PanizzaPresbitero2014, PescatoriSandriSimon2014,
EberhardtPresbitero2015, ChudikMohaddesPesaranRaissi2017}. Collecting 816
estimates from 47 studies, \citet{Heimberger2023} cannot reject a zero
average effect once publication selection is accounted for, and finds
threshold estimates to be sensitive to data and specification choices.

We ask when the association changed rather than at what level of debt it
differs. The two questions are distinct. A threshold model lets the
coefficient depend on where an economy sits in the debt distribution, and
an economy crosses back and forth as its debt ratio moves. Our model lets
the coefficient bend at calendar dates common to all economies, and asks
how fast it was moving on either side of each bend. The sample covers the
global financial crisis and the pandemic, two occasions on which debt and
output moved sharply for very different reasons, and the estimator is built
to separate a movement that reverses from one that does not. The exercise
is descriptive throughout. What it delivers is dates, annual rates of
change, and a distinction between a temporary excursion and a permanent
reorientation, not a causal effect of debt on growth.

\subsection{A debt-only benchmark}\label{sec:dg_benchmark}

Let $g_{i,t}$ denote annual real GDP growth in percentage points and
$d_{i,t}$ general government gross debt as a percentage of GDP. The
benchmark specification is
\begin{align}
    g_{i,t}
    &= \mu_i+\theta_t d_{i,t}+\varepsilon_{i,t},
    \label{eq:dg_model}\\[-2pt]
    \theta_t
    &= \theta_1+\kappa_1(t-1)
       +\sum_{m=1}^{K}\gamma_m(t-T_m)_+,
    \qquad
    \gamma_m:=\kappa_{m+1}-\kappa_m,
    \label{eq:dg_theta}
\end{align}
where $(a)_+:=\max\{a,0\}$. Within regime $\ell$ the slope $\kappa_\ell$ is
the annual change in the debt coefficient and $T_\ell$ is a date at which
that rate changes. Thus $\theta_t=-0.03$ means that a ten point higher debt
ratio is associated with 0.3 percentage points lower contemporaneous
growth, and $\kappa_\ell=0.004$ means that this association becomes 0.04
percentage points less negative each year for the same ten point
difference in debt.

The data are the April 2025 vintage of the IMF \emph{World Economic
Outlook} database \citep{IMFWEO2025}, from which we take
\texttt{NGDP\_RPCH} for real GDP growth and \texttt{GGXWDG\_NGDP} for
general government gross debt. The sample runs from 2000 to 2024. We drop
observations missing either series or with debt outside $(0,400)$ percent
of GDP and retain economies observed in every year, which leaves a balanced
panel of $N=158$ economies over $T=25$ years. Growth is winsorised at the
second and ninety-eighth percentiles, affecting four percent of the
observations. The final years should be read as a vintage result, since the
WEO carries IMF staff estimates where complete historical information is
unavailable and the April 2025 reference forecast was built on information
available as of 4 April 2025.

We set $\zeta_1=2$ and solve the penalised problem along a logarithmic grid
in $\vartheta_1$. The tuning parameter minimises the post-kink information
criterion, whose penalty
$\varrho_{N,T}=0.05\log(NT)/\sqrt{NT}$ is of the BIC type used by
\citet{QianSu2016}. Each selected date is then moved
within a three-year window to the date minimising the post-kink residual
sum of squares, and standard errors are clustered by economy.
Figure~\ref{fig:dg_ic} reports the selection path. The criterion is
minimised at $\widehat\vartheta_1=0.17$ and returns $\widehat K=6$, with
kinks at 2007, 2009, 2010, 2019, 2020 and 2021.

\begin{figure}[htbp]
    \centering
    \includegraphics[width=0.74\linewidth]{fig_dg_ic.pdf}
    \caption{\footnotesize{Selection of the tuning parameter in the
    debt-only benchmark. The solid line is the post-kink information
    criterion along the penalised path, on the left axis. The step function
    is the number of kinks retained, on the right axis. The dashed vertical
    line marks $\widehat\vartheta_1$.}}
    \label{fig:dg_ic}
\end{figure}

Figure~\ref{fig:dg_main} reports the estimated path, the within-regime
slopes and the cross-country debt ratio. The coefficient begins at
$-0.025$ and climbs at $0.0036$ per year until it reaches zero in 2007. It
then falls to $-0.083$ by 2009 and returns to $-0.036$ in 2010, after
which it is flat for nine years, the 2010 to 2019 slope being $0.0002$
with a standard error of $0.0007$. The pandemic produces the same fall and
rebound in compressed form. The coefficient drops to $-0.095$ in 2020,
returns to zero in 2021, and then declines at $-0.0113$ per year to
$-0.034$ in 2024. Pooled OLS and the within estimator give $-0.020$ and
$-0.030$, values that run through the middle of the path and conceal both
movements.

\begin{figure}[htbp]
    \centering
    \includegraphics[width=0.80\linewidth]{fig_dg_main.pdf}
    \caption{\footnotesize{The debt--growth association, 2000--2024. Upper
    panel: the post-kink debt coefficient with its pointwise 95 percent
    economy-clustered confidence band and the pooled and within
    benchmarks. Middle panel: the within-regime slopes with confidence
    intervals and printed estimates, starred at the one, five and ten
    percent levels. Lower panel: the cross-country mean and median debt
    ratio. Shaded regions are the estimated regimes.}}
    \label{fig:dg_main}
\end{figure}

Two clusters of dates therefore carry the crisis and the pandemic, and the
estimator distinguishes them by what happens after the rebound rather than
by the size of the fall. Both reach a similar trough. Only the pandemic
movement closes, with the coefficient back at zero within a year, whereas
after 2010 the coefficient settles near $-0.035$ and stays there for the
rest of the decade. A model that assigns a level to each regime can date
the two movements but cannot make this distinction, because it has no
within-regime rate to compare.

That flat decade is itself the more surprising feature. It runs from 2010
to 2019, years over which the cross-country mean debt ratio rises from 45
to 59 percent of GDP, and the coefficient does not become more negative
anywhere along the way. This is not a test of a threshold, since the model
dates common changes in calendar time rather than crossings by individual
economies. It does show that a large and sustained rise in average
indebtedness need not be accompanied by a progressively more negative
association, which sits closer to the literature emphasising debt
trajectories and heterogeneous environments than to a stable universal
cutoff \citep{PescatoriSandriSimon2014, EberhardtPresbitero2015,
ChudikMohaddesPesaranRaissi2017}. Over much of the same period the decline
in safe interest rates documented by \citet{Blanchard2019} also made the
economic meaning of any given debt ratio more dependent on the surrounding
environment.

\subsection{Coefficient-specific kinks}\label{sec:dg_cw}

A regression carrying one time-varying coefficient must assign to it
whatever common movement in the data happens to correlate with its
regressor. The dates in Figure~\ref{fig:dg_main} are therefore dates at
which \emph{something} in the growth process turned, and debt is only the
variable through which the turn is measured. The component-wise estimator
of Section~\ref{sec:component} lets each coefficient carry its own kink
dates, so the question of which variable turned when becomes an estimable
one.

We add total investment, population growth and inflation. Let
$\*x_{i,t}=[d_{i,t},q_{i,t},n_{i,t},\pi_{i,t}]'$, where $q_{i,t}$ is total
investment as a share of GDP, $n_{i,t}$ is population growth and
$\pi_{i,t}$ is average consumer price inflation, and estimate
\begin{align}
    g_{i,t}
    &=\mu_i+\sum_{j=1}^{4}\theta_{j,t}x_{i,t,j}+\varepsilon_{i,t},
    \label{eq:dg_cw_model}\\[-2pt]
    \theta_{j,t}
    &=\theta_{j,1}+\kappa_{j,1}(t-1)
      +\sum_{m=1}^{K_j}\gamma_{j,m}(t-T_{j,m})_+,
    \qquad
    \gamma_{j,m}:=\kappa_{j,m+1}-\kappa_{j,m},
    \label{eq:dg_cw_theta}
\end{align}
so that coefficient $j$ has its own number of kinks $K_j$, its own dates
and its own regime slopes. Estimation is by the component-wise objective
\eqref{eq:objective_componentwise} with $\zeta_2=2$, everything else as
above. Requiring investment and inflation in every year leaves $N=136$
economies, and regressors are standardised before the penalty is applied.
Figure~\ref{fig:dg_cw_ic} reports the selection, which retains fourteen
kinks: five in debt, at 2006, 2008, 2019, 2020 and 2021, six in
investment, at 2007, 2009, 2010, 2019, 2020 and 2021, two in inflation, at
2008 and 2015, and one in population growth, at 2009.

\begin{figure}[ht]
    \centering
    \includegraphics[width=0.78\linewidth]{fig_cw_ic.pdf}
    \caption{\footnotesize{Selection of the tuning parameter in the
    four-regressor model. Axes as in Figure~\ref{fig:dg_ic}, with one step
    function per coefficient.}}
    \label{fig:dg_cw_ic}
\end{figure}

Investment's six dates are exactly the six the scalar model assigned to
debt. The sharp reversal of 2007 to 2010, and the pandemic movement that
follows it, belong to the investment coefficient once investment is allowed
its own clock, and the debt-only regression was measuring them through the
one regressor it had. Debt keeps neither of the two crisis dates. It bends
in 2006 instead, two years before the crisis, and its 2008 kink opens a
recovery rather than a fall.

\begin{figure}[p]
    \centering
    \begin{minipage}[t]{0.48\linewidth}\centering
        \includegraphics[width=\linewidth]{fig_cw_path_debt.pdf}
    \end{minipage}\hfill
    \begin{minipage}[t]{0.48\linewidth}\centering
        \includegraphics[width=\linewidth]{fig_cw_kappa_debt.pdf}
    \end{minipage}

    \vspace{5pt}
    \begin{minipage}[t]{0.48\linewidth}\centering
        \includegraphics[width=\linewidth]{fig_cw_path_inv.pdf}
    \end{minipage}\hfill
    \begin{minipage}[t]{0.48\linewidth}\centering
        \includegraphics[width=\linewidth]{fig_cw_kappa_inv.pdf}
    \end{minipage}

    \vspace{5pt}
    \begin{minipage}[t]{0.48\linewidth}\centering
        \includegraphics[width=\linewidth]{fig_cw_path_popg.pdf}
    \end{minipage}\hfill
    \begin{minipage}[t]{0.48\linewidth}\centering
        \includegraphics[width=\linewidth]{fig_cw_kappa_popg.pdf}
    \end{minipage}

    \vspace{5pt}
    \begin{minipage}[t]{0.48\linewidth}\centering
        \includegraphics[width=\linewidth]{fig_cw_path_infl.pdf}
    \end{minipage}\hfill
    \begin{minipage}[t]{0.48\linewidth}\centering
        \includegraphics[width=\linewidth]{fig_cw_kappa_infl.pdf}
    \end{minipage}

    \caption{\footnotesize{Component-wise estimates, 2000--2024. Rows are
    gross government debt, total investment, population growth and
    inflation. Left column: the coefficient path with its pointwise 95
    percent country-clustered band and the pooled and within benchmarks.
    Right column: the within-regime slopes with confidence intervals and
    printed values, starred at the one, five and ten percent levels. Shaded
    regions are the estimated regimes.}}
    \label{fig:dg_cw_panels}
\end{figure}

Figure~\ref{fig:dg_cw_panels} gives the four paths and their slopes. Read
row by row, they say that the crisis moved three of the four coefficients
and that only one of the three came back. Debt is flat at about $-0.019$ through the boom, deteriorates from 2006 at
$-0.008$ per year, and then recovers for eleven years at $+0.004$ per year
to a 2019 value at which $\theta_{1,2019}=0$ cannot be rejected
($\mathcal{W}=0.43$, chi-squared with one degree of freedom throughout).
The rate after 2008 differs from the rate before 2006 ($\mathcal{W}=6.01$)
and the 2006 level is not recovered ($\mathcal{W}=10.81$), so the crisis
put the coefficient on a new path rather than displacing it temporarily.
Around the pandemic the two rates again differ ($\mathcal{W}=10.44$) while
the level is restored ($\mathcal{W}=1.34$, not rejected). The 2006 turn
precedes the crisis, which is where the informative variation sits in the
account of \citet{SchularickTaylor2012}, though their leading indicator is
private credit growth rather than the public debt ratio used here.

Investment shows the same asymmetry more sharply. It is flat until 2007,
falls at $-0.089$ per year to 2009 and rebounds at $+0.141$ into 2010, but
the round trip does not close on either margin ($\mathcal{W}=6.86$ on the
rates and $\mathcal{W}=7.13$ on the level). What follows is a new regime
declining at $-0.008$ per year for nine years, the reduced form signature
of the persistent output losses documented by \citet{CerraSaxena2008} and
of the weak capital formation that \citet{Summers2015} places at the centre
of the post-crisis decade. Its pandemic movement, by contrast, closes on
both margins ($\mathcal{W}=0.75$ and $\mathcal{W}=0.00$), which is what
separates a year of suspended investment from a decade of foregone
investment.

Inflation supplies the clearest reversal. The coefficient rises at
$+0.016$ per year to 2008, so that before the crisis an economy with above
average inflation was typically an economy running hot, and then falls at
$-0.043$ for seven years to $-0.213$. A single kink covers those seven
years, which is the substantive point, since a sustained reorientation and
a sequence of annual shocks look nothing alike in this parameterisation.
Its duration is compositional. Advanced economies moved to the effective
lower bound with persistent disinflation and weak growth, over a period in
which \citet{BallMazumder2011} and \citet{CoibionGorodnichenko2015} find
the inflation and activity relation much looser than standard
specifications imply, while emerging and developing economies were
completing a long disinflation of their own in which exchange rate
movements pass through more strongly where policy frameworks are weaker
\citep{HaKoseOhnsorge2019}. Economies crossed between the two
configurations a few at a time, which is what produces a slope rather than
a jump. The 2015 kink falls inside the oil price collapse that began in
June 2014 and that \citet{BaumeisterKilian2016} trace largely to demand
and to market specific developments predating it. After 2015 the
coefficient rises at a rate indistinguishable from its pre-2008 rate
($\mathcal{W}=0.00$) and returns to its initial level
($\mathcal{W}=0.41$). The inflation of 2021 and 2022 falls inside that
final regime and generates no kink of its own.

Population growth carries one weak kink, at 2009. Neither of its two slopes
is significant at the five percent level, the change in rate is significant
only at ten percent ($\mathcal{W}=3.56$), and the initial level is restored
by 2024 ($\mathcal{W}=0.77$). It is the one coefficient here that neither
the crisis nor the pandemic moves.

\begin{figure}[ht]
    \centering
    \includegraphics[width=0.86\linewidth]{fig_cw_overview.pdf}
    \caption{\footnotesize{The debt coefficient across specifications.
    Upper panel: the scalar specification on the full sample and on the
    economies with full covariate coverage. Middle panel: the scalar,
    coefficient-specific and common-kink estimates on the common sample.
    Lower panel: the kink dates each selects. Shaded regions are pointwise
    95 percent country-clustered bands.}}
    \label{fig:dg_cw_overview}
\end{figure}

Figure~\ref{fig:dg_cw_overview} puts the four specifications on one axis
and separates what conditioning does from what the sample does. The sample
restriction is immaterial, since the scalar model estimated on the 136
economies with full covariate coverage selects the same six dates as on the
full sample. Conditioning changes both the timing, as described above, and
the amplitude, halving the 2020 trough from $-0.103$ to $-0.045$. The common-kink penalty, which forces one date set on the whole coefficient
vector, is the more instructive comparison, since it is the restriction our
extension removes. It selects seven dates. Six of them are investment's and
the seventh is inflation's 2015, while debt's own 2006 and 2008 are lost.
The restriction overwrites the timing of precisely the coefficient of
interest, because among the four its signal is the weakest.
Figure~\ref{fig:dg_cw_compare} shows what that costs elsewhere. Population
growth is given six kinks it does not need, and inflation keeps the date
that ends its decline but loses the 2008 date that begins it, which is the
one that organises the whole path.

What survives every specification is the shape and the conclusion drawn
from it. The debt coefficient is kinked, it bends around both global
movements, and no constant coefficient describes it. Two caveats bound the
reading. The specification carries no year effects, so a common
macroeconomic movement loads onto whichever coefficient covaries with it,
and \citet{BlanchardLeigh2013} show how tightly output, fiscal policy and
forecast errors were bound together in exactly the years in which these
coefficients move most. And the estimates remain associations. What the
exercise establishes is when each association changed, how fast, and which
of the changes reversed.

\begin{figure}[h!]
    \centering
    \includegraphics[width=0.94\linewidth]{fig_cw_vs_group.pdf}
    \caption{\footnotesize{Component-wise against common-kink estimates in
    the same four-regressor model. Panels are, in reading order, debt,
    investment, population growth and inflation. Solid paths use
    coefficient-specific kink sets, dashed paths the common dates.}}
    \label{fig:dg_cw_compare}
\end{figure}
\section{Conclusion}\label{sec:conclusion}

This paper studies panel data models in which the coefficients follow continuous piecewise-linear paths with an unknown number of kinks occurring at unknown dates. We estimate the kink dates using an adaptive group fused lasso applied to the second differences of the coefficient paths, show that the true dates are recovered exactly with probability approaching one, and establish the asymptotic normality of the post-kink within estimator. Continuity links adjacent regimes, so each parameter is informed by the regimes on either side of a kink and has its own convergence rate, without requiring balanced regime lengths. A component-wise extension also allows kink dates to differ across regressors. The usefulness of the design and the resulting estimator is demonstrated through an application in macro-finance, specifically by examining whether the relationship between debt and growth changes over calendar time.

\pagebreak

\section*{Declaration of generative AI and AI-assisted technologies in the manuscript preparation process}
During the preparation of this work the authors used Claude and ChatGPT in order to improve the language and presentation of the manuscript, to support the verification of mathematical arguments, and to accelerate the development of simulation code. After using this tool/service, the authors reviewed and edited the content as needed and take full responsibility for the content of the published article.

\pagebreak

\begin{table}[h]
\centering
\small
\caption{Selection performance, common-kink designs.}
\label{tab:mcsel}
\begin{tabular}{@{}cc ccc ccc ccc@{}}
\toprule
&& \multicolumn{3}{c}{DGP~1}
&  \multicolumn{3}{c}{DGP~2}
&  \multicolumn{3}{c}{DGP~3}\\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}\cmidrule(lr){9-11}
$N$ & $T$
& Corr. & $\widehat K\!-\!K^0$ & Miscl.
& Corr. & $\widehat K\!-\!K^0$ & Miscl.
& Corr. & $\widehat K\!-\!K^0$ & Miscl.\\
\midrule
 50 & 20 & 99.3 & 0.01 & 0.4 & 97.8 & 0.01 & 0.9 & 80.0 & 0.07 & 4.3 \\
 50 & 40 & 99.4 & 0.01 & 0.3 & 99.0 & 0.01 & 0.5 & 83.9 & 0.15 & 9.0 \\
 50 & 80 & 99.3 & 0.01 & 0.6 & 98.8 & 0.01 & 0.5 & 82.0 & 0.18 & 8.1 \\
\addlinespace
100 & 20 & 99.4 & 0.01 & 0.3 & 99.5 & 0.00 & 0.3 & 93.9 & 0.05 & 3.0 \\
100 & 40 & 99.9 & 0.00 & 0.0 & 99.5 & 0.01 & 0.2 & 94.5 & 0.06 & 3.4 \\
100 & 80 & 99.8 & 0.00 & 0.1 & 99.1 & 0.01 & 0.6 & 95.6 & 0.04 & 1.7 \\
\addlinespace
200 & 20 & 99.8 & 0.00 & 0.2 & 99.8 & 0.00 & 0.1 & 98.7 & 0.01 & 0.7 \\
200 & 40 & 99.5 & 0.01 & 0.2 & 99.9 & 0.00 & 0.1 & 97.9 & 0.02 & 1.2 \\
200 & 80 & 99.9 & 0.00 & 0.1 & 99.8 & 0.00 & 0.1 & 98.3 & 0.02 & 0.6 \\
\midrule
&& \multicolumn{3}{c}{DGP~5}
&  \multicolumn{3}{c}{DGP~6}
&  \multicolumn{3}{c}{DGP~7}\\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}\cmidrule(lr){9-11}
$N$ & $T$
& Corr. & $\widehat K\!-\!K^0$ & Miscl.
& Corr. & $\widehat K\!-\!K^0$ & Miscl.
& Corr. & $\widehat K\!-\!K^0$ & Miscl.\\
\midrule
 50 & 20 & 81.1 & 0.03 & 1.8 & 94.2 & 0.02 & 1.0 & 70.9 & 0.05 & 3.8 \\
 50 & 40 & 87.1 & 0.11 & 3.5 & 97.5 & 0.02 & 0.9 & 79.8 & 0.14 & 8.7 \\
 50 & 80 & 89.4 & 0.11 & 6.0 & 98.3 & 0.02 & 0.8 & 73.0 & 0.28 & 12.5 \\
\addlinespace
100 & 20 & 94.0 & 0.04 & 1.3 & 99.0 & 0.01 & 0.4 & 88.2 & 0.07 & 3.6 \\
100 & 40 & 93.8 & 0.06 & 1.9 & 99.1 & 0.01 & 0.4 & 85.0 & 0.15 & 9.3 \\
100 & 80 & 95.7 & 0.04 & 3.0 & 98.9 & 0.01 & 0.5 & 88.0 & 0.12 & 5.5 \\
\addlinespace
200 & 20 & 97.5 & 0.03 & 0.9 & 99.3 & 0.01 & 0.3 & 96.3 & 0.03 & 1.9 \\
200 & 40 & 99.2 & 0.01 & 0.2 & 99.6 & 0.00 & 0.2 & 94.2 & 0.06 & 3.7 \\
200 & 80 & 99.0 & 0.01 & 0.7 & 99.7 & 0.00 & 0.2 & 96.2 & 0.04 & 1.5 \\
\bottomrule
\end{tabular}
\par\smallskip
\begin{minipage}{0.92\textwidth}\footnotesize
\emph{Notes}: $1000$ replications. Corr.\ is the frequency, in
percent, of exact recovery of the true kink dates.
$\widehat K-K^0$ is the average kink-count bias. Miscl.\ is the
average percentage of dates assigned to the wrong regime, so that a
kink located one date off misclassifies $1/T$ of the sample while a
spurious interior kink misclassifies the share of dates beyond it.
DGP~5 has uneven regime lengths; DGPs~6 and~7 repeat DGPs~2 and~3
under $(\varpi_1,\varpi_2)=(0.5,0.2)$.
\end{minipage}
\end{table}

\pagebreak

\begin{table}[h]
\centering
\small
\caption{Selection performance of DGP~4 using component-wise estimator.}
\label{tab:mcselcw}
\begin{tabular}{@{}cc ccc ccc@{}}
\toprule
&& \multicolumn{3}{c}{Coordinate 1}
&  \multicolumn{3}{c}{Coordinate 2}\\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
$N$ & $T$
& Corr. & $\widehat K_1\!-\!K_1^0$ & Miscl.
& Corr. & $\widehat K_2\!-\!K_2^0$ & Miscl.\\
\midrule
 50 & 20 & 91.5 & 0.07 & 4.5 & 90.0 & 0.10 & 4.6 \\
 50 & 40 & 92.3 & 0.08 & 4.9 & 90.7 & 0.10 & 4.0 \\
 50 & 80 & 91.2 & 0.09 & 5.4 & 90.9 & 0.10 & 3.4 \\
\addlinespace
100 & 20 & 93.8 & 0.06 & 3.8 & 94.2 & 0.06 & 2.7 \\
100 & 40 & 93.3 & 0.07 & 4.0 & 93.8 & 0.06 & 2.6 \\
100 & 80 & 96.2 & 0.04 & 2.2 & 95.9 & 0.04 & 1.8 \\
\addlinespace
200 & 20 & 97.3 & 0.03 & 1.3 & 96.2 & 0.04 & 1.6 \\
200 & 40 & 97.1 & 0.03 & 1.8 & 97.3 & 0.03 & 1.2 \\
200 & 80 & 98.0 & 0.02 & 1.1 & 97.1 & 0.03 & 1.2 \\
\bottomrule
\end{tabular}
\par\smallskip
\begin{minipage}{0.85\textwidth}\footnotesize
\emph{Notes}: $1000$ replications. Statistics are as in
Table~\ref{tab:mcsel}, computed per coordinate.
\end{minipage}
\end{table}



\pagebreak

\begin{table}[h]
\centering
\small
\caption{Selection performance of DGP~8 using the component-wise
estimator.}
\label{tab:mcselcw8}
\begin{tabular}{@{}cc ccc ccc ccc@{}}
\toprule
&& \multicolumn{3}{c}{Coordinate 1 ($K_1^0=0$)}
&  \multicolumn{3}{c}{Coordinate 2 ($K_2^0=1$)}
&  \multicolumn{3}{c}{Coordinate 3 ($K_3^0=2$)}\\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}\cmidrule(lr){9-11}
$N$ & $T$
& Corr. & $\widehat K_1\!-\!K_1^0$ & Miscl.
& Corr. & $\widehat K_2\!-\!K_2^0$ & Miscl.
& Corr. & $\widehat K_3\!-\!K_3^0$ & Miscl.\\
\midrule
 50 & 20 & 95.0 & 0.05 & 2.7 & 90.7 & 0.10 & 5.2 & 61.2 & 0.20 & 10.6 \\
 50 & 40 & 95.5 & 0.04 & 2.4 & 92.4 & 0.08 & 4.3 & 64.1 & 0.36 & 18.9 \\
 50 & 80 & 96.4 & 0.04 & 1.8 & 90.6 & 0.10 & 5.0 & 69.8 & 0.35 & 15.4 \\
\addlinespace
100 & 20 & 97.7 & 0.03 & 1.4 & 93.7 & 0.07 & 3.6 & 79.4 & 0.17 & 9.1 \\
100 & 40 & 97.3 & 0.03 & 1.2 & 95.1 & 0.05 & 2.4 & 81.7 & 0.20 & 10.3 \\
100 & 80 & 96.9 & 0.03 & 1.6 & 94.5 & 0.06 & 2.7 & 83.7 & 0.18 & 7.5 \\
\addlinespace
200 & 20 & 97.6 & 0.03 & 1.1 & 97.5 & 0.03 & 1.1 & 91.5 & 0.09 & 4.6 \\
200 & 40 & 98.3 & 0.02 & 0.7 & 97.2 & 0.03 & 1.5 & 93.2 & 0.07 & 3.7 \\
200 & 80 & 98.0 & 0.02 & 1.0 & 97.7 & 0.02 & 1.0 & 93.4 & 0.07 & 3.3 \\
\bottomrule
\end{tabular}
\par\smallskip
\begin{minipage}{0.90\textwidth}\footnotesize
\emph{Notes}: $1000$ replications. Statistics are as in
Table~\ref{tab:mcsel}, computed per coordinate. The three coordinates
carry a different number of kinks, so the design tests whether the
estimator recovers heterogeneous kink counts, including none in the
first coordinate.
\end{minipage}
\end{table}


\pagebreak

\begin{table}[h]
\centering
\small
\caption{Bias and RMSE of the slope and path estimates,
DGPs~1 to~3.}
\label{tab:mcest}
\begin{tabular}{@{}cc c ccc cccc@{}}
\toprule
&& DGP~1
&  \multicolumn{3}{c}{DGP~2}
&  \multicolumn{4}{c}{DGP~3}\\
\cmidrule(lr){3-3}\cmidrule(lr){4-6}\cmidrule(lr){7-10}
$N$ & $T$ & Path
& $\+\kappa_1$ & $\+\kappa_2$ & Path
& $\+\kappa_1$ & $\+\kappa_2$ & $\+\kappa_3$ & Path\\
\midrule
\multicolumn{10}{@{}l}{\emph{Panel A: Bias} ($\times1000$)}\\
\midrule
 50 & 20 & $-0.89$ & $-0.11$ & 0.19 & $-0.77$ & $-0.56$ & 0.62 & $-0.81$ & 0.11 \\
 50 & 40 & $-0.22$ & 0.07 & $-0.06$ & 0.57 & $-0.47$ & 0.13 & 0.12 & $-0.38$ \\
 50 & 80 & $-0.22$ & $-0.08$ & 0.08 & $-0.76$ & $-0.15$ & 0.00 & 0.07 & $-0.36$ \\
\addlinespace
100 & 20 & $-0.46$ & $-0.12$ & $-0.30$ & $-0.09$ & 0.23 & 0.36 & $-0.34$ & $-0.07$ \\
100 & 40 & 0.50 & 0.00 & $-0.09$ & $-0.15$ & $-0.29$ & 0.06 & 0.01 & $-0.02$ \\
100 & 80 & 0.23 & 0.01 & $-0.04$ & $-0.03$ & $-0.03$ & 0.04 & $-0.02$ & 0.00 \\
\addlinespace
200 & 20 & $-0.14$ & $-0.14$ & 0.01 & 0.15 & $-0.04$ & $-0.33$ & 0.46 & $-0.81$ \\
200 & 40 & $-0.55$ & 0.07 & $-0.07$ & $-0.44$ & $-0.02$ & $-0.02$ & 0.00 & $-0.47$ \\
200 & 80 & $-0.41$ & 0.00 & 0.00 & 0.11 & 0.00 & 0.02 & 0.00 & 0.09 \\
\midrule
\multicolumn{10}{@{}l}{\emph{Panel B: RMSE}}\\
\midrule
 50 & 20 & 0.049 & 0.014 & 0.012 & 0.062 & 0.030 & 0.019 & 0.021 & 0.082 \\
 50 & 40 & 0.035 & 0.005 & 0.005 & 0.042 & 0.009 & 0.007 & 0.007 & 0.054 \\
 50 & 80 & 0.024 & 0.002 & 0.002 & 0.029 & 0.003 & 0.002 & 0.003 & 0.038 \\
\addlinespace
100 & 20 & 0.035 & 0.010 & 0.009 & 0.041 & 0.021 & 0.013 & 0.014 & 0.051 \\
100 & 40 & 0.024 & 0.003 & 0.003 & 0.029 & 0.006 & 0.005 & 0.005 & 0.035 \\
100 & 80 & 0.017 & 0.001 & 0.001 & 0.021 & 0.002 & 0.002 & 0.002 & 0.025 \\
\addlinespace
200 & 20 & 0.024 & 0.007 & 0.006 & 0.030 & 0.015 & 0.009 & 0.010 & 0.034 \\
200 & 40 & 0.017 & 0.002 & 0.002 & 0.021 & 0.004 & 0.004 & 0.004 & 0.025 \\
200 & 80 & 0.012 & 0.001 & 0.001 & 0.015 & 0.002 & 0.001 & 0.001 & 0.017 \\
\bottomrule
\end{tabular}
\par\smallskip
\begin{minipage}{0.92\textwidth}\footnotesize
\emph{Notes}: $1000$ replications. Bias entries are multiplied by
$1000$. Slope statistics are per regime, pooled across the two
coordinates, and conditional on exact recovery of the kink set. Path
statistics are averaged over dates and coordinates and are
unconditional. In DGP~1 the fitted path under $\widehat K=0$ is a
single linear regime.
\end{minipage}
\end{table}










\pagebreak
\begin{table}[h]
\centering
\footnotesize
\caption{Bias and RMSE of the slope and path estimates, DGPs~5 to~7.}
\label{tab:mcestrob}
\begin{tabular}{@{}cc cccc ccc cccc@{}}
\toprule
&& \multicolumn{4}{c}{DGP~5}
&  \multicolumn{3}{c}{DGP~6}
&  \multicolumn{4}{c}{DGP~7}\\
\cmidrule(lr){3-6}\cmidrule(lr){7-9}\cmidrule(lr){10-13}
$N$ & $T$
& $\+\kappa_1$ & $\+\kappa_2$ & $\+\kappa_3$ & Path
& $\+\kappa_1$ & $\+\kappa_2$ & Path
& $\+\kappa_1$ & $\+\kappa_2$ & $\+\kappa_3$ & Path\\
\midrule
\multicolumn{13}{@{}l}{\emph{Panel A: Bias} ($\times1000$)}\\
\midrule
 50 & 20 & $-0.27$ & 0.15 & $-0.73$ & 0.32 & $-0.23$ & 0.82 & $-1.05$ & $-1.61$ & 1.01 & $-0.17$ & 1.03 \\
 50 & 40 & 0.58 & $-0.14$ & 0.63 & 1.05 & 0.04 & 0.04 & 0.31 & 0.27 & 0.00 & 0.14 & 0.47 \\
 50 & 80 & $-0.13$ & 0.07 & $-0.28$ & $-0.34$ & 0.03 & $-0.11$ & 0.55 & 0.04 & $-0.05$ & $-0.01$ & $-0.19$ \\
\addlinespace
100 & 20 & 0.24 & $-0.10$ & $-0.01$ & 0.27 & $-0.22$ & 0.60 & 0.11 & $-1.60$ & 0.44 & $-0.15$ & $-0.51$ \\
100 & 40 & $-0.09$ & 0.00 & $-0.15$ & 0.50 & 0.00 & $-0.18$ & $-0.28$ & $-0.25$ & $-0.12$ & 0.10 & 0.22 \\
100 & 80 & 0.03 & $-0.01$ & $-0.02$ & 0.25 & $-0.02$ & $-0.02$ & 0.20 & 0.12 & $-0.06$ & $-0.01$ & 0.13 \\
\addlinespace
200 & 20 & 0.48 & 0.13 & $-0.43$ & 0.22 & 0.35 & $-0.12$ & $-0.09$ & $-0.28$ & $-0.22$ & 0.49 & 0.67 \\
200 & 40 & $-0.08$ & 0.01 & $-0.11$ & $-0.17$ & 0.03 & $-0.08$ & $-0.60$ & 0.19 & $-0.04$ & 0.02 & 0.07 \\
200 & 80 & 0.08 & $-0.02$ & 0.07 & $-0.12$ & 0.02 & $-0.01$ & $-0.42$ & $-0.02$ & 0.05 & $-0.04$ & 0.21 \\
\midrule
\multicolumn{13}{@{}l}{\emph{Panel B: RMSE}}\\
\midrule
 50 & 20 & 0.051 & 0.010 & 0.037 & 0.080 & 0.017 & 0.016 & 0.079 & 0.036 & 0.022 & 0.024 & 0.100 \\
 50 & 40 & 0.016 & 0.004 & 0.014 & 0.054 & 0.006 & 0.006 & 0.054 & 0.011 & 0.009 & 0.010 & 0.070 \\
 50 & 80 & 0.005 & 0.001 & 0.005 & 0.036 & 0.002 & 0.002 & 0.038 & 0.004 & 0.003 & 0.004 & 0.051 \\
\addlinespace
100 & 20 & 0.036 & 0.007 & 0.026 & 0.051 & 0.012 & 0.011 & 0.051 & 0.026 & 0.016 & 0.018 & 0.067 \\
100 & 40 & 0.012 & 0.003 & 0.010 & 0.036 & 0.004 & 0.004 & 0.037 & 0.008 & 0.006 & 0.007 & 0.047 \\
100 & 80 & 0.004 & 0.001 & 0.004 & 0.025 & 0.002 & 0.001 & 0.027 & 0.003 & 0.002 & 0.002 & 0.033 \\
\addlinespace
200 & 20 & 0.026 & 0.005 & 0.019 & 0.035 & 0.008 & 0.008 & 0.037 & 0.018 & 0.011 & 0.013 & 0.044 \\
200 & 40 & 0.008 & 0.002 & 0.007 & 0.024 & 0.003 & 0.003 & 0.027 & 0.006 & 0.005 & 0.005 & 0.032 \\
200 & 80 & 0.003 & 0.001 & 0.003 & 0.017 & 0.001 & 0.001 & 0.019 & 0.002 & 0.002 & 0.002 & 0.023 \\
\bottomrule
\end{tabular}
\par\smallskip
\begin{minipage}{0.92\textwidth}\footnotesize
\emph{Notes}: $1000$ replications. Bias entries are multiplied by
$1000$. Slope statistics are per regime, pooled across the two
coordinates, and conditional on exact recovery of the kink set. Path
statistics are averaged over dates and coordinates and are
unconditional. In DGP~5 the outer regimes are short and the middle
regime long, so $\+\kappa_2$ is informed by two long adjacent regimes
while $\+\kappa_1$ and $\+\kappa_3$ are informed by a short own
regime. DGPs~6 and~7 repeat DGPs~2 and~3 under
$(\varpi_1,\varpi_2)=(0.5,0.2)$.
\end{minipage}
\end{table}


\pagebreak

\begin{table}[h]
\centering
\small
\caption{Bias and RMSE of DGP~4 using component-wise estimator.}
\label{tab:mccw}
\begin{tabular}{@{}cc ccc ccc@{}}
\toprule
&& \multicolumn{3}{c}{Coordinate 1}
&  \multicolumn{3}{c}{Coordinate 2}\\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
$N$ & $T$
& Regime 1 & Regime 2 & Path
& Regime 1 & Regime 2 & Path\\
\midrule
\multicolumn{8}{@{}l}{\emph{Panel A: Bias} ($\times1000$)}\\
\midrule
 50 & 20 & $-1.21$ & 0.07 & $-0.31$ & $-0.13$ & 0.34 & $-1.40$ \\
 50 & 40 & $-0.05$ & 0.12 & $-0.62$ & 0.09 & 0.05 & $-0.87$ \\
 50 & 80 & $-0.08$ & 0.05 & 0.33 & 0.01 & $-0.09$ & $-0.58$ \\
\addlinespace
100 & 20 & $-0.35$ & $-0.09$ & 0.74 & 0.17 & $-0.61$ & 0.04 \\
100 & 40 & $-0.09$ & $-0.02$ & 0.69 & $-0.05$ & $-0.13$ & 0.18 \\
100 & 80 & $-0.05$ & 0.02 & $-0.08$ & $-0.02$ & 0.09 & $-0.13$ \\
\addlinespace
200 & 20 & $-0.32$ & 0.07 & $-0.62$ & $-0.06$ & $-0.23$ & 0.24 \\
200 & 40 & $-0.03$ & 0.02 & $-0.64$ & $-0.09$ & 0.22 & $-0.41$ \\
200 & 80 & 0.02 & $-0.01$ & $-0.07$ & $-0.03$ & 0.02 & 0.01 \\
\midrule
\multicolumn{8}{@{}l}{\emph{Panel B: RMSE}}\\
\midrule
 50 & 20 & 0.027 & 0.009 & 0.068 & 0.010 & 0.019 & 0.068 \\
 50 & 40 & 0.008 & 0.003 & 0.046 & 0.003 & 0.007 & 0.048 \\
 50 & 80 & 0.003 & 0.001 & 0.033 & 0.001 & 0.003 & 0.033 \\
\addlinespace
100 & 20 & 0.020 & 0.006 & 0.047 & 0.007 & 0.013 & 0.045 \\
100 & 40 & 0.006 & 0.002 & 0.032 & 0.002 & 0.005 & 0.032 \\
100 & 80 & 0.002 & 0.001 & 0.022 & 0.001 & 0.002 & 0.022 \\
\addlinespace
200 & 20 & 0.014 & 0.004 & 0.031 & 0.005 & 0.009 & 0.031 \\
200 & 40 & 0.004 & 0.002 & 0.022 & 0.002 & 0.003 & 0.022 \\
200 & 80 & 0.001 & 0.001 & 0.015 & 0.001 & 0.001 & 0.015 \\
\bottomrule
\end{tabular}
\par\smallskip
\begin{minipage}{0.85\textwidth}\footnotesize
\emph{Notes}: $1000$ replications. Bias entries are multiplied by
$1000$. Slope statistics are conditional on exact recovery of that
coordinate's kink set. Path statistics are averaged over dates and
are unconditional.
\end{minipage}
\end{table}


\pagebreak

\begin{table}[htbp]
\centering
\caption{Bias and RMSE of DGP~8 using the component-wise estimator.}
\label{tab:mccw8}
\begingroup
\footnotesize
\setlength{\tabcolsep}{3.5pt}
\begin{adjustbox}{max width=\textwidth}
\begin{tabular}{@{}cc cc ccc cccc@{}}
\toprule
&& \multicolumn{2}{c}{Coordinate 1}
&  \multicolumn{3}{c}{Coordinate 2}
&  \multicolumn{4}{c}{Coordinate 3} \\
\cmidrule(lr){3-4}
\cmidrule(lr){5-7}
\cmidrule(lr){8-11}
$N$ & $T$
& Regime 1 & Path
& Regime 1 & Regime 2 & Path
& Regime 1 & Regime 2 & Regime 3 & Path \\
\midrule
\multicolumn{11}{@{}l}{\emph{Panel A: Bias} ($\times1000$)} \\
\midrule
 50 & 20 & $-0.11$ & $-0.74$ & 0.25 & 0.27 & $-0.04$ & $-0.15$ & 1.34 & $-1.08$ & $-0.52$ \\
 50 & 40 & 0.08 & 0.01 & 0.09 & $-0.11$ & 0.81 & $-0.18$ & 0.08 & $-0.20$ & 0.13 \\
 50 & 80 & 0.00 & 0.28 & $-0.04$ & $-0.08$ & $-0.39$ & $-0.10$ & 0.04 & $-0.17$ & $-1.16$ \\
\addlinespace
100 & 20 & 0.03 & 0.58 & $-0.09$ & $-0.25$ & $-0.38$ & $-1.82$ & 0.94 & $-1.30$ & $-0.83$ \\
100 & 40 & 0.00 & $-0.39$ & $-0.13$ & 0.06 & 0.83 & $-0.31$ & 0.31 & $-0.24$ & 0.32 \\
100 & 80 & 0.01 & $-0.13$ & 0.04 & $-0.05$ & $-0.25$ & $-0.15$ & 0.12 & $-0.14$ & $-0.03$ \\
\addlinespace
200 & 20 & $-0.12$ & $-0.39$ & 0.41 & 0.08 & 0.08 & 0.02 & 0.17 & 0.12 & 0.68 \\
200 & 40 & 0.05 & 0.10 & $-0.05$ & 0.05 & $-0.26$ & $-0.05$ & $-0.03$ & 0.15 & 0.23 \\
200 & 80 & $-0.01$ & 0.22 & $-0.01$ & 0.01 & 0.01 & $-0.01$ & $-0.01$ & $-0.01$ & $-0.14$ \\
\midrule
\multicolumn{11}{@{}l}{\emph{Panel B: RMSE}} \\
\midrule
 50 & 20 & 0.006 & 0.055 & 0.014 & 0.012 & 0.067 & 0.030 & 0.017 & 0.020 & 0.096 \\
 50 & 40 & 0.002 & 0.037 & 0.005 & 0.004 & 0.046 & 0.009 & 0.007 & 0.008 & 0.068 \\
 50 & 80 & 0.001 & 0.026 & 0.002 & 0.002 & 0.033 & 0.003 & 0.002 & 0.003 & 0.045 \\
\addlinespace
100 & 20 & 0.004 & 0.036 & 0.010 & 0.009 & 0.046 & 0.021 & 0.012 & 0.014 & 0.062 \\
100 & 40 & 0.001 & 0.025 & 0.003 & 0.003 & 0.032 & 0.006 & 0.005 & 0.005 & 0.041 \\
100 & 80 & 0.001 & 0.018 & 0.001 & 0.001 & 0.023 & 0.002 & 0.002 & 0.002 & 0.028 \\
\addlinespace
200 & 20 & 0.003 & 0.026 & 0.007 & 0.006 & 0.031 & 0.016 & 0.009 & 0.010 & 0.038 \\
200 & 40 & 0.001 & 0.018 & 0.002 & 0.002 & 0.022 & 0.005 & 0.003 & 0.004 & 0.027 \\
200 & 80 & 0.000 & 0.012 & 0.001 & 0.001 & 0.015 & 0.001 & 0.001 & 0.001 & 0.018 \\
\bottomrule
\end{tabular}
\end{adjustbox}
\endgroup
\par\smallskip
\begin{minipage}{\textwidth}
\footnotesize
\emph{Notes}: Based on $1{,}000$ replications. Bias entries are
multiplied by $1000$. Slope statistics are conditional on exact
recovery of the corresponding coordinate's kink set. Path statistics
are averaged over dates and are unconditional. Coordinate~1 is
globally linear and therefore has a single regime slope.
\end{minipage}
\end{table}

\pagebreak


\bibliography{kink_paper}