Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Improved Inference on the Rank of a Matrix
bibunit\pdfbookmark[1]{Title}{title}
\begin{abstract}
This paper develops a general framework for conducting inference on the rank of an unknown matrix $\Pi_0$. A defining feature of our setup is the null hypothesis of the form $\mathrm H_0: \mathrm{rank}(\Pi_0)\le r$. The problem is of first order importance because the previous literature focuses on $\mathrm H_0': \mathrm{rank}(\Pi_0)= r$ by implicitly assuming away $\mathrm{rank}(\Pi_0)<r$, which may lead to invalid rank tests due to over-rejections. In particular, we show that limiting distributions of test statistics under $\mathrm H_0'$ may not stochastically dominate those under $\mathrm{rank}(\Pi_0)<r$. A multiple test on the nulls $\mathrm{rank}(\Pi_0)=0,\ldots,r$, though valid, may be substantially conservative. We employ a testing statistic whose limiting distributions under $\mathrm H_0$ are highly nonstandard due to the inherent irregular natures of the problem, and then construct bootstrap critical values that deliver size control and improved power. Since our procedure relies on a tuning parameter, a two-step procedure is designed to mitigate concerns on this nuisance. We additionally argue that our setup is also important for estimation. We illustrate the empirical relevance of our results through testing identification in linear IV models that allows for clustered data and inference on sorting dimensions in a two-sided matching model with transferrable utility.
\end{abstract}
\begin{center}
Keywords: Matrix rank, Bootstrap, Two-step test, Rank estimation, Identification, Matching dimension
\end{center}
\section{Introduction}
The rank of a matrix plays a number of fundamental roles in economics, not just as crucial technical identification conditions Fisher1966IDBook, but also of central empirical relevance in numerous settings such as inference on cointegration rank Engle_Granger1987Co-In,Johansen1991CoIntegration, specification of finite mixture models McLachlanPeel2004Mixture,KasaharaShimotsu2009NPIDmixture and estimation of matching dimensions DupuyGalichon2014Personality -- more can be found in Supplemental Appendix (ref). These problems reduce to examining the hypotheses: for an unknown matrix $\Pi_0$ of size $m\times k$ with $m\ge k$,
\begin{align}
\mathrm H_0: \mathrm{rank}(\Pi_0)\le r \qquad v.s. \qquad \mathrm H_1: \mathrm{rank}(\Pi_0)> r ,
\end{align}
where $r\in\{0,\ldots,k-1\}$ is some prespecified value and $\mathrm{rank}(\Pi_0)$ denotes the rank of $\Pi_{0}$. If $r=k-1$, then (ref) is concerned with whether $\Pi_0$ has full rank.
Despite a rich set of results in the literature, previous studies instead focus on
\begin{align}
\mathrm H_0': \mathrm{rank}(\Pi_0)= r \qquad v.s. \qquad \mathrm H_1: \mathrm{rank}(\Pi_0)> r .
\end{align}
In effect, the testing problem (ref) assumes away the possibility $\mathrm{rank}(\Pi_0)< r$, which is often unrealistic to be excluded. This, unfortunately, has drastic consequences. As elaborated through an analytic example in Section (ref), a number of popular tests, including Robin_Smith2000rank and Kleibergen_Paap2006rank, may over-reject for some data generating processes and under-reject for others, both having $\mathrm{rank}(\Pi_0)< r$. In particular, contrary to what appears to have been conjectured in the literature (Cragg_Donald1993TestID, Cragg_Donald1993TestID, p.225; Johansen1995likelihood, Johansen1995likelihood, p.168), our analysis suggests that {\it limiting distributions of tests obtained under $\mathrm H_0'$ may not first order stochastically dominate those under $\mathrm{rank}(\Pi_0)< r$}. Hence, ignoring the possibility $\mathrm{rank}(\Pi_0)< r$ may lead to tests that are not even first order valid.
One may nonetheless justify the setup (ref) for two reasons. First, the problem (ref) may be studied by a multiple test on the nulls $\mathrm{rank}(\Pi_0)=0,1,\ldots,r$. Our simulations show, however, that such a procedure, though valid, may be substantially conservative and have trivial power against local alternatives that are close to matrices whose rank is strictly less than $r$. Second, the setup (ref) suits well for estimation by sequentially testing $\mathrm{rank}(\Pi_0)=j$ for $j=0,1,\ldots,k-1$. Crucially, however, all steps except for $j=0$ ignore type I errors (false rejection) potentially made in previous steps, and may have limited capability of controlling type II errors (false acceptance) -- see Supplemental Appendix (ref) for more details. Hence, the setup (ref) is desirable for estimation as well.
We thus conclude that developing a valid and powerful test for (ref) is of first order importance. To the best of our knowledge, no direct tests to date exist in this regard. Our objective in this paper is therefore to develop an inferential framework under the setup (ref). A key insight we exploit to this end is that (ref) is equivalent to
\begin{align}
\mathrm H_0: \phi_r(\Pi_0)=0 \qquad v.s. \qquad \mathrm H_1: \phi_r(\Pi_0)>0 ,
\end{align}
where $\phi_r(\Pi_0)\equiv\sum_{j=r+1}^k\sigma_j^{2}(\Pi_0)$ is the sum of the $k-r$ smallest squared singular values $\sigma_j^{2}(\Pi_0)$ of $\Pi_0$ -- see Supplemental Appendix for a review on singular values. Such a reformulation is attractive because it converts an unwieldy inference problem on an integer-valued parameter (i.e., rank) into a more tractable one on a real-valued functional (i.e., a sum of singular values). Given an estimator $\hat\Pi_n$ of $\Pi_0$, it is thus natural to base the testing statistic on the plug-in estimator $\phi_r(\hat\Pi_n)$ and then invoke the Delta method. As it turns out, the formulation (ref) reveals two crucial irregular natures involved, namely, $\phi_r$ admits a zero first order derivative under $\mathrm{H}_0$ and is second order nondifferentiable precisely when $\mathrm{rank}(\Pi_0)< r$ -- see Proposition (ref) and Lemma (ref). While the null limiting distributions of $\phi_r(\hat\Pi_n)$ can nonetheless be derived by existing generalizations of the Delta method Shapiro2000inference, constructions of critical values are nontrivial because the limits are non-pivotal and highly nonstandard. In particular, they depend on the true rank (among other things), upholding the importance of taking into account the possibility $\mathrm{rank}(\Pi_0)<r$. For this, we appeal to modified bootstrap schemes recently developed by Fang_Santos2014HDD and Chen_Fang2015FOD, which yield tests for (ref) that have asymptotically pointwise exact size control and are consistent. We further characterize analytically classes of local perturbations of the data generating processes under which our tests enjoy size control and nontrivial power.
A common feature of our tests is their dependence on tuning parameters, although we stress that this is only in line with the irregular natures of nonstandard problems CHT2007,AndrewsandSoares2010,Linton2010. While we are unable to offer a general theory guiding their choices, a two-step procedure similar to RomanoShaikhWolf2014TwoStep is proposed to mitigate potential concerns. The intuition is as follows. First, the appearance of $r_0\equiv \mathrm{rank}(\Pi_0)$ in the limits suggests the need of a consistent rank estimator $\hat r_n$, which may be achieved by a sequential testing procedure coupled with a significance level $\alpha_n$ (serving as the tuning parameter) that tends to zero suitably. Although the estimation error of $\hat r_n$, i.e., the probability of false selection, is asymptotically negligible (as $\alpha_n\to 0$), that probability is positive in any finite samples. Thus, we account for false selection by fixing $\alpha_n=\beta$ rather than letting it tend to zero. Given an estimator $\hat r_n$ with $\liminf_{n\to\infty}P(\hat r_n=r_0)\ge 1-\beta$, the two-step procedure at a significance level $\alpha$ is: reject $\mathrm H_0$ if $\hat r_n>r$ in the first step; otherwise in the second step incorporate $\hat r_n$ into our bootstrap and conduct the test at the adjusted significance level $\alpha-\beta>0$. We show in a number of simulation designs that the procedure is quite insensitive to our choices of $\beta$, even for small sample sizes.
The marked size and power properties rest with several attractive features. First, since we rely on the Delta method, the theory is conceptually simple and requires mild assumptions. Essentially, all we need are a matrix estimator $\hat\Pi_n$ that converges weakly and a consistent bootstrap analog. In particular, the data may be non-i.i.d.\ and non-stationary, the convergence rate may be non-$\sqrt n$ and even heterogeneous across entries of $\hat\Pi_n$ -- see Supplemental Appendix (ref), the limit $\mathcal M$ of $\hat\Pi_n$ may be non-Gaussian, the bootstrap for $\mathcal M$ (a crucial ingredient of our method) may be virtually any consistent resampling scheme, and no side rank conditions are directly imposed beyond those entailed by the restrictions on the population quantiles. Second, computation of our testing statistic and the critical values are quite simple as both involve only calculations of singular value decompositions -- we reiterate that the need of resampling only reflects the irregular natures of the problem rather than because of an exclusive attribute of our treatment. Finally, the superior testing properties of our procedure translate to more accurate rank estimators through the aforementioned two channels, namely, reducing type I and type II errors. Simulations confirm that our methods work better when $\mathrm{rank}(\Pi_0)<r$ or when $\mathrm{\Pi_0}$ is close to a matrix whose rank is strictly less than $r$.
We illustrate the application of our framework by testing identification in linear IV models that accommodates clustered data. To draw further attention to the empirical relevance of our results, we study a two-sided bipartite matching model with transferrable utility, building upon the work of DupuyGalichon2014Personality. A central question here is: how many attributes are relevant for the matching? Under a parametric specification of the surplus function, this number is equal to the rank of the so-called affinity matrix. We show that our procedure and Kleibergen_Paap2006rank can produce quite different results with regards to several model specifications, in terms of both $p$-values of the tests and actual estimates of the matching dimension.
As mentioned previously, the literature has been mostly concerned with the hypotheses (ref). In the context of multivariate regression, Anderson1951Estimating develops a likelihood ratio test based on canonical correlations. This test is restrictive in that it crucially depends on the asymptotic variance $\Omega_0$ of $\mathrm{vec}(\hat\Pi_n)$ having a Kronecker product structure. Building upon Gill_Lewbel1992rank, Cragg_Donald1996LDU propose a test that requires nonsingularity of $\Omega_0$ and may be sensitive to the transformations involved. Cragg_Donald1997infer provide a test based on a constrained minimum distance criterion, which, in addition to the nonsingularity requirement of $\Omega_0$, is in general computationally intensive. To relax the nonsingularity condition, Robin_Smith2000rank employ a class of testing statistics which are asymptotically equivalent to ours, but their results only apply to the setup (ref). Kleibergen_Paap2006rank study a Wald-standardized version of our statistic in order to obtain pivotal asymptotic distributions (under $\mathrm H_0'$), but at the expense of a side rank condition. We refer the reader to Camba_Kapetanios2009rank, Portier_Delyon2014 and Sadoon2017Rank for further discussions.
There are a few exceptions that study (ref). Johansen1988CoInt,Johansen1991CoIntegration obtains his likelihood ratio statistics under $ \mathrm H_0$ but only establishes their asymptotic distributions under $\mathrm H_0'$. Shortly after, Johansen1995likelihood presents the limits under $\mathrm H_0$, and essentially argues based on simulations that the asymptotic distributions under $\mathrm{rank}(\Pi_0)< r$ are first order stochastically dominated by those
under $\mathrm H_0'$ and “hence not relevant for calculating the $p$-value”. However, the counterexample given in Section (ref) disproves this conjecture. Cragg_Donald1993TestID recognize the importance of studying (ref), but do not derive the asymptotic distributions under $\mathrm{H}_0$. Instead, they show that their statistic has first order stochastically dominant limiting laws under $\mathrm H_0'$ with somewhat restrictive conditions. Our results suggest that may not be true in general.
We now introduce some notation. The space of $m\times k$ matrices is denoted by $\mathbf M^{m\times k}$. For a matrix $A$, we write its transpose by $A^\text{\scalebox{0.7}{$\intercal$}}$, its trace by $\mathrm{tr}(A)$ if it is square, its vectorization by $\text{vec}(A)$, and its Frobenius norm by $\|A\|\equiv\sqrt{\mathrm{tr}(A^\text{\scalebox{0.7}{$\intercal$}} A)}$. The identity matrix of size $k$ is denoted $I_k$, the $k\times 1$ vectors of zeros and ones are respectively denoted by $\mathbf 0_{k}$ and $\mathbf 1_k$, and the $m\times k$ matrix of zeros is denoted $\mathbf 0_{m\times k}$. We let $\mathrm{diag}(a)$ denote the diagonal matrix whose diagonal entries compose $a$. The $j$th largest singular value of a matrix $A\in\mathbf M^{m\times k}$ is denoted $\sigma_j(A)$. We define the set $\mathbb S^{m\times k}=\{A\in\mathbf M^{m\times k}: A^\text{\scalebox{0.7}{$\intercal$}} A=I_k\}$ and let $\overset{d}{=}$ signify “equal in distribution.” Finally, $\lfloor a\rfloor$ is the integer part of $a\in\mathbf R$.
The remainder of the paper is organized as follows. Section (ref) illustrates the consequences of ignoring $\mathrm{rank}(\Pi_0)<r$, and provides an overview of our tests, together with a step-by-step implementation guide. Section (ref) develops our inferential framework. Section (ref) presents Monte Carlo studies. Section (ref) further illustrates the empirical relevance of our results by studying a matching model. Section (ref) briefly concludes. Proofs are collected in a Supplemental Appendix. We also study the estimation problem, but, due to space limitation, relegate the results to Supplemental Appendix (ref). Finally, we have developed a Stata command bootranktest to test whether a matrix of the form $E[VZ^\text{\scalebox{0.7}{$\intercal$}}]$ has full rank -- see the Supplemental Appendix for a brief description.
\section{Motivations, Overview and Implementation}
In this section, we first motivate the development of our theory by illustrating how serious the issue can be if one ignores the possibility $\mathrm{rank}(\Pi_0)<r$ in conducting rank tests. This is accomplished by examining the influential test proposed by Kleibergen_Paap2006rank, referred to as the KP test hereafter, and its multiple testing version. Then we provide an overview of our tests, together with a step-by-step implementation guide that applies to general settings.
To elucidate the consequences of ignoring $\mathrm{rank}(\Pi_0)<r$, consider an example where $\Pi_0=\mathbf 0_{2\times 2}$ and $r=1$ so that $\mathrm{rank}(\Pi_0)<r$. Suppose $\Pi_0$ admits an estimator $\hat\Pi_n$ such that $\sqrt{n}\hat\Pi_n\overset{d}{=}\mathcal M$ for all $n$ (rather than just asymptotically), where $\mathcal M\in\mathbf M^{2\times 2}$ satisfies $\mathrm{vec}(\mathcal M)\sim N(0,\Omega_0)$ with $\Omega_0$ nonsingular and {\it known}. In this case, the KP test for (ref) employs critical values from $\chi^2(1)$, while the actual distribution of the KP statistic is
\begin{align}
T_{n,\mathrm{kp}}\overset{d}{=} \frac{\sigma_2^2(\mathcal M)}{(\mathcal{Q}_{2}\otimes \mathcal{P}_{2})^{\scalebox{0.7}{$\intercal$}}\Omega_0(\mathcal{Q}_{2}\otimes \mathcal{P}_{2})} ,
\end{align}
where $\mathcal P_2$ and $\mathcal Q_2$ are the left and right singular vectors associated with $\sigma_2(\mathcal M)$, both having unit length. Note the distribution of $T_{n,\mathrm{kp}}$ depends only on $\Omega_0$. Figure (ref) plots (based on simulations) two cdfs $F_1$ and $F_2$ of $T_{n,\mathrm{kp}}$ in (ref) respectively determined by
\begin{align}
\Omega_{1} = \left[
\begin{array}{rrrr}
1&0&0&0\\
0&1&0&0\\
0&0&1&0\\
0&0&0&1\\
\end{array}
\right] \text{ and }
\Omega_{2} = \left[
\begin{array}{rrrr}
1&0&0&-0.9\sqrt{5}\\
0&1&0.9\sqrt{5}&0\\
0&0.9\sqrt{5}&5&0\\
-0.9\sqrt{5}&0&0&5\\
\end{array}
\right] ,
\end{align}
together with the cdf $F_0$ of $\chi^2(1)$. Note that $F_0$ is stochastically dominated by $F_2$ but stochastically dominates $F_1$, both in the first order sense. Hence, the KP test is invalid due to over-rejection when $\Omega_0=\Omega_2$. We have thus {\it disproved} that the limits under $\mathrm{rank}(\Pi_0)=r$ are first order stochastically dominant in general, a conjecture by Cragg_Donald1993TestID for their statistic which they show to hold under somewhat restrictive conditions. These erratic behaviors can also be expected for the test of Robin_Smith2000rank in view of its relation to the KP test -- see Supplemental Appendix (ref).
\pgfplotstableread{
X Y1 Y2 Y3
0 3.73507364540145e-09 2.43226576729300e-09 0
0.0500000000000000 0.00909715271568395 0.00160932765346150 0.00393214000001952
0.100000000000000 0.0365618586579883 0.00635285996713346 0.0157907740934312
0.150000000000000 0.0842008779330303 0.0144885301338314 0.0357657791558976
0.200000000000000 0.151732607950905 0.0259754396467885 0.0641847546673016
0.250000000000000 0.237603608717244 0.0414901201290033 0.101531044267622
0.300000000000000 0.339919101427980 0.0607538991724877 0.148471861832545
0.350000000000000 0.459730566363366 0.0846617946454281 0.205900125227766
0.400000000000000 0.596126853011560 0.113759309645023 0.274995897728455
0.450000000000000 0.751643108364612 0.147976474823992 0.357317168286320
0.500000000000000 0.922756852624857 0.189495334531488 0.454936423119572
0.550000000000000 1.12039054578462 0.237734592980602 0.570651862051189
0.600000000000000 1.34938018483665 0.295545944496175 0.708326300800793
0.650000000000000 1.60197335354828 0.365623241253882 0.873457142989230
0.700000000000000 1.89960737697793 0.451494584493166 1.07419417085759
0.750000000000000 2.25620670728667 0.558068021483900 1.32330369693147
0.800000000000000 2.69908578532006 0.695736867924860 1.64237441514982
0.850000000000000 3.26558877040370 0.884018476842553 2.07225085582223
0.900000000000000 4.09641366876075 1.16204775101438 2.70554345409541
0.950000000000000 5.49430410654744 1.66751998558478 3.84145882069413
}\datainvalid
\begin{figure}[!htp]
\begin{tikzpicture}[scale=0.7]
\begin{axis}[scale = 1.6,
legend style={draw=none},
grid = minor,
xmax=6,xmin=0,
ymax=1,ymin=0,
xtick={0,1.67,3.84,5.49},
ytick={0,0.2,0.4,0.6,0.8,0.95,1},
width=0.65\textwidth,
height=6.8cm,
legend style={at={(0.9,0.28)}}
]
\addplot[smooth,tension=0.3,densely dashdotted, color=blue, line width=1.2pt] table[x = Y2,y=X] from \datainvalid ;
\addlegendentry{$F_1: \Omega_0 = \Omega_1$} ;
\addplot[smooth,tension=0.3, color=Ivory4, line width=1.2pt] table[x = Y3,y=X] from \datainvalid ;
\addlegendentry{$F_0: \chi^{2}(1)$} ;
\addplot[smooth,tension=0.3,loosely dashdotdotted, color=red, line width=1.2pt] table[x = Y1,y=X] from \datainvalid ;
\addlegendentry{$F_2: \Omega_0 = \Omega_2$} ;
\draw[dashed] (0,0.95)--(5.494,0.95);
\draw[dashed,blue] (1.667,0.95)--(1.667,0);
\draw[dashed] (3.841,0.95)--(3.841,0);
\draw[dashed,red] (5.494,0.95)--(5.494,0);
\end{axis}
\end{tikzpicture}
\caption{The cdfs of the KP statistic when $\Pi_0 = \mathbf{0}_{2\times 2}$ and $r=1$}
\end{figure}
Alternatively, one might aim to construct a valid test for (ref) by a multiple test on $\mathrm{rank}(\Pi_0)=0,1,\ldots,r$. However, the validity is achieved at the expense of conservativeness -- see Supplemental Appendix (ref), which may generate substantial power loss. To illustrate, consider the following data generating process:
\begin{align}
Z =\Pi_{0}^\text{\scalebox{0.7}{$\intercal$}} V+u ,
\end{align}
where $V,u\in N(0,I_{6})$ are independent and, for $\delta\ge 0$ and $d\in\{1,\ldots,6\}$,
\begin{align}
\Pi_{0}=\textrm{diag}(\mathbf{1}_{6-d},\mathbf{0}_{d})+\delta I_{6} .
\end{align}
We test the hypotheses in (ref) with $r=5$ at the level $\alpha=5\%$, and note that $\mathrm{H}_{0}$ holds if and only if $\delta=0$. For an i.i.d.\ sample $\{V_i,Z_i\}_{i=1}^{1000}$ generated according to (ref), we conduct tests based on the matrix estimator $\hat{\Pi}_{n}=\frac{1}{1000}\sum_{i=1}^{1000}V_{i}Z_{i}^{\text{\scalebox{0.7}{$\intercal$}}}$ for $\Pi_{0}$.
\pgfplotstableread{
delta alpha d1 d2 d3 d4 d5 d6
0 0.0500000000000000 0.0492000000000000 0.00410000000000000 0.000300000000000000 0 0 0
0.0200000000000000 0.0500000000000000 0.0984000000000000 0.00840000000000000 0.000500000000000000 0 0 0
0.0400000000000000 0.0500000000000000 0.246600000000000 0.0444000000000000 0.00560000000000000 0.000400000000000000 0.000100000000000000 0
0.0600000000000000 0.0500000000000000 0.471500000000000 0.158600000000000 0.0414000000000000 0.00740000000000000 0.00140000000000000 0.000200000000000000
0.0800000000000000 0.0500000000000000 0.717800000000000 0.399000000000000 0.179600000000000 0.0674000000000000 0.0215000000000000 0.00610000000000000
0.100000000000000 0.0500000000000000 0.887300000000000 0.678300000000000 0.444100000000000 0.257000000000000 0.126200000000000 0.0538000000000000
0.120000000000000 0.0500000000000000 0.963700000000000 0.876300000000000 0.732000000000000 0.567300000000000 0.396100000000000 0.243700000000000
0.140000000000000 0.0500000000000000 0.992000000000000 0.966400000000000 0.915300000000000 0.831000000000000 0.715300000000000 0.570400000000000
0.160000000000000 0.0500000000000000 0.998800000000000 0.995500000000000 0.983300000000000 0.958200000000000 0.917600000000000 0.843900000000000
0.180000000000000 0.0500000000000000 0.999800000000000 0.999500000000000 0.998300000000000 0.993300000000000 0.982400000000000 0.964400000000000
0.200000000000000 0.0500000000000000 1 1 1 0.999600000000000 0.998000000000000 0.994700000000000
}\motivation
\pgfplotstableread{
delta alpha d1 d2 d3 d4 d5 d6
0 0.0500000000000000 0.0502000000000000 0.0536000000000000 0.0516000000000000 0.0535000000000000 0.0529000000000000 0.0541000000000000
0.0200000000000000 0.0500000000000000 0.100900000000000 0.0832000000000000 0.0739000000000000 0.0668000000000000 0.0621000000000000 0.0636000000000000
0.0400000000000000 0.0500000000000000 0.248300000000000 0.197300000000000 0.150900000000000 0.124900000000000 0.110700000000000 0.0985000000000000
0.0600000000000000 0.0500000000000000 0.472100000000000 0.434500000000000 0.345800000000000 0.278500000000000 0.222000000000000 0.184400000000000
0.0800000000000000 0.0500000000000000 0.716400000000000 0.701300000000000 0.620600000000000 0.545100000000000 0.454300000000000 0.370200000000000
0.100000000000000 0.0500000000000000 0.883900000000000 0.885800000000000 0.848700000000000 0.795600000000000 0.723300000000000 0.632300000000000
0.120000000000000 0.0500000000000000 0.963700000000000 0.963200000000000 0.952700000000000 0.932000000000000 0.898300000000000 0.851100000000000
0.140000000000000 0.0500000000000000 0.991900000000000 0.990100000000000 0.988600000000000 0.982400000000000 0.970500000000000 0.956200000000000
0.160000000000000 0.0500000000000000 0.998600000000000 0.997400000000000 0.997300000000000 0.995700000000000 0.993900000000000 0.990700000000000
0.180000000000000 0.0500000000000000 0.999800000000000 0.999700000000000 0.999500000000000 0.999300000000000 0.998700000000000 0.998100000000000
0.200000000000000 0.0500000000000000 1 1 1 0.999800000000000 0.999800000000000 0.999500000000000
}\dominationcf
{
\begin{figure}[!h]
\begin{subfigure}[b]{0.48\textwidth}
\resizebox{\linewidth}{!}{
\begin{tikzpicture}
\begin{axis}[
legend style={draw=none},
grid = minor,
xmax=0.2,xmin=0,
ymax=1,ymin=0,
xtick={0,0.1,0.2},
ytick={0.5,1},
tick label style={/pgf/number format/fixed},
legend style={at={(0.82,0.47)},anchor=north,
row sep = 3pt}]
\addlegendimage{empty legend};
\addlegendentry{ $d=1$}
\addplot[smooth,tension=0.5,color=SeaGreen4, line width=0.75pt] table[x = delta,y=d1] from \dominationcf ;
\addlegendentry{ CF-A} ;
\addplot[smooth,tension=0.5,color=black, line width=0.75pt, dashdotted] table[x = delta,y=d1] from \motivation ;
\addlegendentry{ KP-M} ;
\addplot[smooth,tension=0.5,no markers, color=NavyBlue, line width=0.75pt, densely dotted] table[x = delta,y=alpha] from \motivation ;
\addlegendentry{ $5\%$ level} ;
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.48\textwidth}
\resizebox{\linewidth}{!}{
\begin{tikzpicture}
\begin{axis}[
legend style={draw=none},
grid = minor,
xmax=0.2,xmin=0,
ymax=1,ymin=0,
xtick={0,0.1,0.2},
ytick={0.5,1},
tick label style={/pgf/number format/fixed},
legend style={at={(0.82,0.47)},anchor=north,
row sep = 3pt}]
\addlegendimage{empty legend};
\addlegendentry{ $d=2$}
\addplot[smooth,tension=0.5,color=SeaGreen4, line width=0.75pt] table[x = delta,y=d2] from \dominationcf ;
\addlegendentry{ CF-A} ;
\addplot[smooth,tension=0.5,color=black, line width=0.75pt,dashdotted] table[x = delta,y=d2] from \motivation ;
\addlegendentry{ KP-M} ;
\addplot[smooth,tension=0.5,no markers, color=NavyBlue, line width=0.75pt, densely dotted] table[x = delta,y=alpha] from \motivation ;
\addlegendentry{ $5\%$ level} ;
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.48\textwidth}
\resizebox{\linewidth}{!}{
\begin{tikzpicture}
\begin{axis}[
legend style={draw=none},
grid = minor,
xmax=0.2,xmin=0,
ymax=1,ymin=0,
xtick={0,0.1,0.2},
ytick={0.5,1},
tick label style={/pgf/number format/fixed},
legend style={at={(0.82,0.47)},anchor=north,
row sep = 3pt}]
\addlegendimage{empty legend};
\addlegendentry{ $d=3$} ;
\addplot[smooth,tension=0.5,color=SeaGreen4, line width=0.75pt] table[x = delta,y=d3] from \dominationcf ;
\addlegendentry{ CF-A} ;
\addplot[smooth,tension=0.5,color=black, line width=0.75pt, dashdotted] table[x = delta,y=d3] from \motivation ;
\addlegendentry{ KP-M} ;
\addplot[smooth,tension=0.5,no markers, color=NavyBlue, line width=0.75pt, densely dotted] table[x = delta,y=alpha] from \motivation ;
\addlegendentry{ $5\%$ level} ;
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.48\textwidth}
\resizebox{\linewidth}{!}{
\begin{tikzpicture}
\begin{axis}[
legend style={draw=none},
grid = minor,
xmax=0.2,xmin=0,
ymax=1,ymin=0,
xtick={0,0.1,0.2},
ytick={0.5,1},
tick label style={/pgf/number format/fixed},
legend style={at={(0.82,0.47)},anchor=north,
row sep = 3pt}]
\addlegendimage{empty legend};
\addlegendentry{ $d=4$} ;
\addplot[smooth,tension=0.5,color=SeaGreen4, line width=0.75pt] table[x = delta,y=d4] from \dominationcf ;
\addlegendentry{ CF-A} ;
\addplot[smooth,tension=0.5,color=black, line width=0.75pt,dashdotted] table[x = delta,y=d4] from \motivation ;
\addlegendentry{ KP-M} ;
\addplot[smooth,tension=0.5,no markers, color=NavyBlue, line width=0.75pt, densely dotted] table[x = delta,y=alpha] from \motivation ;
\addlegendentry{ $5\%$ level} ;
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.48\textwidth}
\resizebox{\linewidth}{!}{
\begin{tikzpicture}
\begin{axis}[
legend style={draw=none},
grid = minor,
xmax=0.2,xmin=0,
ymax=1,ymin=0,
xtick={0,0.1,0.2},
ytick={0.5,1},
tick label style={/pgf/number format/fixed},
legend style={at={(0.82,0.47)},anchor=north,
row sep = 3pt}]
\addlegendimage{empty legend};
\addlegendentry{ $d=5$}
\addplot[smooth,tension=0.5,color=SeaGreen4, line width=0.75pt] table[x = delta,y=d5] from \dominationcf ;
\addlegendentry{ CF-A} ;
\addplot[smooth,tension=0.5,color=black, line width=0.75pt, dashdotted] table[x = delta,y=d5] from \motivation ;
\addlegendentry{ KP-M} ;
\addplot[smooth,tension=0.5,no markers, color=NavyBlue, line width=0.75pt, densely dotted] table[x = delta,y=alpha] from \motivation ;
\addlegendentry{ $5\%$ level} ;
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.48\textwidth}
\resizebox{\linewidth}{!}{
\begin{tikzpicture}
\begin{axis}[
legend style={draw=none},
grid = minor,
xmax=0.2,xmin=0,
ymax=1,ymin=0,
xtick={0,0.1,0.2},
ytick={0.5,1},
tick label style={/pgf/number format/fixed},
legend style={at={(0.82,0.47)},anchor=north,
row sep = 3pt}]
\addlegendimage{empty legend};
\addlegendentry{ $d=6$}
\addplot[smooth,tension=0.5,color=SeaGreen4, line width=0.75pt] table[x = delta,y=d6] from \dominationcf ;
\addlegendentry{ CF-A} ;
\addplot[smooth,tension=0.5,color=black, line width=0.75pt,dashdotted] table[x = delta,y=d6] from \motivation ;
\addlegendentry{ KP-M} ;
\addplot[smooth,tension=0.5,no markers, color=NavyBlue, line width=0.75pt, densely dotted] table[x = delta,y=alpha] from \motivation ;
\addlegendentry{ $5\%$ level} ;
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\caption{Conservativeness of the KP-M test. The number of Monte Carlo simulations is 10,000, the number of bootstrap repetitions is 500, and $\kappa_n=n^{-1/4}$ (for CF-A).}
\end{figure}
}
\pgfplotstableread{
delta alpha kp kpm cft cfa cfn
0 0.0500000000000000 0.00500000000000000 0.00460000000000000 0.0444000000000000 0.0514000000000000 0.0482000000000000
0.0200000000000000 0.0500000000000000 0.00990000000000000 0.00900000000000000 0.0629000000000000 0.0733000000000000 0.0693000000000000
0.0400000000000000 0.0500000000000000 0.0386000000000000 0.0374000000000000 0.160300000000000 0.187200000000000 0.180600000000000
0.0600000000000000 0.0500000000000000 0.154900000000000 0.153000000000000 0.338200000000000 0.430900000000000 0.411500000000000
0.0800000000000000 0.0500000000000000 0.392300000000000 0.390500000000000 0.540400000000000 0.703500000000000 0.685500000000000
0.100000000000000 0.0500000000000000 0.678900000000000 0.678500000000000 0.734800000000000 0.890200000000000 0.883000000000000
0.120000000000000 0.0500000000000000 0.885000000000000 0.885000000000000 0.890300000000000 0.967900000000000 0.969000000000000
0.140000000000000 0.0500000000000000 0.972600000000000 0.972600000000000 0.969200000000000 0.991100000000000 0.994400000000000
0.160000000000000 0.0500000000000000 0.995800000000000 0.995800000000000 0.994800000000000 0.997800000000000 0.999200000000000
0.180000000000000 0.0500000000000000 0.999300000000000 0.999300000000000 0.999200000000000 0.999600000000000 0.999900000000000
0.200000000000000 0.0500000000000000 0.999900000000000 0.999900000000000 0.999900000000000 0.999900000000000 1
}\motivationa
\pgfplotstableread{
delta alpha kp kpm cft cfa cfn
0 0.0500000000000000 0.115100000000000 0.0290000000000000 0.0469000000000000 0.0501000000000000 0.0420000000000000
0.0400000000000000 0.0500000000000000 0.239800000000000 0.119700000000000 0.157700000000000 0.169600000000000 0.150100000000000
0.0800000000000000 0.0500000000000000 0.446800000000000 0.402200000000000 0.416900000000000 0.483200000000000 0.455800000000000
0.1200000000000000 0.0500000000000000 0.602000000000000 0.600600000000000 0.594300000000000 0.756700000000000 0.730700000000000
0.1600000000000000 0.0500000000000000 0.744600000000000 0.744600000000000 0.734100000000000 0.908100000000000 0.889600000000000
0.200000000000000 0.0500000000000000 0.873700000000000 0.873700000000000 0.866600000000000 0.953300000000000 0.962300000000000
0.240000000000000 0.0500000000000000 0.952400000000000 0.952400000000000 0.947800000000000 0.964200000000000 0.989100000000000
0.280000000000000 0.0500000000000000 0.986600000000000 0.986600000000000 0.984600000000000 0.987100000000000 0.997600000000000
0.320000000000000 0.0500000000000000 0.997200000000000 0.997200000000000 0.996800000000000 0.997100000000000 0.999700000000000
0.360000000000000 0.0500000000000000 0.999500000000000 0.999500000000000 0.999500000000000 0.999500000000000 0.999900000000000
0.400000000000000 0.0500000000000000 0.999900000000000 0.999900000000000 0.999900000000000 0.999900000000000 1
}\motivationb
{
\begin{figure}[!h]
\begin{subfigure}[b]{0.48\textwidth}
\resizebox{\linewidth}{!}{
\begin{tikzpicture}
\begin{axis}[
legend style={draw=none},
grid = minor,
xmax=0.2,xmin=0,
ymax=1,ymin=0,
xtick={0,0.1,0.2},
ytick={0.5,1},
tick label style={/pgf/number format/fixed},
legend style={at={(0.98,0.6)}}]
\addlegendimage{empty legend};
\addlegendentry{ $\Omega_0=\Omega_1$}
\addplot[smooth,tension=0.5,mark=*, color=blue, line width=0.75pt] table[x = delta,y=cfa] from \motivationa ;
\addlegendentry{ CF-A} ;
\addplot[smooth,tension=0.5,mark=pentagon, color=green, line width=0.75pt] table[x = delta,y=cfn] from \motivationa ;
\addlegendentry{ CF-N} ;
\addplot[smooth,tension=0.5,mark=triangle, color={rgb:red,4;green,2;yellow,1}, line width=0.75pt] table[x = delta,y=cft] from \motivationa ;
\addlegendentry{ CF-T} ;
\addplot[smooth,tension=0.5,mark =square , color=magenta, line width=0.75pt] table[x = delta,y=kp] from \motivationa ;
\addlegendentry{ KP} ;
\addplot[smooth,tension=0.5,mark=star, color=red, line width=0.75pt] table[x = delta,y=kpm] from \motivationa ;
\addlegendentry{ KP-M} ;
\addplot[smooth,tension=0.5,no markers, color=NavyBlue, line width=0.75pt, densely dotted] table[x = delta,y=alpha] from \motivationa ;
\addlegendentry{ $5\%$ level} ;
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.48\textwidth}
\resizebox{\linewidth}{!}{
\begin{tikzpicture}
\begin{axis}[
legend style={draw=none},
grid = minor,
xmax=0.4,xmin=0,
ymax=1,ymin=0,
xtick={0,0.2,0.4},
ytick={0.5,1},
tick label style={/pgf/number format/fixed},
legend style={at={(0.98,0.6)}}]
\addlegendimage{empty legend};
\addlegendentry{ $\Omega_0=\Omega_2$}
\addplot[smooth,tension=0.5,mark=*, color=blue, line width=0.75pt] table[x = delta,y=cfa] from \motivationb ;
\addlegendentry{ CF-A} ;
\addplot[smooth,tension=0.5,mark=pentagon, color=green, line width=0.75pt] table[x = delta,y=cfn] from \motivationb ;
\addlegendentry{ CF-N} ;
\addplot[smooth,tension=0.5,mark=triangle, color={rgb:red,4;green,2;yellow,1}, line width=0.75pt] table[x = delta,y=cft] from \motivationb ;
\addlegendentry{ CF-T} ;
\addplot[smooth,tension=0.5,mark =square , color=magenta, line width=0.75pt] table[x = delta,y=kp] from \motivationb ;
\addlegendentry{ KP} ;
\addplot[smooth,tension=0.5,mark=star, color=red, line width=0.75pt] table[x = delta,y=kpm] from \motivationb ;
\addlegendentry{ KP-M} ;
\addplot[smooth,tension=0.5,no markers, color=NavyBlue, line width=0.75pt, densely dotted] table[x = delta,y=alpha] from \motivationb ;
\addlegendentry{ $5\%$ level} ;
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\caption{Comparisons with the KP and the KP-M tests. The number of Monte Carlo simulations is 10,000, the number of bootstrap repetitions is 1000, $\kappa_n=n^{-1/4}$ (for both CF-A and CF-N), and $\beta=\alpha/10$ (for CF-T).}
\end{figure}
}
Figure (ref) plots the power functions (against $\delta$) of the multiple KP test, labelled KP-M. For $d=1$ (and so $\mathrm{rank}(\Pi_0)=r$), the null rejection rate is 5%, while the power increases to unity as $\delta$ increases. As soon as $d>1$ (so that $\mathrm{rank}(\Pi_0)<r$), the power curves shift downward dramatically: the null rejection rates are close to zero and the power is well below 5% when $\delta$ is close to zero. Moreover, the power deteriorates as $\Pi_0$ becomes more degenerate in the sense that $\Pi_0$ is close to a matrix whose rank becomes smaller as $d$ increases. This reinforces the critical importance to accommodate $\mathrm{rank}(\Pi_0)<r$.
To compare, we first show that three versions of our test -- CF-A, CF-N and CF-T (see below) -- control size even when the KP test does not. Let $\{Z_i\}_{i=1}^{1000}$ be an i.i.d.\ sample in $\mathbf M^{2\times 2}$ such that $\mathrm{vec}(Z_1)\sim N(\mathrm{vec}(\Pi_0), \Omega_0)$, where $\mathrm{vec}(\Pi_0) = \delta \Omega_0^{1/2} \mathrm{vec}(I_{2})$ with $\delta\geq 0$ and $\Omega_0\in\{\Omega_1,\Omega_2\}$ as in (ref). We test (ref) with $r = 1$ based on $\hat{\Pi}_{n} = \frac{1}{1000}\sum_{i=1}^{1000}Z_{i}$, at $\alpha = 5\%$. Figure (ref) shows our tests indeed control size for both choices of $\Omega_0$, while the KP test under-rejects when $\Omega_0=\Omega_1$ and over-rejects when $\Omega_0=\Omega_2$. Note also that the KP-M test is conservative. Next, for the designs in (ref) and (ref), Figure (ref) depicts the power curves of CF-A. For $d=1$, CF-A and KP-M have virtually the same rejection rates across $\delta$. Whenever $d>1$, our test effectively raises the power curves of the KP-M test so that the null rejection rates equal $5\%$, and the power becomes nontrivial. But it is more than that. The power improvement increases when $d$ gets larger.
To describe our test, let $\hat\Pi_n$ be an estimator of $\Pi_0\in\mathbf M^{m\times k}$ with $\tau_n\{\hat\Pi_n-\Pi_0\}\xrightarrow{L} \mathcal M$. The exact characterization of $\mathcal M$ (e.g., the covariance structure) is not required. Here, $\tau_n$ is typically $\sqrt n$ in cross-sectional and stationary time series settings, and may be non-$\sqrt n$ with non-stationary time series. Then our test statistic for (ref) is $\tau_n^2\phi_r(\hat\Pi_n)\equiv \tau_n^2\sum_{j=r+1}^{k}\sigma_j^2(\hat\Pi_n)$. It turns out that, under $\mathrm{H}_0$, we have: for $r_0\equiv\mathrm{rank}(\Pi_0)$,
\begin{align}
\tau_{n}^{2}\phi_r(\hat{\Pi}_{n})\overset{L}{\rightarrow} \sum_{j=r-r_{0}+1}^{k-r_{0}}\sigma^{2}_j(P_{0,2}^\text{\scalebox{0.7}{$\intercal$}} \mathcal{M} Q_{0,2}) ,
\end{align}
where $P_{0,2}\in\mathbb S^{m\times(m-r_0)}$ and $Q_{0,2}\in\mathbb S^{k\times (k-r_0)}$ whose columns are respectively the left and the right singular vectors of $\Pi_{0}$ associated with its zero singular values. Since the limit in (ref) depends on the true rank $r_0$ (crucially), $P_{0,2}$, $Q_{0,2}$ and $\mathcal M$, we estimate its law by first estimating these unknown objects, towards constructing critical values.
The rank $r_0$ may be consistently (under $\mathrm H_0$) estimated by: for $\kappa_n\to 0$ and $\tau_n\kappa_n\to\infty$,
\begin{align}
\hat r_n=\max\{j=1,\ldots,r: \sigma_j(\hat\Pi_n)\ge \kappa_n\}
\end{align}
if the set is nonempty and $\hat r_n=0$ otherwise. Heuristically, $\kappa_n$ may be thought of as testing which population singular values are zero. Note that by estimating $r_0$ we take into account the possibility $r_0<r$. Next, for a singular value decomposition $\hat\Pi_n=\hat P_n\hat\Sigma_n\hat Q_n^\text{\scalebox{0.7}{$\intercal$}}$, we may respectively estimate $P_{0,2}$ and $Q_{0,2}$ by $\hat{P}_{2,n}$ and $\hat{Q}_{2,n}$, which are respectively formed by the last $(m-\hat r_n)$ and $(k-\hat r_n)$ columns of $\hat P_n$ and $\hat Q_n$. The law of $\mathcal M$ may be consistently estimated by a bootstrap, say, $\hat{\mathcal M}_n^*$. Often, $\hat{\mathcal M}_n^*=\sqrt n\{\hat\Pi_n^*-\hat\Pi_n\}$ with $\hat\Pi_n^*$ computed in the same way as $\hat\Pi_n$ but based on a bootstrap sample. Finally, the law of the limit in (ref) is estimated by the conditional distribution (given the data) of
\begin{align}
\sum_{j=r-\hat r_n+1}^{k-\hat{r}_{n}}\sigma^{2}_j(\hat{P}_{2,n}^\text{\scalebox{0.7}{$\intercal$}} \hat{\mathcal M}_n^* \hat{Q}_{2,n}) .
\end{align}
Given a significance level $\alpha$, the CF-A test rejects $\mathrm H_0$ whenever $\tau_n^2\phi_r(\hat\Pi_n)>\hat c_{n,1-\alpha}$, where $\hat c_{n,1-\alpha}$ is the $1-\alpha$ conditional quantile of (ref) given the data.
While we are unable to provide an optimal choice of $\kappa_n$, a two-step test, CF-T, is proposed to mitigate potential concerns. In the first step, we obtain an estimator $\hat r_n$ satisfying $\liminf_{n\to\infty}P(\hat r_n=r_0)\ge 1-\beta$ for some $\beta<\alpha$, and then reject $\mathrm{H}_0$ if $\hat r_n>r$ and move on to the next step if $\hat r_n\le r$. In the second step, we plug $\hat r_n$ into (ref) and reject $\mathrm{H}_0$ if $\tau_n^2\phi_r(\hat\Pi_n)>\hat c_{n,1-\alpha+\beta}$, where the significance level is adjusted to be $\alpha-\beta$. The estimator $\hat r_n$ in (ref) now may not be appropriate as it appears challenging to control $P(\hat r_n=r_0)$. Instead, a desired estimator $\hat r_n$ may be obtained by a sequential testing procedure as actually employed in the literature and formalized in Supplemental Appendix (ref). In this regard, we stress that the KP test may be utilized and is recommended as it is tuning parameter free and does not require additional simulations.
Below we provide an implementation guide for testing (ref) at significance level $\alpha$.
{
\begin{tcolorbox}[enhanced,breakable,boxrule=0.7pt,width=0.95\textwidth]
\underline{\sc Step 1:} Compute a singular value decomposition $\hat\Pi_n=\hat P_n\hat\Sigma_n\hat Q_n^\text{\scalebox{0.7}{$\intercal$}}$.
\underline{\sc Step 2:} Obtain $\hat r_n$ as in (ref) for a chosen $\kappa_n$ (e.g.\ $\kappa_n=n^{-1/4}$).
\underline{\sc Step 3:} Bootstrap $B$ times and compute copies of $\hat{\mathcal M}_n^*$, denoted $\{\hat{\mathcal M}_{n,b}^*\}_{b=1}^B$.
\underline{\sc Step 4:} For $\hat P_{2,n}$ and $\hat Q_{2,n}$ formed by the last $(m-\hat r_n)$ and $(k-\hat r_n)$ columns of $\hat P_n$ and $\hat Q_n$ respectively, set $\hat c_{n,1-\alpha}$ to be the $\lfloor B(1-\alpha)\rfloor$-th largest value in
\begin{align}
\sum_{j=r-\hat r_n+1}^{k-\hat{r}_{n}}\sigma^{2}_j(\hat{P}_{2,n}^\text{\scalebox{0.7}{$\intercal$}} \hat{\mathcal M}_{n,1}^* \hat{Q}_{2,n}) ,\,\ldots ,\sum_{j=r-\hat r_n+1}^{k-\hat{r}_{n}}\sigma^{2}_j(\hat{P}_{2,n}^\text{\scalebox{0.7}{$\intercal$}} \hat{\mathcal M}_{n,B}^* \hat{Q}_{2,n}) .
\end{align}
\underline{\sc Step 5:} Reject $\mathrm{H}_0$ if $\tau_n^2\sum_{j=r+1}^{k}\sigma_j^2(\hat\Pi_n)>\hat c_{n,1-\alpha}$.
\end{tcolorbox}
}
Compared to CF-N which is based on the numerical differentiation Hong_Li2017numerical (see Sections (ref) and (ref) for more details), CF-A is somewhat insensitive to the choice of $\kappa_n$ even in small samples. The two-step test CF-T, on the other hand, is overall the least sensitive, but may be over-sized in small samples $(n\le 100)$. Thus, for practical purpose, we recommend the latter when the sample size is reasonably large. To implement it, one replaces {\sc Steps 2} and 5 with
{
\begin{tcolorbox}[enhanced,breakable,boxrule=0.7pt,width=0.95\textwidth]
\underline{\sc Step 2':} Obtain $\hat r_n$ by sequentially testing $\mathrm{rank}(\Pi_0)=0,1,\ldots,k-1$ at level $\beta$ (e.g., $\beta=\alpha/10$) using the KP test (based on $\hat\Pi_n$), i.e., $\hat r_n=j^*$ if accepting $\mathrm{rank}(\Pi_0)=j^*$ is the first acceptance in the procedure, and $\hat r_n=k$ if all nulls are rejected. Reject $\mathrm H_0$ if $\hat r_n>r$ and move on to Step 3 otherwise.
\underline{\sc Step 5':} Reject $\mathrm{H}_0$ if $\tau_n^2\sum_{j=r+1}^{k}\sigma_j^2(\hat\Pi_n)>\hat c_{n,1-\alpha+\beta}$.
\end{tcolorbox}
}
\section{The Inferential Framework}
In this section, we develop our inferential framework in three steps. First, we derive the differential properties of the map $\phi_r$ given in (ref), which is nontrivial and the key to our theory. Second, given an estimator $\hat\Pi_n$ of $\Pi_0$, we derive the asymptotic distributions for the plug-in estimator $\phi_r(\hat\Pi_n)$ by invoking the Delta method. These limits turn out to be highly nonstandard whenever $\mathrm{rank}(\Pi_0)<r$. Thus, in the third step, we construct valid and powerful rank tests by appealing to recent advances on bootstrap in irregular problems Fang_Santos2014HDD,Chen_Fang2015FOD,Hong_Li2017numerical. A two-step test is proposed to mitigate potential concerns on sensitivity of our tests to the choices of tuning parameters. Local properties of our tests will also be discussed.
\subsection{Differential Properties}
Let $\Pi_0\in\mathbf{M}^{m\times k}$ be an unknown matrix with $m\ge k$ and $\sigma_1(\Pi_0)\ge\cdots\ge\sigma_k(\Pi_0)\ge 0$ be singular values of $\Pi_0$. Then the rank of $\Pi_0$ is equal to the number of nonzero singular values of $\Pi_0$ -- see, for example, Bhatia1997Matrix and also Supplemental Appendix for a brief review. Hence, the hypotheses in (ref) are equivalent to
\begin{align}
\mathrm{H}_{0}: \phi_r(\Pi_0)=0 \qquad \text{v.s.} \qquad \mathrm H_1: \phi_r(\Pi_0)>0 ,
\end{align}
where $\phi_r: \mathbf M^{m\times k}\to\mathbf R$ is given by
\begin{align}
\phi_r(\Pi)\equiv\sum_{j=r+1}^k\sigma^{2}_j(\Pi) .
\end{align}
Heuristically, $\phi_r(\Pi)$ simply gives us the sum of the $k-r$ smallest squared singular values of $\Pi$. One may also consider other $L_p$-type functionals such as $\sum_{j=r+1}^k\sigma_j(\Pi)$. Our current focus, however, allows us to uncover $\chi^2$-type limiting distributions when $\mathrm{rank}(\Pi_0)=r$ and in this way facilitates comparisons with existing rank tests.
Towards deriving the asymptotic distributions of the plug-in estimator $\phi_r(\hat\Pi_n)$ for a given estimator $\hat\Pi_n$ of $\Pi_0$, we need to first establish suitable differentiability for the map $\phi_r$. The following lemma shall prove useful in this regard.
\begin{lem}
For the map $\phi_r$ in (ref), we have:
\begin{align}
\phi_r(\Pi)=\min_{U\in\mathbb S^{k\times(k-r)}}\|\Pi U\|^2 .
\end{align}
\end{lem}
Lemma (ref) shows that $\phi_r(\Pi)$ can be represented as the minimum of a quadratic form over the space of orthonormal matrices in $\mathbf M^{m\times (k-r)}$. The special case when $r=k-1$ (corresponding to the test of $\Pi$ having full rank) is a well known implication of the classical Courant-Fischer theorem, i.e., $\sigma^{2}_k(\Pi)=\min_{\|U\|=1}\|\Pi U\|^2$. Note that the minimum in (ref) is attained and hence well defined. It turns out that $\phi_r$ is not fully differentiable in general but belongs to a class of directionally differentiable maps. For completeness, we next introduce the relevant notions of directional differentiability.
\begin{defn}
Let $\phi:\mathbf{M}^{m\times k}\to\mathbf R$ be a generic function.
\begin{itemize}
• The map $\phi$ is said to be {\it Hadamard directionally differentiable} at $\Pi \in\mathbf{M}^{m\times k}$ if there is a map $\phi_\Pi':\mathbf{M}^{m\times k}\to\mathbf R$ such that:
\begin{equation}
\lim_{n\rightarrow \infty}\frac{\phi(\Pi +t_n M_n)-\phi(\Pi)}{t_n} = \phi_\Pi'(M) ,
\end{equation}
whenever $M_n\to M$ in $\mathbf{M}^{m\times k}$ and $t_n\downarrow 0$ for $\{t_n\}$ all strictly positive.
• If $\phi:\mathbf{M}^{m\times k}\to\mathbf R$ is Hadamard directionally differentiable at $\Pi \in\mathbf{M}^{m\times k}$, then we say that $\phi$ is {\it second order Hadamard directionally differentiable} at $\Pi\in\mathbf{M}^{m\times k}$ if there is a map $\phi_\Pi'':\mathbf{M}^{m\times k} \to\mathbf R$ such that:
\begin{align}
\lim_{n\rightarrow \infty}\frac{\phi(\Pi +t_n M_n)-\phi(\Pi)-t_n\phi_\Pi'(M_n)}{t_n^2} = \phi_\Pi”(M) ,
\end{align}
whenever $M_n\to M$ in $\mathbf{M}^{m\times k}$ and $t_n\downarrow 0$ for $\{t_n\}$ all strictly positive.
\end{itemize}
\end{defn}
For simplicity, we shall drop the qualifier “Hadamard” in what follows, with the understanding that both full differentiability and directional differentiability (both first and second order) are meant in the Hadamard sense. Definition (ref)(i) generalizes (full) differentiability which additionally requires the derivative $\phi_\Pi'$ to be linear. By Proposition 2.1 in Fang_Santos2014HDD, linearity is precisely the gap between these two notions of differentiability -- see also Shapiro1990 for more discussions. Despite the relaxation, the Delta method remains valid even when $\phi$ is only directionally differentiable Shapiro1991,Dumbgen1993. Unfortunately, as shall be proved, the asymptotic distributions of our statistic $\phi(\hat\Pi_n)$ implied by this generalized Delta method are degenerate under the null. In turn, Definition (ref)(ii) formulates a suitable second order analog of the directional differentiability, which permits us to obtain nondegenerate asymptotic distributions by a (generalized) second order Delta method Shapiro2000inference,Chen_Fang2015FOD. The second order directional differentiability becomes second order full differentiability precisely when $\phi_\Pi''$ corresponds to a bilinear form.
The following proposition formally establishes the differentiability of $\phi_r$.
\begin{pro}
Let $\phi_r: \mathbf M^{m\times k}\to\mathbf R$ be defined as in (ref).
\begin{itemize}
• $\phi_r$ is first order directionally differentiable at any $\Pi\in \mathbf M^{m\times k}$ with the derivative $\phi_{r,\Pi}': \mathbf M^{m\times k}\to\mathbf R$ given by
\begin{align}
\phi_{r,\Pi}'(M)=\min_{U\in\Psi(\Pi)} 2\mathrm{tr}\big((\Pi U)^\text{\scalebox{0.7}{$\intercal$}} MU\big) ,
\end{align}
where $\Psi(\Pi)\equiv\operatornamewithlimits{arg\,min}_{U\in\mathbb S^{k\times (k-r)}}\|\Pi U\|^{2}$.
• $\phi_r$ is second order directionally differentiable at any $\Pi\in \mathbf M^{m\times k}$ satisfying $\phi_r(\Pi)=0$ with the derivative
$\phi_{r,\Pi}^{\prime\prime}: \mathbf M^{m\times k}\to\mathbf R$ given by: for $r_0\equiv\mathrm{rank}(\Pi)$,
\begin{align}
\phi_{r,\Pi}^{\prime\prime}(M) = \sum_{j=r-r_{0}+1}^{k-r_{0}}\sigma^{2}_j(P_2^\text{\scalebox{0.7}{$\intercal$}} MQ_2) ,
\end{align}
where the columns of $P_2\in\mathbb S^{m\times (m-r_0)}$ and $Q_2\in \mathbb S^{k\times (k-r_0)}$ are left and right singular vectors associated with the zero singular values of $\Pi$.
\end{itemize}
\end{pro}
Proposition (ref)(i) shows that $\phi_r$ is not fully differentiable in general but only directionally differentiable. Moreover, the first order derivative is degenerate at zero whenever $\phi_r(\Pi)=0$ as in this case $\Pi U=0$ for any $U\in\Psi(\Pi)$. Proposition (ref)(ii) indicates that $\phi_r$ is second order directionally differentiable whenever the degeneracy occurs, and, interestingly, the derivative evaluated at $M$ is simply the sum of the $k-r$ smallest squared singular values of the $(m-r_0)\times (k-r_0)$ matrix $P_2^\text{\scalebox{0.7}{$\intercal$}} MQ_2$. In general, $\phi_r$ is not second order fully differentiable precisely when $\mathrm{rank}(\Pi)<r$, reflecting a critical irregular nature of our setup -- see Lemma (ref) for more details. To gain further intuition, suppose that $\Pi_0=\mathrm{diag}(\pi_{0,1},\pi_{0,2})$ and we want to test if $\mathrm{rank}(\Pi_0)\le 1$. Then by definition
\begin{align}
\phi_r(\Pi_0)=\min\{\pi_{0,1}^2,\pi_{0,2}^2\} .
\end{align}
Note that if $\mathrm{rank}(\Pi_0)\le 1$, then $\pi_{0,1}^2=\pi_{0,2}^2$ if and only if $\mathrm{rank}(\Pi_0)<1$ in which case $\pi_{0,1}=\pi_{0,2}=0$. Hence, $\phi_r$ is not second order differentiable at $\Pi_0$ if and only if $\mathrm{rank}(\Pi_0)<1$ as the map $(\pi_1,\pi_2)\mapsto\min\{\pi_1,\pi_2\}$ is not differentiable precisely when $\pi_1=\pi_2$. In any case, fortunately, $\phi_r$ is second order directionally differentiable, which is sufficient to invoke the second order Delta method as we elaborate next.
\subsection{The Asymptotic Distributions}
With the differentiability established in Proposition (ref), we now derive the asymptotic distributions for the plug-in statistic $\phi_r(\hat\Pi_n)$ where $\hat\Pi_n$ is a generic estimator of $\Pi_0$. This is achieved by appealing to a generalized Delta method for second order directionally differentiable maps Shapiro2000inference,Chen_Fang2015FOD. Towards this end, we impose the following assumption.
\begin{ass}
There is an estimator $\hat{\Pi}_{n}: \{X_i\}_{i=1}^n\to\mathbf M^{m\times k}$ of $\Pi_0\in\mathbf{M}^{m\times k}$ (with $m\ge k$) satisfying ${\tau_{n}}\{\hat{\Pi}_{n}-\Pi_{0}\}\overset{L}{\rightarrow}\mathcal{M}$ for some $\tau_{n}\uparrow \infty$ and random matrix $\mathcal{M}\in\mathbf{M}^{m\times k}$.
\end{ass}
Assumption (ref) simply requires an estimator $\hat\Pi_n$ of $\Pi_0$ that admits an asymptotic distribution. Note that the data need not be i.i.d., $\tau_n$ may be non-$\sqrt n$ and $\mathcal M$ can be non-Gaussian, which is important in, for example, nonstationary time series settings. Moreover, as in Robin_Smith2000rank but in contrast to Cragg_Donald1997infer, the covariance matrix of $\mathrm{vec}(\mathcal M)$ is not required to be nonsingular. Assumption (ref) can be relaxed to accommodate settings where convergence rates across entries of $\hat\Pi_n$ are not homogeneous, as in cointegratoin settings -- see Supplemental Appendix (ref). For ease of exposition, however, we stick to Assumption (ref) in the main text.
Given Proposition (ref) and Assumption (ref), the following theorem delivers the asymptotic distributions of $\phi_r(\hat{\Pi}_{n})$ by the Delta method.
\begin{thm}
If Assumption (ref) holds, then we have, for any $\Pi_0\in\mathbf M^{m\times k}$,
\begin{align}
\tau_{n}\{\phi_r(\hat{\Pi}_{n})-\phi_r(\Pi_{0})\}\overset{L}{\rightarrow}\min_{U\in\Psi(\Pi_{0})} 2\mathrm{tr}(U^\text{\scalebox{0.7}{$\intercal$}}\Pi_{0}^\text{\scalebox{0.7}{$\intercal$}} \mathcal{M} U) .
\end{align}
If in addition $r_0\equiv \mathrm{rank}(\Pi_0)\le r$, then
\begin{align}
\tau_{n}^{2}\phi_r(\hat{\Pi}_{n})\overset{L}{\rightarrow} \sum_{j=r-r_{0}+1}^{k-r_{0}}\sigma^{2}_j(P_{0,2}^\text{\scalebox{0.7}{$\intercal$}} \mathcal{M} Q_{0,2}) ,
\end{align}
where the columns of $P_{0,2}\in\mathbb S^{m\times(m-r_0)}$ and $Q_{0,2}\in\mathbb S^{k\times (k-r_0)}$ are respectively the left and the right singular vectors of $\Pi_{0}$ associated with its zero singular values.
\end{thm}
Theorem (ref) implies that, under $\mathrm{H}_0$ (and so $\tau_{n}\phi_r(\hat{\Pi}_{n})$ is degenerate), the statistic $\tau_{n}^{2}\phi_r(\hat{\Pi}_{n})$ converges in law to a nondegenerate second order limit. Towards constructing critical values, we would then like to estimate the law of the limit. Unfortunately, as shown by Chen_Fang2015FOD, bootstrapping a nondegenerate second order limit is nontrivial; in particular, standard bootstrap schemes such as the nonparametric bootstrap of Efron1979 are necessarily inconsistent even if they are consistent for $\mathcal M$. This predicament is further intensified by the nondifferentiability nature of the map $\phi_r$ Dumbgen1993,Fang_Santos2014HDD, which renders the limits in (ref) highly nonstandard in general. We shall thus present a consistent bootstrap shortly.
We emphasize that the limit of $\tau_n^2\phi_r(\hat\Pi_n)$ in Theorem (ref) is obtained pointwise in each $\Pi_0$ under the {\it entire} null, regardless of whether the truth rank of $\Pi_0$ is strictly less than $r$ or not. To the best of our knowledge, this is the first distributional result for a rank test statistic that accommodates the possibility $\mathrm{rank}(\Pi_0)<r$, at the generality of our setup. In turn, such a result permits us to develop a test that has asymptotic null rejection rates exactly equal to the significance level, and hence is more powerful.
In relating our work to the literature, we note that, if $\tau_n=\sqrt n$, then the plug-in statistic $\tau_n^2\phi_r(\hat\Pi_n)$ is precisely a Robin-Smith statistic (see (ref)), while the KP statistic is simply a Wald-type standardization of it. Though standardization can help obtain pivotal asymptotic distributions under $r_0=r$, this is generally not hopeful whenever $r_0<r$. Since we shall reply on bootstrap for inference, non-pivotalness creates no problems for us. Perhaps more importantly, one may be better off without standardization because it entails invertibility of the weighting matrix in the limit, which may be hard to justify. One might nonetheless interpret the inverse in the KP statistic as a generalized inverse, but consistency of the inverse does not automatically follow from consistency of the covariance matrix estimator without further conditions Andrews1987Wald.
Finally, the limit of $\tau_n^2\phi_r(\hat\Pi_n)$ obtained under $\mathrm{H}_0$ is in fact a weighted sum of independent $\chi^2(1)$ variables if $r_0=r$ and $\mathcal M$ is centered Gaussian, showing consistency of our work with Robin_Smith2000rank. To see this, note that
\begin{align}
\sum_{j=r-r_{0}+1}^{k-r_{0}}\sigma^{2}_j(P_{0,2}^\text{\scalebox{0.7}{$\intercal$}} \mathcal{M} Q_{0,2})=\sum_{j=1}^{k-r}\sigma^{2}_j(P_{0,2}^\text{\scalebox{0.7}{$\intercal$}} \mathcal{M} Q_{0,2}) ,
\end{align}
which is simply the sum of all squared singular values of the $(m-r)\times(k-r)$ matrix $P_{0,2}^\text{\scalebox{0.7}{$\intercal$}} \mathcal{M} Q_{0,2}$, or equivalently the squared Frobenius norm of $P_{0,2}^\text{\scalebox{0.7}{$\intercal$}} \mathcal{M} Q_{0,2}$ Bhatia1997Matrix. Consequently, the limit in (ref) can be rewritten as
\begin{align}
\text{vec}(P_{0,2}^{\text{\scalebox{0.7}{$\intercal$}}}\mathcal{M}Q_{0,2})^{\text{\scalebox{0.7}{$\intercal$}}}\text{vec}(P_{0,2}^{\text{\scalebox{0.7}{$\intercal$}}}\mathcal{M}Q_{0,2})
=\mathrm{vec}(\mathcal M)^\text{\scalebox{0.7}{$\intercal$}} (Q_{0,2}\otimes P_{0,2}) (Q_{0,2}\otimes P_{0,2})^\text{\scalebox{0.7}{$\intercal$}} \mathrm{vec}(\mathcal M) ,
\end{align}
as claimed, where we exploited a property of the vec operator Hamilton1994. Our general limit in (ref) characterizes the channels through which the true rank plays its role, and thus highlights the importance of studying the problem (ref).
\subsection{The Bootstrap Inference}
Since asymptotic distributions of our statistic $\tau_n^2\phi_r(\hat\Pi_n)$ are not pivotal and highly nonstandard in general, in this section we thus aim to develop a consistent bootstrap. This turns out to be quite challenging due to two complications involved.
First, since under $\mathrm H_0$ the first order derivative of $\phi_r$ is degenerate while second order derivative is not (by Proposition (ref)), $\phi_r(\hat\Pi_n^*)$ is necessarily inconsistent even if $\hat\Pi_n^*$ is a consistent bootstrap (in a sense defined below) in estimating the law of $\mathcal M$ Chen_Fang2015FOD, and this remains true in the conventional setup where $\mathrm{rank}(\Pi_0)=r$. Second, the possibility $\mathrm{rank}(\Pi_0)<r$ makes the map $\phi_r$ nondifferentiable -- see Lemma (ref), and hence further complicates the inference Dumbgen1993,Fang_Santos2014HDD. One may resort to the $m$ out of $n$ resampling Shao1994bootstrap or subsampling Politis_Romano1994subsample. However, both methods can be viewed as special cases of our general bootstrap procedure, and that, more importantly, such a perspective enables us to improve upon these existing resampling schemes and to analyze the local properties in a unified and transparent way -- see Remark (ref) and Section (ref).
The insight our bootstrap builds on is that the limit $\phi_{r,\Pi_0}''(\mathcal M)$ in Theorem (ref) is a composition of two unknown components, namely, the limit $\mathcal M$ and the derivative $\phi_{r,\Pi_0}''$. Heuristically, one may therefore obtain a consistent estimator for the law of $\phi_{r,\Pi_0}''(\mathcal M)$ by composing a consistent bootstrap $\hat{\mathcal M}_n^*$ for $\mathcal M$ with an estimator $\hat\phi_{r,n}''$ of $\phi_{r,\Pi_0}''$ that is suitably “consistent.” This is precisely the bootstrap initially proposed in Fang_Santos2014HDD and further developed in Chen_Fang2015FOD and Hong_Li2017numerical. In what follows, we thus commence by estimating the two components separately.
Starting with $\mathcal M$, we note that the law of $\mathcal M$ may be estimated by standard bootstrap or variants of it that suit particular settings. To formalize the notion of bootstrap consistency, we employ the bounded Lipschitz metric Vaart1996 and consider estimating the law of a general random element $\mathbb G$ in a normed space $\mathbb D$ with norm $\|\cdot\|_{\mathbb D}$ -- the space $\mathbb D$ is either $\mathbf M^{m\times k}$ or $\mathbf R$ in this paper. Let $\mathbb G_n^*:\{X_i,W_{ni}\}_{i=1}^n\to\mathbb D$ be a generic bootstrap estimator where $\{W_{ni}\}_{i=1}^{n}$ are bootstrap weights independent of the data $\{X_{i}\}_{i=1}^{n}$. Then we say that the conditional law of $\mathbb G_n^*$ given the data is consistent for the law of $\mathbb G$, or simply $\mathbb G_n^*$ is a consistent bootstrap for $\mathbb G$, if
\begin{align}
\sup_{f\in\mathrm{BL}_1(\mathbb D)}\big |E_W[f(\mathbb G_n^*)]-E[f(\mathbb G)]\big|=o_p(1) ,
\end{align}
where $E_W$ denotes expectation with respect to $\{W_{ni}\}_{i=1}^{n}$ holding $\{X_{i}\}_{i=1}^{n}$ fixed, and
\begin{align}
\mathrm{BL}_1(\mathbb D)\equiv\{f:\mathbb D\to\mathbf R: \sup_{x\in\mathbb D}|f(x)|<\infty,|f(x)-f(y)|\le\|x-y\|_{\mathbb D}\,\forall\,x,y\in\mathbb D\} .
\end{align}
Given the metric, we now proceed by imposing
\begin{ass}
(i) $\hat{\mathcal M}_{n}^{\ast}: \{X_{i},W_{ni}\}_{i=1}^{n}\to\mathbf{M}^{m\times k}$ is a bootstrap estimator with $\{W_{ni}\}_{i=1}^{n}$ independent of $\{X_{i}\}_{i=1}^{n}$; (ii) $\hat{\mathcal M}_{n}^{\ast}$ is a consistent bootstrap for $\mathcal M$.
\end{ass}
Assumption (ref)(i) introduces the bootstrap estimator $\hat{\mathcal M}_{n}^{\ast}$, which may be constructed from nonparametric bootstrap, multiplier bootstrap, general exchangeable bootstrap, block bootstrap, score bootstrap, the $m$ out of $n$ resampling or subsampling. The presence of $\{W_{ni}\}_{i=1}^{n}$ simply characterizes the bootstrap randomness given the data -- see Praestgaard_Wellner1993. For $\hat\Pi_n^*$ a bootstrap analog of $\hat\Pi_n$, it is common to have $\hat{\mathcal M}_{n}^{\ast}=\tau_n\{\hat\Pi_n^*-\hat\Pi_n\}$; if $\hat\Pi_{m_n}^*$ is an analog of $\hat\Pi_n$ constructed based on a subsample of size $m_n$, then one may instead have $\hat{\mathcal M}_{n}^{\ast}=\tau_{m_n}\{\hat\Pi_{m_n}^*-\hat\Pi_n\}$. Assumption (ref)(ii) requires that $\hat{\mathcal M}_{n}^{\ast}$ be consistent in estimating the law of the target limit $\mathcal M$.
Turning to the estimation of $\phi_{r,\Pi_0}''$, we recall by Chen_Fang2015FOD that, given Assumption (ref), the composition $\hat\phi_{r,n}''(\hat{\mathcal M}_n^*)$ is a consistent bootstrap for $\phi_{r,\Pi_0}''(\mathcal M)$ provided $\hat\phi_{r,n}^{\prime\prime}$ is consistent for $\phi_{r,\Pi_0}''$ in the sense that, whenever $M_n\to M$ as $n\to\infty$,
\begin{align}
\hat\phi_{r,n}^{\prime\prime}(M_n)\xrightarrow{p} \phi_{r,\Pi_0}”(M) .
\end{align}
In this regard, there are two general constructions, namely, the numerical estimator and the analytic estimator, as we elaborate next.
The numerical estimator is simply a finite sample analog of (ref) in the definition of second order derivative, i.e., we estimate $\phi^{\prime\prime}_{r,\Pi_{0}}$ by: for any $M\in\mathbf M^{m\times k}$,
\begin{align}
\hat\phi_{r,n}^{\prime\prime}(M)=\frac{\phi_r(\hat\Pi_{n}+\kappa_{n}M)-\phi_r(\hat\Pi_{n})}{\kappa_{n}^{2}} ,
\end{align}
for a suitable $\kappa_n\downarrow 0$, where we have exploited $\phi_{r,\Pi_0}'=0$ under the null. By Chen_Fang2015FOD, (ref) meets the requirement (ref) if $\kappa_n\downarrow 0$ and $\tau_n\kappa_n\to\infty$. Numerical differentiation in the general context of the Delta method dates back to Dumbgen1993, and is recently extended by Hong_Li2017numerical. The numerical estimator enjoys marked simplicity and wide applicability, because it merely requires a sequence $\{\kappa_n\}$ of step sizes satisfying certain rate conditions. There is, however, no general theory to date guiding the choice of $\kappa_n$, a problem that appears challenging Hong_Li2017numerical. In this regard, it may be sensible to employ the analytic estimator instead.
The analytic estimator heavily exploits the analytic structure of the derivative $\phi_{r,\Pi_0}''$, which, by Proposition (ref)(ii), involves three unknown objects, namely, the true rank $r_0$, $P_{0,2}$ and $Q_{0,2}$ -- note that the columns of $P_{0,2}$ and $Q_{0,2}$ are the left and the right singular vectors associated with the zero singular values of $\Pi_0$. We may thus estimate $\phi_{r,\Pi_0}''$ by replacing these unknowns with their estimated counterparts. The key is consistent estimation of $r_0$: given a consistent estimator $\hat r_n$ of $r_0$, we may then obtain estimators $\hat{P}_{2,n}$ and $\hat{Q}_{2,n}$ of $P_{0,2}$ and $Q_{0,2}$ respectively in a straightforward manner as described in Section (ref). One possible construction of $\hat r_n$ is given by (ref). Alternatively, $\hat r_n$ may also be constructed by sequential testing, and the tuning parameter then becomes an adjusted significance level -- see Supplemental Appendix (ref). In any case, by Lemma (ref), we may then obtain a consistent estimator for $\phi_{r,\Pi_0}''$: for any $M\in\mathbf M^{m\times k}$,
\begin{align}
\hat\phi_{r,n}^{\prime\prime}(M)= \sum_{j=r-\hat r_n+1}^{k-\hat{r}_{n}}\sigma^{2}_j(\hat{P}_{2,n}^\text{\scalebox{0.7}{$\intercal$}} M \hat{Q}_{2,n}) .
\end{align}
Similar to the numerical estimator, the analytic estimator (ref) also depends on a tuning parameter, but now through consistent estimation of the rank. An advantage of the latter over the former is that the choice of the tuning parameter is easier to motivate. For example, if $\hat r_n$ is given by (ref), then $\kappa_n$ has a meaningful interpretation, namely, it measures the parsimoniousness in selecting the rank.
Given a significance level $\alpha$, we now formally define our critical value $\hat c_{n,1-\alpha}$ as
\begin{align}
\hat{c}_{n,1-\alpha}\equiv\inf\{c\in\mathbf{R}: P_{W}(\hat\phi_{r,n}^{\prime\prime}(\hat{\mathcal M}_n^*)\leq c)\geq 1-\alpha\} ,
\end{align}
where $P_{W}$ denotes the probability evaluated with respect to $\{W_{ni}\}_{i=1}^n$ holding the data fixed. In practice, we often approximate $\hat{c}_{n,1-\alpha}$ using the following algorithm:
{\sc Step 1:} Compute the derivative estimator $\hat\phi_{r,n}''$ by either (ref), or (ref) and (ref).
{\sc Step 2:} Generate $B$ realizations $\{\hat{\mathcal M}_{n,b}^*\}_{b=1}^B$ of $\hat{\mathcal M}_{n}^*$ based on $B$ bootstrap samples.
{\sc Step 3:} Approximate $\hat c_{n,1-\alpha}$ by the $\lfloor B(1-\alpha)\rfloor$ largest number in $\{\hat\phi_{r,n}''(\hat{\mathcal M}_{n,b}^*)\}_{b=1}^B$.
Our simulations suggest that the analytic method tends to enjoy better size control.
The following theorem establishes that our test has pointwise {\it exact} asymptotic size control under the entire null $\mathrm{H}_0$, and is consistent against any fixed alternatives.
\begin{thm}
Let Assumptions (ref) and (ref) hold, and $\hat{c}_{n,1-\alpha}$ be as in (ref) where $\hat\phi_{r,n}^{\prime\prime}$ is given by either (ref) with $\{\kappa_n\}$ satisfying $\kappa_n\downarrow 0$ and $\tau_n\kappa_n\to\infty$, or (ref) with $\hat r_n\xrightarrow{p} r_0$ under $\mathrm H_0$. If the cdf of the limiting distribution in (ref) is continuous and strictly increasing at its $(1-\alpha)$-quantile for $\alpha\in(0,1)$, then under $\mathrm{H}_{0}$,
\[\lim_{n\to\infty}P(\tau_{n}^{2}\phi_r(\hat{\Pi}_{n})>\hat{c}_{n,1-\alpha})=\alpha~.\]
Furthermore, under $\mathrm{H}_{1}$,
\[\lim_{n\to\infty}P(\tau_{n}^{2}\phi_r(\hat{\Pi}_{n})>\hat{c}_{n,1-\alpha})=1~.\]
\end{thm}
Theorem (ref) shows that our test is not conservative in the pointwise sense while accommodating the possibility $\mathrm{rank}(\Pi_0)<r$. This roots in the simple fact that our critical values are constructed for the pointwise distributions obtained under $\mathrm{H}_0$. By the same token, the power is nontrivial and tends to one against any fixed alternative. We shall further examine the local power properties in Section (ref) and provide numerical evidences in Section (ref). Overall, the theoretical and numerical results manifest superiority of our test in terms of size control and power performance.
In addition to the attractive features mentioned after Assumption (ref), we stress that the bootstrap for $\mathcal M$ may be virtually any consistent resampling scheme, and that no side rank conditions whatsoever are directly imposed beyond those entailed by the restriction that the limiting cdf is continuous and strictly increasing at $c_{1-\alpha}$. Such a quantile restriction is standard as consistent estimation of the limiting laws does not guarantee consistency of critical values -- see, for example, Lemma 11.2.1 in TSH2005. To appreciate how weak this condition is, consider the conventional setup (ref) when $\mathcal M$ is Gaussian. Then each limit under $\mathrm{H}_0'$ is a weighted sum of independent $\chi^2(1)$ random variables -- see our discussions towards the end of Section (ref). Consequently, the quantile condition is automatically satisfied provided the covariance matrix of $\text{vec}(P_{0,2}^{\text{\scalebox{0.7}{$\intercal$}}}\mathcal{M}Q_{0,2})$ is nonzero (i.e., nonzero rank), which is precisely Assumption 2.4 in Robin_Smith2000rank. In contrast, Kleibergen_Paap2006rank require nonsingularity of the same matrix (i.e., full rank).
Despite the irregular natures of the problem, computation of our testing statistic and the critical values are quite simple as both involve only calculations of singular value decompositions, for which there are commands in common computation softwares. In particular, $\hat c_{n,1-\alpha}$ in practice is set to be the $(1-\alpha)$-quantile of
\begin{align}
\hat\phi_{r,n}^{\prime\prime}(\hat{\mathcal M}_{n,1}^*), \hat\phi_{r,n}^{\prime\prime}(\hat{\mathcal M}_{n,2}^*), \ldots, \hat\phi_{r,n}^{\prime\prime}(\hat{\mathcal M}_{n,B}^*) .
\end{align}
Therefore, in each repetition, the numerical and the analytic approaches simply entail singular value decompositions of $\hat{\Pi}_{n}+\kappa_{n}\hat{\mathcal M}_{n,b}^*$ and $\hat{P}_{2,n}^\text{\scalebox{0.7}{$\intercal$}} \hat{\mathcal M}_{n,b}^*\hat{Q}_{2,n}$ respectively.
A common feature of our previous two tests is their dependence on a tuning parameter -- see (ref) and (ref). To mitigate concerns on sensitivity to the choice of tuning parameters, we next develop a two-step test by exploiting the structure in (ref). The intuition is as follows. The estimator (ref), though consistent, may differ from the truth in finite samples. We would thus like to control $P(\hat r_n= r_0)$, for which (ref) may not be appropriate as it appears challenging to bound $P(\hat r_n= r_0)$. Instead, we may obtain a suitable estimator $\hat r_n$ by a sequential testing procedure -- see Theorem (ref). Specifically, we sequentially test $\mathrm{rank}(\Pi_0)=0,1,\ldots,k-1$ at level $\beta<\alpha$, and set $\hat r_n=j^*$ if accepting $\mathrm{rank}(\Pi_0)=j^*$ is the first acceptance, and $\hat r_n=k$ if no acceptance occurs. In this regard, we recommend the KP test as it is tuning parameter free and does not require additional simulations.\footnote{If estimation of $r_0$ is one's {\it ultimate} goal (rather than an intermediate step for test), then it may be desirable to instead employ our tests in the sequential procedure, as existing tests may lead to estimators that are not as accurate when $\Pi_0$ is “local to degeneracy” -- see Section (ref) for simulation evidences.} Table (ref) compares the empirical probabilities of $\{\hat r_n=r_0\}$ for $\hat r_n$ obtained by (ref) and the sequential KP test respectively, based on the same simulation data from Section (ref) when $d>1$. The empirical probabilities for (ref) are close to one when $\kappa_n=n^{-1/4}$ (as chosen in Section 2) or $\in\{n^{-1/4}, 1.5n^{-1/4}, n^{-1/5}, 1.5n^{-1/5}\}$ (omitted due to space limitation), but may be far away from one or even close to zero for other choices. On the other hand, the sequential approach leads to rank estimators with empirical probabilities approximately $1-\beta$ across our choices of $\beta$.
{
{3pt}
\begin{table}[!htbp]
\caption{Estimation of $\mathrm{rank}(\Pi_0)$ Defined by the Model (ref)-(ref)}
\begin{tabular}{ccccccccccccccc}
\hline\hline
\multirow{2}{*}{\makecell{$d$}} & &\multicolumn{5}{c}{Choices of $\kappa_n$ in (ref)}&&\multicolumn{6}{c}{Choices of $\beta$ for the Sequential Method} \\
\cline{3-7} \cline{9-14}
& & $n^{-1/4}$ & $n^{-1/3}$&$1.5n^{-1/3}$&$n^{-2/5}$&$1.5n^{-2/5}$& & $\alpha/5$&$\alpha/10$&$\alpha/15$&$\alpha/20$&$\alpha/25$&$\alpha/30$ \\
\cmidrule{1-14}
$2$ &&1.0000&0.9975&1.0000&0.6679&0.9618&& 0.9902&0.9947&0.9965&0.9974&0.9975&0.9979\\
$3$ &&1.0000&0.8516&0.9988&0.2246&0.7862&&0.9908&0.9951&0.9958&0.9963&0.9976&0.9980\\
$4$ &&0.9995&0.5550&0.9922&0.0249&0.4474&&0.9877&0.9949&0.9963&0.9972&0.9976&0.9981\\
$5$ &&0.9977&0.2176&0.9581&0.0003&0.1420&&0.9861&0.9933&0.9958&0.9968&0.9976&0.9979\\
$6$ &&0.9899&0.0422&0.8557&0.0000&0.0203&&0.9840&0.9916&0.9946&0.9960&0.9967&0.9967\\
\hline\hline
\end{tabular}
\end{table}
}
Given an estimator $\hat r_n$ with $P(\hat r_n=r_0)\ge 1-\beta$ (approximately) for some $\beta<\alpha$, the two-step test now goes as follows. In the first step, we reject $\mathrm{H}_0$ if $\hat r_n>r$; otherwise we plug $\hat r_n$ into (ref) in the second step and reject $\mathrm{H}_0$ if $\tau_n^2\phi_r(\hat\Pi_n)>\hat c_{n,1-\alpha+\beta}$. Note that the significance level in the second step is adjusted to be $\alpha-\beta$ in order to take into account the event of false selection (which has probability $\beta$). Formally, letting
\begin{align}
\psi_n=1\{\hat r_n>r\text{ or }\tau_n^2\phi_r(\hat\Pi_n)>\hat c_{n,1-\alpha+\beta}\} ,
\end{align}
we then reject the null $\mathrm H_0$ if $\psi_n=1$ and fail to reject otherwise. Our next theorem shows that the two-step procedure controls size and is consistent.
\begin{thm}
Suppose that Assumptions (ref) and (ref) hold, and that the cdf of the limit distribution in (ref) is continuous and strictly increasing at its $(1-\alpha+\beta)$-quantile for $\alpha\in(0,1)$ and $\beta\in(0,\alpha)$. Let $\psi_n$ be the test given by (ref). Then, under $\mathrm{H}_{0}$,
\[
\limsup_{n\to\infty}E[\psi_n]\le \alpha
\]
provided $\liminf_{n\to\infty} P(\hat r_n=r_0)\ge 1-\beta$, and, under $\mathrm{H}_{1}$,
\[
\lim_{n\to\infty}E[\psi_n]=1~.
\]
\end{thm}
The idea of the two-step test may be found in Loh1985New, BergerBoos1994Pvalues, and Silvapulle1996Nuisance, and has recently been employed in the context of moment inequality models Andrews_Barwick2012Inference,RomanoShaikhWolf2014TwoStep. A common feature that our test shares here is that the size control is not exact, i.e., we cannot show the size is equal to $\alpha$. This raises the concern that the test may be potentially conservative. Nonetheless, it is possible to derive a lower bound of the asymptotic size which is close to $\alpha$ by choosing a small $\beta$ -- see RomanoShaikhWolf2014TwoStep for a similar feature. Summarizing, there are two (types of) test procedures: one rejects $\mathrm{H}_0$ if $\tau_n^2\phi_r(\hat\Pi_n)>\hat c_{n,1-\alpha}$ with $\hat c_{n,1-\alpha}$ computed according to (ref), and the other one applies when one has control over $P(\hat r_n=r_0)$: if $\liminf_{n\to\infty} P(\hat r_n=r_0)\ge 1-\beta$, we reject if $\hat r_n>r$ or $\tau_n^2\phi_r(\hat\Pi_n)>\hat c_{n,1-\alpha+\beta}$. Our simulation results in Section (ref) show that the two-step procedure produces results that are quite insensitive to our choice of $\beta$.
\begin{rem}
The $m$ out of $n$ bootstrap and the subsampling are special cases of our bootstrap procedure. For example, the former amounts to $\hat{\mathcal M}_{n}^{\ast}=\tau_{m_n}\{\hat\Pi_{m_n}^*-\hat\Pi_n\}$ with $\hat\Pi_{m_n}^*$ constructed based on subsamples of size $m_n$ (obtained through resampling with replacement), and the derivative estimator $\hat\phi_{r,n}''$ given by (ref) with $\kappa_n=m_n^{-1}$. Subsampling is precisely the same procedure except that the subsamples are obtained without replacement. In other words, these procedures estimate the derivative through (ref) implicitly and automatically when the subsample size is properly chosen, combining the two-steps into one single step. By disentangling estimation of the two ingredients, however, we may better estimate both the derivative $\phi_{r,\Pi_0}''$ (through exploiting the structure of the derivative and a choice of the tuning parameter) and the law of the limit $\mathcal M$ (using full samples), which may in turn lead to efficiency improvement. Moreover, such a perspective enables us to establish conditions under which tests based on these resampling schemes have local size control and nontrivial power, properties not guaranteed in general and nontrivial to analyze otherwise AndrewsandGuggen2010ET. \qed
\end{rem}
\subsubsection{Local Power Properties}
Having established size control and consistency, we next aim to obtain a more precise characterization of the quality of our tests by studying the local power properties Neyman1937LocalPower. Following Cragg_Donald1997infer, we thus proceed by imposing
\begin{assprime}{Ass: Weaklimit: Pihat}
(i) $\mathrm{rank}(\Pi_{0,n})> r$ for all $n$; (ii) $\tau_n\{\Pi_{0,n}-\Pi_0\}\to \Delta$ for some $\Pi_0$ with $\mathrm{rank}(\Pi_0)\le r$ and nonrandom $\Delta$; (iii) $\tau_n\{\hat\Pi_n-\Pi_{0,n}\}\overset{L_n}{\to}\mathcal M$ for some $\tau_n\uparrow \infty$, where $\overset{L_n}{\to}$ denotes convergence in law along distributions of the data associated with $\{\Pi_{0,n}\}$.
\end{assprime}
Assumption (ref)(i)(ii) formally defines $\{\Pi_{0,n}\}$ as a sequence of local alternatives that approaches some $\Pi_0$ in the null at the convergence rate $\tau_n$, while Assumption (ref)(iii) formalizes the notion that the {\it asymptotic} distributions of $\hat\Pi_n$ should remain unchanged in response to {\it small} (finite sample) perturbations of the data generating processes, a property that may be verified through, for example, the framework of limits of statistical experiments Vaart1998,HallinAkkerWerker2016Coint.
Our next result characterizes the asymptotic behaviors of the testing statistic $\tau_n^2\phi_r(\hat\Pi_n)$ under local alternatives that satisfy Assumption (ref).
\begin{pro}
If Assumption (ref) holds, then it follows that
\begin{align}
\tau_n^2\phi_r(\hat\Pi_n)\overset{L_n}{\to} \sum_{j=r-r_{0}+1}^{k-r_{0}}\sigma^{2}_j(P_{0,2}^\text{\scalebox{0.7}{$\intercal$}} (\mathcal M+\Delta) Q_{0,2}) .
\end{align}
\end{pro}
Proposition (ref) includes Theorem (ref) as a special case with $\Pi_{0,n}=\Pi_0$ for all $n$ so that $\Delta=0$. The main utility of this result is to analyze the asymptotic local power function. In what follows, we focus on the one-step tests for conciseness and transparency, though analogous results hold for the two-step test $\psi_n$. Thus, if the local alternatives $\{\Pi_{0,n}\}$ in Assumption (ref) approach $\Pi_0$ in the sense of contiguity Roussas1972contiguity,Rothenberg1984Handbook,\footnote{This means that if (any) $T_n$ is negligible (i.e., of order $o_p(1)$) under $\Pi_0$ then it remains so under $\Pi_{0,n}$. Thus, contiguity simply formalizes the notion that the effect of “small” perturbations is negligible.} then we may obtain a lower bound as follows:
\begin{align}
\liminf_{n\to\infty}P_{n}(\tau_n^2\phi_r(\hat\Pi_n)>\hat c_{n,1-\alpha}) \ge P(\sum_{j=r-r_{0}+1}^{k-r_{0}}\sigma^{2}_j(P_{0,2}^\text{\scalebox{0.7}{$\intercal$}} (\mathcal M+\Delta) Q_{0,2})>c_{1-\alpha}) ,
\end{align}
where $P_n$ denotes probability evaluated under $\Pi_{0,n}$. While it appears challenging to prove that the asymptotic local power is nontrivial under arbitrary local alternatives, there is, nonetheless, an interesting case under which the asymptotic local power can be proven to be nontrivial. This is the conventional setup where $\mathrm{rank}(\Pi_0)$ is exactly equal to the hypothesized value $r$ and $\mathcal M$ is centered Gaussian. Since the derivative $\phi_{r,\Pi_0}''$ then coincides with the squared Frobenius norm -- see Proposition (ref)(ii), we have along contiguous local alternatives that
\begin{align}
\liminf_{n\to\infty}P_n(\tau_n^2\phi_r(\hat\Pi_n)>\hat c_{n,1-\alpha})\ge P(\|P_{0,2}^\text{\scalebox{0.7}{$\intercal$}}(\mathcal M+\Delta)Q_{0,2}\|^2 >c_{1-\alpha}) .
\end{align}
An application of Anderson's lemma -- see, for example, Lemma 3.11.4 in Vaart1996 -- then yields
\begin{align}
P(\|P_{0,2}^\text{\scalebox{0.7}{$\intercal$}}(\mathcal M+\Delta)Q_{0,2}\|^2 >c_{1-\alpha})\ge P(\|P_{0,2}^\text{\scalebox{0.7}{$\intercal$}} \mathcal M Q_{0,2}\|^2 >c_{1-\alpha})=\alpha .
\end{align}
If the localization parameter $\Delta$ is nontrivial (i.e., $\Delta\neq 0$) and belongs to the support of $\mathcal M$ -- which is the case, for example, if the covariance matrix of $\mathrm{vec}(\mathcal M)$ is nonsingular, then by Lemma B.4 in Chen_Santos2015 (a strengthening of Anderson's lemma), the asymptotic local lower is in fact nontrivial, i.e.,
\begin{align}
P(\|P_{0,2}^\text{\scalebox{0.7}{$\intercal$}}(\mathcal M+\Delta)Q_{0,2}\|^2 >c_{1-\alpha})>\alpha .
\end{align}
In view of the irregularities of the problem (ref), one may also be interested in the size control of our test. Under Assumption (ref) but with (i) replaced by $\mathrm{rank}(\Pi_{0,n})\le r$ for all $n\in\mathbf N$ so that the contiguous perturbations occur under the null, we may obtain
\begin{align}
\limsup_{n\to\infty}P_n(\tau_n^2\phi_r(\hat\Pi_n)>\hat c_{n,1-\alpha}) \le P(\sum_{j=r-r_{0}+1}^{k-r_{0}}\sigma^{2}_j(P_{0,2}^\text{\scalebox{0.7}{$\intercal$}} (\mathcal M+\Delta) Q_{0,2})\ge c_{1-\alpha}) .
\end{align}
Now suppose $\mathrm{rank}(\Pi_0)=r$ but without requiring $\mathcal M$ to be centered nor Gaussian. Since $\phi_r(\Pi_{0,n})=\phi_r(\Pi_0)=0$, it follows by Assumption (ref)(ii) and Proposition (ref) that
\begin{align}
0=\lim_{n\to\infty}\tau_n^2\{\phi_r(\Pi_{0,n})-\phi_r(\Pi_0)\}=\phi_{r,\Pi_0}”(\Delta)=\|P_{0,2}^\text{\scalebox{0.7}{$\intercal$}} \Delta Q_{0,2}\|^2 .
\end{align}
Hence, we have $P_{0,2}^\text{\scalebox{0.7}{$\intercal$}} \Delta Q_{0,2}=0$ and consequently,
\begin{align}
P(\sum_{j=r-r_{0}+1}^{k-r_{0}}\sigma^{2}_j(P_{0,2}^\text{\scalebox{0.7}{$\intercal$}} (\mathcal M+\Delta) Q_{0,2})\ge c_{1-\alpha})= \alpha ,
\end{align}
if the quantile restrictions on $c_{1-\alpha}$ as in Theorem (ref) hold. Size control under arbitrary local perturbations in $\mathrm H_0$, unfortunately, appears (to us) as challenging as establishing nontrivial local power under arbitrary local alternatives. We pose these as open questions, and leave them for future study.
\subsubsection{Illustration: Identification in Linear IV Models}
We now illustrate how to apply our framework by testing identification in linear IV models due to their simplicity and popularity. Let $(Y,Z^\text{\scalebox{0.7}{$\intercal$}})^\text{\scalebox{0.7}{$\intercal$}}\in\mathbf R^{1+k}$ satisfy:
\begin{align}
Y=Z^\text{\scalebox{0.7}{$\intercal$}}\beta_0+u ,
\end{align}
where $\beta_0\in\mathbf R^k$ and $u$ is an error term. Let $V\in\mathbf R^m$ be an instrument variable with $E[Vu]=0$ and $m\ge k$. Then global identification of $\beta_0$ requires $E[VZ^\text{\scalebox{0.7}{$\intercal$}}]$ to be of full rank. Thus, identification of $\beta_0$ may be tested by examining (ref) with
\begin{align}
\Pi_0=E[VZ^\text{\scalebox{0.7}{$\intercal$}}] \text{ and } r=k-1 .
\end{align}
The hypotheses in (ref) may be restrictive since it is generally unknown if $\text{rank}(\Pi_{0})\geq k-1$. Analogous rank conditions also arise for global identification in simultaneous linear equation models KoopmansHood1953Simultaneous,Fisher1961ID and in models with misclassification errors Hu2008Identification, and for local identification in nonlinear/nonparametric models Rothenberg1971identification,Roehrig1988Conditions,Chesher2003ID,Matzkin2008NPSimul,Chen_Chernozhukov_Lee_Newey2014LocalID and in DSGE models Canov_Sala2009,Komunjer_Ng2011DSGE.
To apply our framework, let $\{V_{i},Z_{i}\}_{i=1}^{n}$ be an i.i.d.\ sample. Then the estimator
\begin{align}
\hat{\Pi}_{n} = \frac{1}{n}\sum_{i=1}^{n}V_{i}Z_{i}^{\text{\scalebox{0.7}{$\intercal$}}}
\end{align}
satisfies Assumption (ref) for $\tau_n=\sqrt n$ and some centered Gaussian matrix $\mathcal M$ under suitable moment restrictions. In turn, let $\{Z^{\ast}_{i},V^{\ast}_{i}\}_{i=1}^{n}$ be an i.i.d.\ sample drawn with replacement from $\{Z_{i},V_{i}\}_{i=1}^{n}$. Then $\hat{\mathcal M}_n^*\equiv\sqrt n\{\hat\Pi_n^*-\hat\Pi_n\}$ with $\hat{\Pi}_{n}^{\ast}$ given by
\begin{align}
\hat{\Pi}_{n}^{\ast} \equiv \frac{1}{n}\sum_{i=1}^{n}V^{\ast}_{i}Z_{i}^{\ast\text{\scalebox{0.7}{$\intercal$}}}= \frac{1}{n}\sum_{i=1}^{n}W_{ni}V_{i}Z_{i}^{\text{\scalebox{0.7}{$\intercal$}}} ,
\end{align}
where $(W_{n1},\ldots,W_{nn})$ is multinomial over $n$ categories with probabilities $(n^{-1},\ldots,n^{-1})$, satisfies Assumption (ref) -- see, for example, Theorem 23.4 in Vaart1998. We have thus verified the main assumptions.
Empirical research, however, is often faced with clustered data; e.g., micro-level data often cluster on geographical regions such as cities or states. To illustrate, suppose that there are $G$ clusters where $G$ is large, and the $g$th cluster has observations $\{V_{gi},Z_{gi}\}_{i=1}^{n_g}$. The data are independent across clusters but may otherwise be correlated within each cluster. Let $n\equiv\sum_{g=1}^{G}n_g$. In these settings, $\Pi_0$ is identified as the probability limit of
\begin{align}
\hat\Pi_n\equiv \frac{1}{n}\sum_{g=1}^{G}V_{g}^\text{\scalebox{0.7}{$\intercal$}} Z_{g}
\end{align}
as $G\to\infty$, where $V_g\equiv[V_{g1},\ldots,V_{gn_g}]^\text{\scalebox{0.7}{$\intercal$}}$ and $Z_g\equiv[Z_{g1},\ldots,Z_{gn_g}]^\text{\scalebox{0.7}{$\intercal$}}$. Assumption (ref) holds for $\tau_n=\sqrt n$ and some centered Gaussian matrix $\mathcal M$, by the Lindeberg-Feller type central limit theorem. Following CameronGelbachMiller2008BootCluster, we may construct
\begin{align}
\hat{\mathcal M}_n^*\equiv \frac{1}{n}\sum_{g=1}^{G}W_g\{V_{g}^{\text{\scalebox{0.7}{$\intercal$}}} Z_{g}-\hat\Pi_n\} ,
\end{align}
where $(W_1,\ldots,W_G)$ may be a multinomial vector over $G$ categories with probabilities $(1/G,\ldots,/1G)$ (corresponding to the pairs cluster bootstrap) or other weights (such as those leading to the cluster wild bootstrap); see also DjogbenouMackinnonNielsen2018WCB.
For the convenience of practitioners, we next provide an implementation guide of our two-step test at significance level $\alpha$.
{\sc Step 1:} (a) Sequentially test $\mathrm{rank}(\Pi_0)=0,1,\ldots,k-1$ at level $\beta$ (e.g., $\beta=\alpha/10$) based on $\hat\Pi_n$ using the KP test and obtain the rank estimator $\hat r_n$; (b) Reject $\mathrm{H}_0$ if $\hat r_n=k$ and move on to the next step otherwise.
{\sc Step 2:} (a) Draw $B$ bootstrap samples by either the empirical bootstrap or the cluster bootstrap depending on if clustering is present, construct $\{\hat{\mathcal M}_{n,b}^*\}_{b=1}^B$ accordingly (i.e., as in (ref) or (ref)), and set $\hat c_{1-\alpha+\beta}$ to be the $\lfloor B(1-\alpha+\beta)\rfloor$ largest number in
\begin{align}
\sum_{j=r-\hat r_n+1}^{k-\hat r_n}\sigma_{j}^2(\hat P_{2,n}^\text{\scalebox{0.7}{$\intercal$}}\hat{\mathcal M}_{n,1}^* \hat Q_{2,n}) ,\ldots,\sum_{j=r-\hat r_n+1}^{k-\hat r_n}\sigma_{j}^2(\hat P_{2,n}^\text{\scalebox{0.7}{$\intercal$}}\hat{\mathcal M}_{n,B}^* \hat Q_{2,n}) ,
\end{align}
where $\hat P_{2,n}$ and $\hat Q_{2,n}$ are from the singular value decomposition of $\hat\Pi_n$ as before; (b) Reject $\mathrm{H}_0$ if $n\sigma_{\min}^2(\hat\Pi_n)>\hat c_{1-\alpha+\beta}$ with $\sigma_{\min}(\hat\Pi_n)$ the smallest singular value of $\hat\Pi_n$.
For our one-step test based on (ref) and (ref), one may directly proceed with Step 2, but with $\hat r_n$ constructed from (ref) and reject if $n\sigma_{\min}^2(\hat\Pi_n)>\hat c_{1-\alpha}$.
\section{Simulation Studies}
In this section, we examine the finite sample performance of our inferential framework by Monte Carlo simulations. First, we compare our tests with the multiple KP test in more complicated data environments with heteroskedasticity, serial correlation and different sample sizes. We shall pay special attention to the choices of tuning parameters. We refer the reader to Supplemental Appendix (ref) where we provide additional comparisons with Kleibergen_Paap2006rank based on their simulation designs and a real dataset that they use. Second, we also conduct simulations to assess the performance of our rank estimators, obtained by a sequential testing procedure employed in the literature and formalized in Supplemental Appendix (ref).
We commence by considering the following linear model
\begin{align}
Z_{t} =\Pi_0^\text{\scalebox{0.7}{$\intercal$}} V_{t} +V_{1,t}u_{t} ,
\end{align}
where $Z_t\in\mathbf R^4$ for all $t$, $\{V_{t}\}\overset{\text{i.i.d.}}{\sim}N(0,I_{4})$ and $\{u_t\}$ are generated according to
\begin{align}
u_{t}=\epsilon_{t}-\frac{1}{4}\mathbf{1}_{4}\mathbf{1}_{4}^{\text{\scalebox{0.7}{$\intercal$}}}\epsilon_{t-1}
\end{align}
with $\{\epsilon_{t}\}\overset{\text{i.i.d.}}{\sim}N(0,I_{4})$ independent of $\{V_{t}\}$, and $V_{1,t}$ the first entry of $V_{t}$. Moreover, we configure $\Pi_0$ as: for $\delta\in\{0, 0.1, 0.3, 0.5\}$,
\begin{align}
\Pi_{0}=\textrm{diag}(\mathbf{1}_{2},\mathbf{0}_{2})+\delta I_{4} .
\end{align}
We test the hypotheses in (ref) for $r\in\{2,3\}$ at level $\alpha=5\%$. Thus, for both cases, $\mathrm{H}_0$ is true if and only if $\delta=0$, and they respectively correspond to $\mathrm{rank}(\Pi_0)=r$ and $\mathrm{rank}(\Pi_0)<r$ under $\mathrm H_0$. We estimate $\Pi_{0}$ by $\hat{\Pi}_{n}=\frac{1}{n}\sum_{t=1}^{n}V_{t}Z_{t}^{\text{\scalebox{0.7}{$\intercal$}}}$ for sample sizes $n\in\{50,100,300,1000\}$, and for each $n$, the number of simulation replications is set to be 5,000 with 500 bootstrap repetitions for each replication. As the data exhibit first order autocorrelation, we adopt the circular block bootstrap Politis_Romano1992circular with block size $b=2$. To implement the multiple KP test, labelled KP-M, we estimate the variance of $\mathrm{vec}(\hat\Pi_n)$ by the HACC estimator with one lag West1997AnotherHAC. To carry out our tests, we choose $\kappa_n\in\{n^{-2/5},1.5n^{-2/5}, n^{-1/5}, n^{-1/4}, n^{-1/3}, 1.5n^{-1/5}$, $1.5n^{-1/4}, 1.5n^{-1/3}\}$ for both the numerical estimator in (ref) and the analytic estimator in (ref) and (ref), and $\beta\in\{\alpha/5,\alpha/10,\alpha/15,\alpha/20$, $\alpha/25,\alpha/30\}$ for the two-step test. As in Section (ref), we respectively label these three tests as CF-N, CF-A, and CF-T.
{3pt}
\begin{table}[!htbp]
\begin{threeparttable}
\caption{Rejection rates of rank tests for the model (ref) at $\alpha=5\%$\tnote{\dag}}
\begin{tabular}{cccccccccccccccc}
\hline\hline
&\multirow{2}{*}{\makecell{Sample\\ size}} &\multicolumn{3}{c}{CF-T}&&\multicolumn{3}{c}{CF-A}&&\multicolumn{3}{c}{CF-N}&&\multirow{2}{*}{KP-M}&\\
\cline{3-5}\cline{7-9}\cline{11-13}
&&$\alpha/10$&$\alpha/15$&$\alpha/20$ &&$n^{-1/5}$&$n^{-1/4}$&$n^{-1/3}$ &&$n^{-1/5}$&$n^{-1/4}$&$n^{-1/3}$ &&\\
\cmidrule{2-15}
&&\multicolumn{13}{c}{Rejection rates for $r=2$}&\\
\cline{2-15}
&&\multicolumn{13}{c}{$\delta=0$}&\\
\cline{2-15}
&$50$ &0.17&0.17&0.17&&0.04&0.04&0.04&&0.29&0.28&0.21&&0.08&\\
&$100$ &0.08&0.08&0.08&&0.04&0.04&0.04&&0.23&0.20&0.12&&0.08&\\
&$300$ &0.04&0.04&0.04&&0.05&0.05&0.05&&0.16&0.12&0.05&&0.06&\\
&$1000$ &0.04&0.04&0.04&&0.05&0.05&0.05&&0.11&0.08&0.04&&0.05&\\
\cline{2-15}
&&\multicolumn{13}{c}{$\delta=0.1$}&\\
\cline{2-15}
&$50$ &0.23&0.23&0.23&&0.08&0.08&0.08&&0.37&0.35&0.27&&0.13&\\
&$100$ &0.18&0.17&0.17&&0.12&0.12&0.12&&0.38&0.34&0.23&&0.19&\\
&$300$ &0.34&0.34&0.34&&0.35&0.35&0.35&&0.57&0.51&0.36&&0.44&\\
&$1000$ &0.89&0.90&0.90&&0.90&0.90&0.90&&0.95&0.92&0.88&&0.92&\\
\cline{2-15}
&&\multicolumn{13}{c}{$\delta=0.3$}&\\
\cline{2-15}
&$50$ &0.67&0.67&0.66&&0.48&0.48&0.48&&0.80&0.79&0.72& &0.40&\\
&$100$ &0.85&0.85&0.85&&0.80&0.80&0.80&&0.95&0.93&0.89&&0.77&\\
&$300$ &1.00&1.00&1.00&&1.00&1.00&1.00&&1.00&1.00&1.00&&1.00&\\
&$1000$ &1.00&1.00&1.00&&1.00&1.00&1.00&&1.00&1.00&1.00&&1.00&\\
\cline{2-15}
&&\multicolumn{13}{c}{$\delta=0.5$}&\\
\cline{2-15}
&$50$ &0.95&0.95&0.94&&0.89&0.89&0.89&&0.98&0.98&0.97&&0.55&\\
&$100$ &1.00&1.00&1.00&&1.00&1.00&1.00&&1.00&1.00&1.00&&0.87&\\
&$300$ &1.00&1.00&1.00&&1.00&1.00&1.00&&1.00&1.00&1.00&&1.00&\\
&$1000$ &1.00&1.00&1.00&&1.00&1.00&1.00&&1.00&1.00&1.00&&1.00&\\
\cmidrule{1-16}
&&\multicolumn{13}{c}{Rejection rates for $r=3$}&\\
\cline{2-15}
&&\multicolumn{13}{c}{$\delta=0$}&\\
\cline{2-15}
&$50$ &0.09&0.09&0.10&&0.07&0.06&0.04&&0.14&0.14&0.12&&0.01&\\
&$100$ &0.06&0.07&0.07&&0.06&0.06&0.03&&0.12&0.12&0.09&&0.01&\\
&$300$ &0.04&0.05&0.05&&0.05&0.05&0.03&&0.09&0.08&0.06&&0.01&\\
&$1000$ &0.05&0.05&0.05&&0.06&0.06&0.05&&0.08&0.07&0.05&&0.00&\\
\cline{2-15}
&&\multicolumn{13}{c}{$\delta=0.1$}&\\
\cline{2-15}
&$50$ &0.12&0.12&0.12&&0.10&0.09&0.05&&0.18&0.18&0.16&&0.01&\\
&$100$ &0.12&0.13&0.13&&0.13&0.11&0.06&&0.21&0.19&0.16&&0.02&\\
&$300$ &0.25&0.26&0.27&&0.32&0.29&0.16&&0.38&0.36&0.31&&0.09&\\
&$1000$ &0.63&0.65&0.67&&0.82&0.81&0.59&&0.84&0.82&0.77&&0.54&\\
\cline{2-15}
&&\multicolumn{13}{c}{$\delta=0.3$}&\\
\cline{2-15}
&$50$ &0.43&0.44&0.45&&0.39&0.33&0.25&&0.57&0.56&0.52&&0.12&\\
&$100$ &0.61&0.63&0.64&&0.66&0.57&0.50&&0.80&0.79&0.74&&0.43&\\
&$300$ &0.96&0.96&0.96&&0.98&0.96&0.96&&1.00&0.99&0.99&&0.96&\\
&$1000$ &1.00&1.00&1.00&&1.00&1.00&1.00&&1.00&1.00&1.00&&1.00&\\
\cline{2-15}
&&\multicolumn{13}{c}{$\delta=0.5$}&\\
\cline{2-15}
&$50$ &0.76&0.77&0.78&&0.68&0.64&0.63&&0.88&0.88&0.84&&0.37&\\
&$100$ &0.92&0.93&0.93&&0.92&0.91&0.91&&0.99&0.99&0.98&&0.79&\\
&$300$ &1.00&1.00&1.00&&1.00&1.00&1.00&&1.00&1.00&1.00&&1.00&\\
&$1000$ &1.00&1.00&1.00&&1.00&1.00&1.00&&1.00&1.00&1.00&&1.00&\\
\hline\hline
\end{tabular}
\begin{tablenotes}
• The three values under CF-T are the choices of $\beta$, and those under CF-A and CF-N are the choices of $\kappa_n$ as in (ref) and (ref) respectively.
\end{tablenotes}
\end{threeparttable}
\end{table}
{4pt}
\begin{sidewaystable}[!htbp]
\begin{threeparttable}
\caption{Additional results on rejection rates of rank tests for the model (ref) with $r=2$, at $\alpha=5\%$\tnote{\dag}}
\begin{tabular}{cccccccccccccccccccc}
\hline\hline
&\multirow{2}{*}{\makecell{Sample\\ size}} &\multicolumn{3}{c}{CF-T}&&\multicolumn{5}{c}{CF-A}&&\multicolumn{5}{c}{CF-N}&&\multirow{2}{*}{KP-M}&\\
\cline{3-5}\cline{7-11}\cline{13-17}
&&$\alpha/5$&$\alpha/25$&$\alpha/30$ &&$1.5n^{-1/5}$&$1.5n^{-1/4}$&$1.5n^{-1/3}$&$n^{-2/5}$&$1.5n^{-2/5}$ &&$1.5n^{-1/5}$&$1.5n^{-1/4}$&$1.5n^{-1/3}$&$n^{-2/5}$&$1.5n^{-2/5}$ & & & \\
\cline{2-19}
&&\multicolumn{17}{c}{$\delta=0$}&\\
\cline{2-19}
&$50$ &0.16&0.17&0.17&&0.08&0.05&0.04&0.04&0.04&&0.31&0.30&0.29&0.13&0.25&&0.08&\\
&$100$ &0.08&0.08&0.08&&0.04&0.04&0.04&0.04&0.04&&0.26&0.24&0.20&0.06&0.15&&0.08&\\
&$300$ &0.03&0.04&0.04&&0.05&0.05&0.05&0.05&0.05&&0.20&0.18&0.11&0.02&0.06&&0.06&\\
&$1000$ &0.04&0.05&0.05&&0.05&0.05&0.05&0.05&0.05&&0.15&0.12&0.06&0.02&0.03&&0.05&\\
\cline{2-19}
&&\multicolumn{17}{c}{$\delta=0.1$}&\\
\cline{2-19}
&$50$ &0.23&0.23&0.24&&0.11&0.09&0.08&0.08&0.08&&0.39&0.39&0.36&0.17&0.31&&0.13&\\
&$100$ &0.17&0.17&0.17&&0.12&0.12&0.12&0.12&0.12&&0.41&0.40&0.34&0.12&0.26&&0.19&\\
&$300$ &0.34&0.34&0.34&&0.35&0.35&0.35&0.35&0.35&&0.63&0.60&0.50&0.21&0.37&&0.44&\\
&$1000$ &0.88&0.90&0.90&&0.90&0.90&0.90&0.90&0.90&&0.97&0.96&0.91&0.80&0.87&&0.92&\\
\cline{2-19}
&&\multicolumn{17}{c}{$\delta=0.3$}&\\
\cline{2-19}
&$50$ &0.67&0.66&0.66&&0.49&0.48&0.48&0.48&0.48&&0.81&0.81&0.79&0.61&0.76&&0.40&\\
&$100$ &0.86&0.84&0.84&&0.80&0.80&0.80&0.80&0.80&&0.95&0.95&0.94&0.79&0.90&&0.77&\\
&$300$ &1.00&1.00&1.00&&1.00&1.00&1.00&1.00&1.00&&1.00&1.00&1.00&1.00&1.00&&1.00&\\
&$1000$ &1.00&1.00&1.00&&1.00&1.00&1.00&1.00&1.00&&1.00&1.00&1.00&1.00&1.00&&1.00&\\
\cline{2-19}
&&\multicolumn{17}{c}{$\delta=0.5$}&\\
\cline{2-19}
&$50$ &0.95&0.94&0.94&&0.89&0.89&0.89&0.89&0.89&&0.98&0.99&0.98&0.93&0.97 &&0.55&\\
&$100$ &1.00&1.00&1.00&&1.00&1.00&1.00&1.00&1.00&&1.00&1.00&1.00&1.00&1.00 &&0.87&\\
&$300$ &1.00&1.00&1.00&&1.00&1.00&1.00&1.00&1.00&&1.00&1.00&1.00&1.00&1.00 &&1.00&\\
&$1000$ &1.00&1.00&1.00&&1.00&1.00&1.00&1.00&1.00&&1.00&1.00&1.00&1.00&1.00 &&1.00&\\ \hline\hline
\end{tabular}
\begin{tablenotes}
• The three values under CF-T are the choices of $\beta$, and those under CF-A and CF-N are the choices of $\kappa_n$ as in (ref) and (ref) respectively.
\end{tablenotes}
\end{threeparttable}
\end{sidewaystable}
{4pt}
\begin{sidewaystable}[!htbp]
\begin{threeparttable}
\caption{Additional results on rejection rates of rank tests for the model (ref) with $r=3$, at $\alpha=5\%$\tnote{\dag}}
\begin{tabular}{cccccccccccccccccccc}
\hline\hline
&\multirow{2}{*}{\makecell{Sample\\ size}} &\multicolumn{3}{c}{CF-T}&&\multicolumn{5}{c}{CF-A}&&\multicolumn{5}{c}{CF-N}&&\multirow{2}{*}{KP-M}&\\
\cline{3-5}\cline{7-11}\cline{13-17}
&&$\alpha/5$&$\alpha/25$&$\alpha/30$ &&$1.5n^{-1/5}$&$1.5n^{-1/4}$&$1.5n^{-1/3}$&$n^{-2/5}$&$1.5n^{-2/5}$ &&$1.5n^{-1/5}$&$1.5n^{-1/4}$&$1.5n^{-1/3}$&$n^{-2/5}$&$1.5n^{-2/5}$ & & &\\
\cline{2-19}
&&\multicolumn{17}{c}{$\delta=0$}&\\
\cline{2-19}
&$50$ &0.08&0.10&0.10&&0.08&0.08&0.06&0.02&0.05&&0.14&0.14&0.14&0.10&0.14&&0.01&\\
&$100$ &0.05&0.07&0.07&&0.06&0.06&0.06&0.01&0.04&&0.13&0.13&0.12&0.07&0.10&&0.01&\\
&$300$ &0.04&0.05&0.05&&0.05&0.05&0.05&0.01&0.03&&0.10&0.09&0.07&0.04&0.06&&0.01&\\
&$1000$ &0.04&0.05&0.05&&0.06&0.06&0.05&0.01&0.05&&0.10&0.08&0.06&0.04&0.05&&0.00&\\
\cline{2-19}
&&\multicolumn{17}{c}{$\delta=0.1$}&\\
\cline{2-19}
&$50$ &0.10&0.13&0.13&&0.11&0.17&0.09&0.03&0.07&&0.18&0.18&0.18&0.13&0.17 &&0.01&\\
&$100$ &0.10&0.13&0.14&&0.13&0.13&0.11&0.04&0.07&&0.21&0.21&0.20&0.11&0.17 &&0.02&\\
&$300$ &0.21&0.28&0.28&&0.32&0.32&0.28&0.10&0.16&&0.41&0.39&0.35&0.23&0.31 &&0.09&\\
&$1000$ &0.58&0.68&0.68&&0.88&0.82&0.76&0.54&0.58&&0.86&0.84&0.81&0.68&0.77 &&0.54&\\
\cline{2-19}
&&\multicolumn{17}{c}{$\delta=0.3$}&\\
\cline{2-19}
&$50$ &0.40&0.45&0.45&&0.47&0.44&0.35&0.23&0.28&&0.57&0.57&0.56&0.44&0.54&&0.12&\\
&$100$ &0.56&0.65&0.65&&0.75&0.72&0.58&0.49&0.51&&0.81&0.80&0.79&0.65&0.76&&0.43&\\
&$300$ &0.95&0.96&0.96&&0.99&0.99&0.96&0.96&0.96&&1.00&1.00&0.99&0.97&0.99&&0.96&\\
&$1000$ &1.00&1.00&1.00&&1.00&1.00&1.00&1.00&1.00&&1.00&1.00&1.00&1.00&1.00&&1.00&\\
\cline{2-19}
&&\multicolumn{17}{c}{$\delta=0.5$}&\\
\cline{2-19}
&$50$ &0.74&0.78&0.79&&0.80&0.74&0.65&0.63&0.63&&0.89&0.89&0.88&0.77&0.87&&0.37&\\
&$100$ &0.91&0.93&0.93&&0.96&0.94&0.92&0.91&0.91&&0.99&0.99&0.99&0.95&0.98&&0.79&\\
&$300$ &1.00&1.00&1.00&&1.00&1.00&1.00&1.00&1.00&&1.00&1.00&1.00&1.00&1.00&&1.00&\\
&$1000$ &1.00&1.00&1.00&&1.00&1.00&1.00&1.00&1.00&&1.00&1.00&1.00&1.00&1.00&&1.00&\\
\hline\hline
\end{tabular}
\begin{tablenotes}
• The three values under CF-T are the choices of $\beta$, and those under CF-A and CF-N are the choices of $\kappa_n$ as in (ref) and (ref) respectively.
\end{tablenotes}
\end{threeparttable}
\end{sidewaystable}
Table (ref) summarizes the simulation results for tuning parameters in the middle range of the choices, while Tables (ref) and (ref) collect results for the remaining choices. For the case of $r=2$ (so that $\mathrm{rank}(\Pi_0)=r$ under $\mathrm H_0$), the performance of CF-A and CF-T is comparable with that of KP-M especially when the sample size is large, though CF-T exhibits more size distortion than KP-M for $n=50$ and CF-N appears to be somewhat sensitive to the choice of $\kappa_n$. For the case of $r=3$ (so that $\mathrm{rank}(\Pi_0)<r$ under $\mathrm H_0$), KP-M is markedly under-sized even in large samples, while its local power is uniformly dominated by our three tests, across all the choices of the tuning parameters, sample sizes, and the local parameter $\delta$. With regards to comparisons among our three tests, there are also some persistent patterns. First, CF-N overall tends to be the most over-sized especially in small samples, and the most sensitive to the choice of the tuning parameters. Second, between CF-A and CF-T, one does not seem to dominate the other. The former appears to perform better overall in terms of size control and local power in small samples, though the differences become smaller as the sample size increases. The latter, on the other hand, seems to be the least sensitive to the choice of the tuning parameters especially in the irregular case when $r=3$, as desired. Thus, it seems sensible to employ CF-A in small samples and CF-T instead in large samples.
{\setlength\intextsep{0pt}
\begin{figure}[!h]
\begin{subfigure}[b]{0.48\textwidth}
\resizebox{\linewidth}{!}{
\pgfplotsset{title style={at={(0.5,0.91)}}}
\begin{tikzpicture}
\begin{axis}[
ybar,
height=10cm,
width=13cm,
enlarge y limits=false,
axis lines*=left,
ymin=0,
ymax=100,
title={$d=1$},
legend style={at={(0.2,0.9)},
anchor=north,legend columns=-1,
/tikz/every even column/.append style={column sep=0.5cm}
},
ylabel={Coverge percentages},
symbolic x coords={0,1,2,3,4,5,6},
xtick=data,
nodes near coords,
every node near coord/.append style={
anchor=mid west,
rotate=70
}
]
\addplot [fill=RoyalBlue4] coordinates {(0,0) (1,0) (2,0)
(3,0) (4,0) (5,10.80)(6,89.20)};
\addplot [fill=LightGrey] coordinates {(0,0) (1,0) (2,0)
(3,0) (4,0) (5,10.64) (6,89.36)};
\legend{CF-A,KP}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.48\textwidth}
\resizebox{\linewidth}{!}{
\pgfplotsset{title style={at={(0.5,0.91)}}}
\begin{tikzpicture}
\begin{axis}[
ybar,
height=10cm,
width=13cm,
enlarge y limits=false,
axis lines*=left,
ymin=0,
ymax=100,
title={$d=2$},
legend style={at={(0.2,0.9)},
anchor=north,legend columns=-1,
/tikz/every even column/.append style={column sep=0.5cm}
},
ylabel={Coverge percentages},
symbolic x coords={0,1,2,3,4,5,6},
xtick=data,
nodes near coords,
every node near coord/.append style={
anchor=mid west,
rotate=70
}
]
\addplot [fill=RoyalBlue4] coordinates {(0,0) (1,0) (2,0)
(3,0) (4,3.36) (5,8.78)(6,87.86)};
\addplot [fill=LightGrey] coordinates {(0,0) (1,0) (2,0)
(3,0) (4,3.44) (5,28.48) (6,68.08)};
\legend{CF-A,KP}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.48\textwidth}
\resizebox{\linewidth}{!}{
\pgfplotsset{title style={at={(0.5,0.91)}}}
\begin{tikzpicture}
\begin{axis}[
ybar,
height=10cm,
width=13cm,
enlarge y limits=false,
axis lines*=left,
ymin=0,
ymax=100,
title={$d=3$},
legend style={at={(0.2,0.9)},
anchor=north,legend columns=-1,
/tikz/every even column/.append style={column sep=0.5cm}
},
ylabel={Coverge percentages},
symbolic x coords={0,1,2,3,4,5,6},
xtick=data,
nodes near coords,
every node near coord/.append style={
anchor=mid west,
rotate=70
}
]
\addplot [fill=RoyalBlue4] coordinates {(0,0) (1,0) (2,0)
(3,1.54) (4,3.26) (5,12.08)(6,83.12)};
\addplot [fill=LightGrey] coordinates {(0,0) (1,0) (2,0)
(3,1.34) (4,16.22) (5,37.94) (6,44.50)};
\legend{CF-A,KP}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.48\textwidth}
\resizebox{\linewidth}{!}{
\pgfplotsset{title style={at={(0.5,0.91)}}}
\begin{tikzpicture}
\begin{axis}[
ybar,
height=10cm,
width=13cm,
enlarge y limits=false,
axis lines*=left,
ymin=0,
ymax=100,
title={$d=4$},
legend style={at={(0.2,0.9)},
anchor=north,legend columns=-1,
/tikz/every even column/.append style={column sep=0.5cm}
},
ylabel={Coverge percentages},
symbolic x coords={0,1,2,3,4,5,6},
xtick=data,
nodes near coords,
every node near coord/.append style={
anchor=mid west,
rotate=70
}
]
\addplot [fill=RoyalBlue4] coordinates {(0,0) (1,0) (2,0.72)
(3,1.80) (4,4.14) (5,14.92)(6,78.42)};
\addplot [fill=LightGrey] coordinates {(0,0) (1,0) (2,0.66)
(3,8.84) (4,32.48) (5,32.22) (6,25.80)};
\legend{CF-A,KP}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.48\textwidth}
\resizebox{\linewidth}{!}{
\pgfplotsset{title style={at={(0.5,0.91)}}}
\begin{tikzpicture}
\begin{axis}[
ybar,
height=10cm,
width=13cm,
enlarge y limits=false,
axis lines*=left,
ymin=0,
ymax=100,
title={$d=5$},
legend style={at={(0.2,0.9)},
anchor=north,legend columns=-1,
/tikz/every even column/.append style={column sep=0.5cm}
},
ylabel={Coverge percentages},
symbolic x coords={0,1,2,3,4,5,6},
xtick=data,
nodes near coords,
every node near coord/.append style={
anchor=mid west,
rotate=70
}
]
\addplot [fill=RoyalBlue4] coordinates {(0,0) (1,0.38) (2,1.84)
(3,1.74) (4,5.68) (5,20.06)(6,70.30)};
\addplot [fill=LightGrey] coordinates {(0,0) (1,0.36) (2,5.22)
(3,25.80) (4,35.08) (5,21.08) (6,12.46)};
\legend{CF-A,KP}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.48\textwidth}
\resizebox{\linewidth}{!}{
\pgfplotsset{title style={at={(0.5,0.91)}}}
\begin{tikzpicture}
\begin{axis}[
ybar,
height=10cm,
width=13cm,
enlarge y limits=false,
axis lines*=left,
ymin=0,
ymax=100,
title={$d=6$},
legend style={at={(0.2,0.9)},
anchor=north,legend columns=-1,
/tikz/every even column/.append style={column sep=0.5cm}
},
ylabel={Coverge percentages},
symbolic x coords={0,1,2,3,4,5,6},
xtick=data,
nodes near coords,
every node near coord/.append style={
anchor=mid west,
rotate=70
}
]
\addplot [fill=RoyalBlue4] coordinates {(0,0) (1,2.16) (2,2.32)
(3,3.12) (4,7.84) (5,24.12)(6,60.44)};
\addplot [fill=LightGrey] coordinates {(0,0) (1,3.36) (2,20.42)
(3,34.04) (4,26.84) (5,9.88) (6,5.46)};
\legend{CF-A,KP}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\caption{The rank estimation: $\mathrm{rank}(\Pi_0)=6$, $\alpha = 5\%$ and $\delta=0.1$}
\end{figure}
}
{\setlength\intextsep{0pt}
\begin{figure}[!h]
\begin{subfigure}[b]{0.48\textwidth}
\resizebox{\linewidth}{!}{
\pgfplotsset{title style={at={(0.5,0.91)}}}
\begin{tikzpicture}
\begin{axis}[
ybar,
height=10cm,
width=13cm,
enlarge y limits=false,
axis lines*=left,
ymin=0,
ymax=100,
title={$d=1$},
legend style={at={(0.2,0.9)},
anchor=north,legend columns=-1,
/tikz/every even column/.append style={column sep=0.5cm}
},
ylabel={Coverge percentages},
symbolic x coords={0,1,2,3,4,5,6},
xtick=data,
nodes near coords,
every node near coord/.append style={
anchor=mid west,
rotate=70
}
]
\addplot [fill=RoyalBlue4] coordinates {(0,0) (1,0) (2,0)
(3,0) (4,0) (5,3.28)(6,96.72)};
\addplot [fill=LightGrey] coordinates {(0,0) (1,0) (2,0)
(3,0) (4,0) (5,3.24) (6,96.76)};
\legend{CF-A,KP}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.48\textwidth}
\resizebox{\linewidth}{!}{
\pgfplotsset{title style={at={(0.5,0.91)}}}
\begin{tikzpicture}
\begin{axis}[
ybar,
height=10cm,
width=13cm,
enlarge y limits=false,
axis lines*=left,
ymin=0,
ymax=100,
title={$d=2$},
legend style={at={(0.2,0.9)},
anchor=north,legend columns=-1,
/tikz/every even column/.append style={column sep=0.5cm}
},
ylabel={Coverge percentages},
symbolic x coords={0,1,2,3,4,5,6},
xtick=data,
nodes near coords,
every node near coord/.append style={
anchor=mid west,
rotate=70
}
]
\addplot [fill=RoyalBlue4] coordinates {(0,0) (1,0) (2,0)
(3,0) (4,0.34) (5,3.36)(6,96.30)};
\addplot [fill=LightGrey] coordinates {(0,0) (1,0) (2,0)
(3,0) (4,0.42) (5,11.56) (6,88.02)};
\legend{CF-A,KP}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.48\textwidth}
\resizebox{\linewidth}{!}{
\pgfplotsset{title style={at={(0.5,0.91)}}}
\begin{tikzpicture}
\begin{axis}[
ybar,
height=10cm,
width=13cm,
enlarge y limits=false,
axis lines*=left,
ymin=0,
ymax=100,
title={$d=3$},
legend style={at={(0.2,0.9)},
anchor=north,legend columns=-1,
/tikz/every even column/.append style={column sep=0.5cm}
},
ylabel={Coverge percentages},
symbolic x coords={0,1,2,3,4,5,6},
xtick=data,
nodes near coords,
every node near coord/.append style={
anchor=mid west,
rotate=70
}
]
\addplot [fill=RoyalBlue4] coordinates {(0,0) (1,0) (2,0)
(3,0.16) (4,0.66) (5,4.10)(6,95.08)};
\addplot [fill=LightGrey] coordinates {(0,0) (1,0) (2,0)
(3,0.12) (4,2.74) (5,23.26) (6,73.88)};
\legend{CF-A,KP}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.48\textwidth}
\resizebox{\linewidth}{!}{
\pgfplotsset{title style={at={(0.5,0.91)}}}
\begin{tikzpicture}
\begin{axis}[
ybar,
height=10cm,
width=13cm,
enlarge y limits=false,
axis lines*=left,
ymin=0,
ymax=100,
title={$d=4$},
legend style={at={(0.2,0.9)},
anchor=north,legend columns=-1,
/tikz/every even column/.append style={column sep=0.5cm}
},
ylabel={Coverge percentages},
symbolic x coords={0,1,2,3,4,5,6},
xtick=data,
nodes near coords,
every node near coord/.append style={
anchor=mid west,
rotate=70
}
]
\addplot [fill=RoyalBlue4] coordinates {(0,0) (1,0) (2,0.08)
(3,0.20) (4,0.88) (5,5.92)(6,92.92)};
\addplot [fill=LightGrey] coordinates {(0,0) (1,0) (2,0.08)
(3,0.88) (4,10.40) (5,30.94) (6,57.70)};
\legend{CF-A,KP}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.48\textwidth}
\resizebox{\linewidth}{!}{
\pgfplotsset{title style={at={(0.5,0.91)}}}
\begin{tikzpicture}
\begin{axis}[
ybar,
height=10cm,
width=13cm,
enlarge y limits=false,
axis lines*=left,
ymin=0,
ymax=100,
title={$d=5$},
legend style={at={(0.2,0.9)},
anchor=north,legend columns=-1,
/tikz/every even column/.append style={column sep=0.5cm}
},
ylabel={Coverge percentages},
symbolic x coords={0,1,2,3,4,5,6},
xtick=data,
nodes near coords,
every node near coord/.append style={
anchor=mid west,
rotate=70
}
]
\addplot [fill=RoyalBlue4] coordinates {(0,0) (1,0.04) (2,0.08)
(3,0.42) (4,1.40) (5,8.98)(6,89.08)};
\addplot [fill=LightGrey] coordinates {(0,0) (1,0.02) (2,0.38)
(3,4.56) (4,21.56) (5,33.10) (6,40.38)};
\legend{CF-A,KP}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.48\textwidth}
\resizebox{\linewidth}{!}{
\pgfplotsset{title style={at={(0.5,0.91)}}}
\begin{tikzpicture}
\begin{axis}[
ybar,
height=10cm,
width=13cm,
enlarge y limits=false,
axis lines*=left,
ymin=0,
ymax=100,
title={$d=6$},
legend style={at={(0.2,0.9)},
anchor=north,legend columns=-1,
/tikz/every even column/.append style={column sep=0.5cm}
},
ylabel={Coverge percentages},
symbolic x coords={0,1,2,3,4,5,6},
xtick=data,
nodes near coords,
every node near coord/.append style={
anchor=mid west,
rotate=70
}
]
\addplot [fill=RoyalBlue4] coordinates {(0,0) (1,0.08) (2,0.28)
(3,0.70) (4,2.00) (5,12.78)(6,84.16)};
\addplot [fill=LightGrey] coordinates {(0,0) (1,0.10) (2,2.02)
(3,14.34) (4,30.32) (5,27.92) (6,25.30 )};
\legend{CF-A,KP}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\caption{The rank estimation: $\mathrm{rank}(\Pi_0)=6$, $\alpha = 5\%$ and $\delta=0.12$}
\end{figure}
}
We now compare with Kleibergen_Paap2006rank in terms of estimation by making use of the same data generating process as specified by (ref) and (ref) with $\delta=0.1$ and $0.12$ so that $\mathrm{rank}(\Pi_0)=6$ (i.e., full rank) in both cases for all $d=1,\ldots, 6$. Our estimation is based on the analytic derivative estimator (ref) with $\hat r_n$ given by (ref) and $\kappa_n=n^{-1/4}$ -- the results for $\kappa_n=n^{-1/3}$ are similar and available upon request. In each configuration, we depict the empirical distributions of the estimators based on 5,000 simulations, 500 bootstrap repetitions for each simulation, and $\alpha=5\%$. As shown by Figures (ref) and (ref), our rank estimators, labelled CF-A, pick up the truth with probabilities higher than the KP estimators, uniformly over $d\in\{2,\ldots, 6\}$ and $\delta\in\{0.1,0.12\}$; when $d=1$, the two sets of estimators are very similar. Note that the empirical probabilities of $\hat r_n=r_0$ are lower in Figure (ref) than in Figure (ref) because $\Pi_0$ is closer to a lower rank matrix (due to a smaller value of $\delta$), and in each figure, the probabilities for both sets of estimators decrease as $\Pi_0$ becomes more degenerate (as $d$ increases). There are two additional interesting persistent patterns. First, the distributions of the KP estimators are more spread out and tend to underestimate the true rank, especially when $d$ is large, i.e., when $\Pi_0$ is local to a matrix whose rank is small. This is in accord with the trivial power of the KP test in this scenario -- see Figure (ref). Second, the probability of our rank estimators equal to the truth can exceed that of the KP rank estimator by as high as 57.84%, and in 5 out of the 12 data generating processes considered, the probabilities of our rank estimator covering the truth are at least 48.70% higher. Once again, this happens especially when $\Pi_0$ is local to a matrix whose rank is small. These observations suggest that our estimators are more robust to local-to-degeneracy.
\section{Saliency Analysis in Matching Models}
In this section, we study a one-to-one, bipartite matching model with transferable utility, where a central question is how many attributes are statistically relevant for the sorting of agents DupuyGalichon2014Personality,CiscatoGalichonGousse2015Like. As shall be seen shortly, this question can be answered by appealing to our framework developed previously. Following the literature, we shall call the two sets of agents men and women, though the theory obviously extends under the general setup.
\subsection{The Model Setup and Saliency Analysis}
Let $X\in\mathcal X\subset\mathbf R^m$ and $Y\in\mathcal Y\subset\mathbf R^k$ be vectors of attributes of men and women respectively, with $P_0$ and $Q_0$ the probability distributions of $X$ and $Y$ respectively. A matching is then characterized by a probability distribution $\pi$ on $\mathcal X\times\mathcal Y$ such that its density $f_\pi(x,y)$ describes the probability of occurrence of a couple with attributes $(x,y)$. Since we only consider matched couples and matching is one-to-one, $\pi$ must have marginals $P_0$ and $Q_0$. A defining feature of the transferrable utility framework is that matched couples behave unitarily, i.e., there is a single surplus function $s: \mathcal X\times\mathcal Y\to\mathbf R$ generated by the matching, and how the surplus is shared between the spouses is endogenous. A final ingredient crucial to the matching game is the equilibrium concept. As standard in the literature, we employ the notion of stability GaleShapley1962College, and call a matching stable if (i) no matched individual would rather be single and (ii) no pair of individuals would {\it both} like being matched together better than their current situation. It is well known that stability (a game theoretical concept) and surplus maximization (a social planner's problem) are equivalent ShapleyShubik1971Assignment,ChiapporiMcCannNesheim2010Hedonic. Consequently, the matching $\pi_0$ in equilibrium can be characterized by the centralized problem:
\begin{align}
\max_{\pi\in\bm\Pi(P_0,Q_0)} E_\pi[s(X,Y)] ,
\end{align}
where $\bm\Pi(P_0,Q_0)$ is the family of distributions on $\mathcal X\times\mathcal Y$ with marginals $P_0$ and $Q_0$.
Without further appropriate modelling, the optimal transport problem (ref) implies pure matching under regularity conditions Becker1973Marriage,ChiapporiMcCannNesheim2010Hedonic, i.e., a certain type of men is for sure going to be matched with a certain type of women. One empirical strategy to reconcile such unrealistic predictions with data is to incorporate unobserved heterogeneity into the surplus function. Following ChooSiow2006Marries and ChiapporiSalanieWeiss2017Partner, we assume that
\begin{align}
s(x,y)=\Phi(x,y)+\epsilon_m(y)+\epsilon_w(x) ,
\end{align}
where $\Phi(x,y)$ is the systematic part of the surplus, and $\epsilon_m(y)$ and $\epsilon_w(x)$ are unobserved random shocks. Note that $\epsilon_m(y)$ and $\epsilon_w(x)$ enter the surplus function additively and separably, which is by no means a haphazard restriction: it makes an otherwise extremely difficult problem more tractable ChiapporiSalanie2016Matching,Chiappori2017Matching. Nonparametric identification of both $\Phi$ and the error distributions, however, remains a challenging task. Following Dagsvik2000Aggregation and ChooSiow2006Marries, we thus further assume that the errors follow the type-I extreme value distribution, though we note that such distributional assumption can be completely dispensed with GalichonSalanie2015Cupid. The matching distribution $\pi_0$ can in turn be characterized by
\begin{align}
\max_{\pi\in\bm\Pi(P_0,Q_0)} E_\pi[\Phi(X,Y)]-E_\pi[\log f_\pi(X,Y)] ,
\end{align}
and $\Phi$ is nonparametrically identified GalichonSalanie2015Cupid. For the purpose of estimation, we further assume that, for some $A_0\in\mathbf M^{m\times k}$ and any $(x,y)\in\mathcal X\times\mathcal Y$,
\begin{align}
\Phi(x,y)\equiv\Phi_{A_0}(x,y)=x^\text{\scalebox{0.7}{$\intercal$}} A_0 y ,
\end{align}
where $A_0$ is called the affinity matrix. Such a parametric specification has also been employed by GalichonSalanie2010Tradeoff,GalichonSalanie2015Cupid and DupuyGalichon2014Personality.
Heuristically, the $(i,j)$th entry $a_{ij}$ of $A_0$ measures the strength of mutual attractiveness between attributes $x_i$ and $y_j$. The rank of $A_0$ provides valuable information on the number of dimensions on which sorting occurs, and helps construct indices of mutual attractiveness DupuyGalichon2014Personality,DupuyGalichon2015Canonical. Following DupuyGalichon2014Personality and GalichonSalanie2015Cupid, we estimate $A_0$ by matching moments:
\begin{align}
E_{\pi(A_0,P_0,Q_0)}[XY^\text{\scalebox{0.7}{$\intercal$}}]= E[XY^\text{\scalebox{0.7}{$\intercal$}}] ,
\end{align}
where $\pi_0\equiv \pi(A_0,P_0,Q_0)$ is the matching distribution in equilibrium. By Lemma (ref), if $X$ and $Y$ are finitely discrete-valued with probability mass functions $p_0$ and $q_0$, then equation (ref) defines not only a unique $A_0$, but also an implicit map $(p_0,q_0,E[XY^\text{\scalebox{0.7}{$\intercal$}}])\mapsto A(p_0,q_0,E[XY^\text{\scalebox{0.7}{$\intercal$}}])\equiv A_0$ which is differentiable. This has two immediate implications. First, the estimator $\hat A_n$ defined by the sample analog of (ref), i.e.,
\begin{align}
E_{\pi(\hat A_n,\hat p_n,\hat q_n)}[XY^\text{\scalebox{0.7}{$\intercal$}}]= \frac{1}{n}\sum_{i=1}^{n}X_iY_i^\text{\scalebox{0.7}{$\intercal$}} ,
\end{align}
where $\hat p_n$ and $\hat q_n$ are sample analogs of $p_0$ and $q_0$ respectively, is asymptotically normal. Second, the bootstrap estimator $\hat A_n^*$ defined by the bootstrap analog of (ref), i.e.,
\begin{align}
E_{\pi(\hat A_n^*,\hat p_n^*,\hat q_n^*)}[XY^\text{\scalebox{0.7}{$\intercal$}}]= \frac{1}{n}\sum_{i=1}^{n}X_i^*Y_i^{*\text{\scalebox{0.7}{$\intercal$}}} ,
\end{align}
where $\hat p_n^*$ and $\hat q_n^*$ are bootstrap analogs of $\hat p_n$ and $\hat q_n$ respectively, is consistent in estimating the asymptotic distribution of $\hat A_n$. We have thus verified the main assumptions in order to apply our framework. We note in passing that it appears challenging to verify Assumption (ref) when $X$ and $Y$ are continuous, and we believe it should be based on arguments different from those above.
Alternatively, DupuyGalichon2014Personality estimate the rank of $A_0$ by employing the test of Kleibergen_Paap2006rank, which they call the saliency analysis. There are two motivations of using our inferential procedure. First, as argued previously, the KP test is designed for the more restrictive setup (ref) and can be invalid and/or conservative for the hypotheses in (ref). Consequently, estimation of $\mathrm{rank}(A_0)$ by sequentially conducting the KP tests may be less accurate. Second, the KP test relies on an estimator of the asymptotic variance of $\hat A_n$ which appears to be somewhat complicated -- see the formula (B18) in DupuyGalichon2014Personality, while one generic merit of bootstrap inference is to avoid analytic complications by repetitive resampling HorowitzBoot.
\subsection{Data and Empirical Results}
We use the same data source as DupuyGalichon2014Personality, i.e., the 1993-2002 waves of the DNB Household Survey, to estimate preferences in the marriage market in Dutch. The panel contains rich information about individual attributes such as demographic variables (e.g., education), anthropometry parameters (e.g., height and weight), personality traits (e.g., emotional stability, extraversion, conscientiousness, agreeableness, autonomy) and risk attitude -- see Nyhus1996VSB for more detailed descriptions of the data. In order to apply our framework, we have discretized the variables in the following way: (i) BMI\footnote{The body mass index (BMI) is defined as the body mass divided by the square of the body height, which provides a simple numeric measure of a person's thinness.} is converted into a trinary variable according to the international BMI classification, i.e., BMI is set to be 1 if BMI $<$18.50, 2 if $18.50\le$ BMI $< 24.99$, and 3 if BMI $\ge 24.99$; (ii) Five personal traits variables and risk aversion are also converted into trinary data by taking the value 1 if they are below the corresponding 25% quantiles, 2 if they are between the 25% and the 75% quantiles, and 3 if they are strictly larger than the 75% quantiles; (iii) Education remains unchanged since it is discrete as it is. We make use of the same sample as DupuyGalichon2014Personality which has 1158 couples, but only include subsets of the 10 attribute variables that they considered to reduce the computational burden -- see Table (ref). Following DupuyGalichon2014Personality still, we demean and standardize the data beforehand, and then compute the optimal matching distribution by the iterative projection fitting procedure Ruschendorf1995Iterative.
\begin{table}[!h]
\caption{Model specifications}
\begin{center}
\begin{tabular}{ccl}
\hline\hline
Model && Attributes included\\
\hline
(1) && Education, BMI, Risk aversion\\
(2) && Education, BMI, Risk aversion, Conscientiousness\\
(3) && Education, BMI, Risk aversion, Extraversion\\
(4) && Education, BMI, Risk aversion, Agreeableness\\
(5) && Education, BMI, Risk aversion, Emotional stability\\
(6) && Education, BMI, Risk aversion, Autonomy\\
(7) && Education, BMI, Risk aversion, Conscientiousness, Extraversion\\
(8) && Education, BMI, Risk aversion, Conscientiousness, Autonomy\\
(9) && Education, BMI, Risk aversion, Extraversion, Autonomy\\
\hline\hline
\end{tabular}
\end{center}
\end{table}
{2pt}
\begin{table}[!h]
\caption{Empirical results}
\begin{center}
\begin{threeparttable}
\begin{tabular}{ccccccccccccccccc}
\hline\hline
& & & \multicolumn{13}{c}{The $p$-values for full rank tests\tnote{\dag}}\\
\cmidrule{2-16}
&\multirow{2}{*}{Model} &\multirow{2}{*}{\makecell{Maximum\\ rank}} & \multicolumn{3}{c}{CF-T}& & \multicolumn{3}{c}{CF-A}& &\multicolumn{3}{c}{CF-N} & &\multirow{2}{*}{KP-M\tnote{\ddag}}&\\
\cmidrule{4-6}\cmidrule{8-10} \cmidrule{12-14}
&&&$\alpha/10$&$\alpha/15$&$\alpha/20$ & &$n^{-1/5}$&$n^{-1/4}$&$n^{-1/3}$ & &$n^{-1/5}$& $n^{-1/4}$ & $n^{-1/3}$&&&\\
\cmidrule{2-16}
&$(1)$ & 3 &0.00 &0.00& 0.00 &&0.00& 0.00 & 0.00 &&0.00&0.00 &0.00 &&0.00 &\\
&$(2)$ & 4 &0.01 &0.01& 0.01 &&0.00& 0.00 &0.01 &&0.00&0.00 &0.00 &&0.03 &\\
&$(3)$ & 4 &0.04 &0.04& 0.04 &&0.01& 0.04 &0.18 &&0.01&0.02 &0.04 &&0.25 &\\
&$(4)$ & 4 &0.88 &0.88& 0.88 &&0.86& 0.88 &0.92 &&0.85&0.86 &0.87 &&0.94 &\\
&$(5)$ & 4 &0.23 &0.08& 0.08 &&0.03& 0.08 &0.23 &&0.02&0.03 &0.06 &&0.35 &\\
&$(6)$ & 4 &0.01 &0.01& 0.01 &&0.00& 0.00 &0.01 &&0.00&0.00 &0.00 &&0.03 &\\
&$(7)$ & 5 &0.02 &0.02& 0.02 &&0.00& 0.02 &0.14 &&0.00&0.00 &0.01 &&0.19 &\\
&$(8)$ & 5 &0.00 &0.00& 0.00 &&0.00& 0.00 &0.01 &&0.00&0.00 &0.00 &&0.03 &\\
&$(9)$ & 5 &0.00 &0.00& 0.00 &&0.00& 0.00 &0.03 &&0.00&0.00 &0.00 &&0.22 &\\
\cmidrule{2-16}
&& & \multicolumn{13}{c}{Estimates of the true rank ($\alpha=5\%$)}&\\
\cmidrule{2-16}
&$(1)$ & 3 &3 &3 & 3 && 3& 3&3 && 3& 3&3 &&3&\\
&$(2)$ & 4 &4 &4 & 4 && 4& 4&4 && 4& 4&4 &&4&\\
&$(3)$ & 4 &4 &4 & 4 && 4& 4&3 && 4& 4&4 &&3&\\
&$(4)$ & 4 &3 &3 & 3 && 3& 3&3 && 3& 3&3 &&3&\\
&$(5)$ & 4 &3 &3 & 3 && 4& 3&3 && 4& 4&3 &&3&\\
&$(6)$ & 4 &4 &4 & 4 && 4& 4&4 && 4& 4&4 &&4&\\
&$(7)$ & 5 &5 &5 & 5 && 5& 5&4 && 5& 5&5 &&3&\\
&$(8)$ & 5 &5 &5 & 5 && 5& 5&5 && 5& 5&5 &&5&\\
&$(9)$ & 5 &5 &5 & 5 && 5& 5&5 && 5& 5&5 &&3&\\
\hline\hline
\end{tabular}
\begin{tablenotes}
• The three values under CF-T are the choices of $\beta$, and those under CF-A and CF-N are the choices of $\kappa_n$ as in (ref) and (ref) respectively.
• The $p$-value for KP-M is given by the smallest significance level such that the null hypothesis is rejected, which is equal to the maximum $p$-value of all Kleibergen_Paap2006rank's tests implemented by the multiple testing method.
\end{tablenotes}
\end{threeparttable}
\end{center}
\end{table}
For each model specification, we study two problems: testing singularity of the corresponding affinity matrix and estimating its true rank. In carrying out our inferential procedures, we estimate the derivative through either (ref) or (ref), for which we choose the tuning parameter $\kappa_n\in\{n^{-1/5},n^{-1/4},n^{-1/3}\}$. The corresponding results are labelled as CF-A and CF-N respectively. We also implement the two-step procedure with $\beta\in\{\alpha/10,\alpha/15,\alpha/20\}$, labelled as CF-T. The significance level is $\alpha=5\%$. As shown by Table (ref), our three inferential procedures yield overall consistent results, with the exception of models (3), (5) and (7). For example, for model (3), all our procedures estimate the rank to be 4, except CF-A with $\kappa_n=n^{-1/3}$ which estimates the rank to be $3$. Such discrepancies may be due to the choices of tuning parameters or finite sample variations. Nonetheless, what is comforting to us is that, in the three models, the majority of the 9 estimates point to the same rank. We also note that the $p$-values and estimates of the rank based on CF-T are the same across all three choices of $\beta$, for all model specifications except for model (5).
There are, however, noticeable differences between our results and those obtained by the KP test. First, there are sizable discrepancies between the $p$-values of our tests and those for the KP-M tests, especially for model specifications (3), (5), (7) and (9). Second, in terms of estimation, there are also marked differences. For example, for model (9), our tests unanimously estimate the rank to be 5, while the KP test estimates the rank to be 3. Similar patterns occur for models (3) and (7) for which the KP test provides a smaller rank estimator. Inspecting these differences, it seems that Extraversion is not important for matching in the Dutch marriage market according to the KP results, while our results show that it is important. Overall, we obtain estimates different from those based on Kleibergen_Paap2006rank in 3 out of the 9 model specifications.
\section{Conclusion}
In this paper, we have developed a general framework for conducting inference on the rank of a matrix $\Pi_0$. The problem is of first order importance because we have shown, through an analytic example and simulation evidences, that existing tests may be invalid due to over-rejections when in truth $\mathrm{rank}(\Pi_0)$ is strictly less than the hypothesized value $r$, while their multiple testing versions, though valid, can be substantially conservative. We have then developed a testing procedure that has asymptotic exact size control, is consistent, and meanwhile accommodates the possibility $\mathrm{rank}(\Pi_0)<r$. A two-step test is proposed to mitigate the concern on tuning parameters. We also characterized classes of local perturbations under which our tests have local size control and nontrivial local power. These attractive testing properties in turn lead to more accurate rank estimators. We illustrated the empirical relevance of our results by conducting inference on the rank of an affinity matrix in a two-sided matching model.
We stress that our framework is limited to matrices of fixed dimensions and inapplicable to examples where the dimensions diverge as sample size increases. This is because Assumption (ref) is being violated in these settings, as $\Pi_0$ typically does not admit weakly convergent estimators. While we find extensions allowing varying dimensions important in, for example, many IV problems and high dimensional factor models, a thorough treatment is beyond the scope of this paper and hence left for future study.
\addcontentsline{toc}{section}{References}
\putbib
bibunit\setcounter{page}{1}