EconBase
← Back to paper

Randomization Tests for Equality in Dependence Structure

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

94,207 characters · 8 sections · 1 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Randomization Tests for Equality in Dependence Structure

\sloppy

center[center omitted — 33 chars of source]

We develop a new statistical procedure to test whether the dependence structure is identical between two groups. Rather than relying on a single index such as Pearson's correlation coefficient or Kendall's $\tau$, we consider the entire dependence structure by investigating the dependence functions (copulas). The critical values are obtained by a modified randomization procedure designed to exploit asymptotic group invariance conditions. Implementation of the test is intuitive and simple, and does not require any specification of a tuning parameter or weight function. At the same time, the test exhibits excellent finite sample performance, with the null rejection rates almost equal to the nominal level even when the sample size is extremely small. Two empirical applications concerning the dependence between income and consumption, and the Brexit effect on European financial market integration are provided. \\ Keywords: copula; randomization test; permutation test

\doublespacing

Introduction

As the most fundamental measure of dependence, Pearson's coefficient of correlation has been widely used for centuries in numerous empirical studies concerning the dependence between variables. Despite its lasting popularity, however, Pearson's coefficient of correlation certainly has some limitations in characterizing dependence structures because it only captures pairwise and linear dependence. When the variables of interest are Gaussian where the linearity is implied, the correlation coefficient serves as the best measure of dependence in the sense that the dependence structure is fully described by the correlation. Nevertheless, it is not always reasonable to presume that the underlying dependence structure is linear and in fact, many economic and financial data frequently exhibit nonlinear relationships.

Alternatively, we may aim to provide a meaningful description of the entire dependence structure rather than summarize it using a single index, particularly if nonlinear features of the variables such as asymmetric dependence or tail dependence are of interest. In this context, dependence functions, also known as copulas, have proved to be useful in the studies of dependence structures. Suppose that $X_{1},...,X_{d}$ are random variables with joint distribution function $H$ and univariate margins $F_{1},...,F_{d}$. Sklar's theorem (1959) ensures the existence of a joint distribution function $C:[0,1]^{d}\rightarrow \lbrack 0,1]$ that has uniform margins and satisfies

equation[equation omitted — 76 chars of source]

for all $(x_{1},...,x_{d})\in \mathbb{R}^{d}$. The function $C$ is called the copula associated with $H$. This result suggests that any joint distribution can be decomposed into two parts, the univariate marginal distributions which determine the behavior of individual variables, and the copula function which determines the dependence structure between variables. For the reason, copulas have been major concerns in various applications of dependence modelling.\footnote{For instance, Patton (2006) used time-varying copulas to capture the asymmetric dependence between exchange rates. Zimmer and Trivedi (2006) employed trivariate copulas for an application to family health care demand, and Bonhomme and Robin (2009) explored copulas in the modelling of the\ transitory component in earnings. In finance, Embrechts et al.\ (2002) and Rosenberg and Schuermann (2006) studied risk management using the copula models, while Oh and Patton (2013, 2017) used copulas to model the dependence between stock returns. Copula-based models of serial dependence or heteroskedasticity have been studied by Chen and Fan (2006), Chen et al.\ (2009), Lee and Long (2009), Ibragimov (2009), Smith et al. (2010), Beare (2010), and Beare and Seo (2014, 2015), Loaiza-Maya et al. (2018). See also Creal and Tsay (2015) for an application to panel data.}

In this paper, we use the copulas to propose a statistical procedure for testing homogeneity of dependence structure between different groups. By comparing copulas, our test detects any arbitrary form of dissimilarity in the dependence structure. The test statistic is constructed based on the $L_{p}$ distance between two empirical copulas and critical values are obtained from a novel randomization procedure. Implementation of our test is intuitive and straightforward, and does not require any specification of a tuning parameter or weight function that can be arbitrarily chosen by a researcher. At the same time, the test exhibits substantially excellent performance in finite samples, with the null rejection rates almost equal to the nominal level even when the sample size is extremely small.

In developing our methodology, we adapt the permutation method used to test distributional equality. A modification is required because the null set of copula equality is strictly larger than the null set of distributional equality. When the problem is to test the equality of distributions, classical permutation tests deliver exact size control and those tests are as powerful as standard parametric tests under general conditions (Hoeffding, 1952). However, when the null hypothesis to be tested is larger than distributional equality, the usual permutation tests generally do not control size even when the sample size is large (Romano, 1990; Chung and Romano, 2013). Therefore, we expect that a naive application of a permutation test to the hypothesis of copula equality may fail to deliver valid inference even asymptotically.

To resolve this problem, recent studies on randomization tests focus on the studentization of a test statistic. See, for instance, Neuhaus (1993), Janssen (1997, 1999, 2005), Janssen and Pauls (2003, 2005), Neubert and Brunner (2007), Omelka and Pauly (2012) and Chung and Romano (2016). These modifications provide asymptotic level $\alpha $ tests of relevant null hypotheses, with exact size control under distributional equality. However, this approach is not applicable in our context because our test statistic has more or less complicated form and its sampling distribution involves several Brownian bridges determined by underlying copulas, and the derivatives of those copulas. Henceforth, we take a different approach introducing theorems of conditional convergence in the context of randomization test literature. Although we only focus on the copula equality, our technical ingredients may be useful for handling other null hypotheses where the test statistic is not linear and the delta method is to be invoked.

The remainder of the paper is structured as follows. In Section 2, we explain why the classical permutation method is invalid when used to test copula equality. Our main results are presented in Section 3, where we introduce modified randomization procedures and discuss their asymptotic properties. In Section 4, we report some numerical simulation results. In Section 5, two empirical applications concerning the dependence between income and consumption, and the Brexit effect on European financial market integration are provided. Proofs of lemmas and theorems are collected together in the Appendix.

Application of the permutation method

Suppose that $X^{1},...,X^{n}$ are i.i.d. draws from a distribution $ H_{1}\,$and $Y^{1},...,Y^{m}$ are i.i.d. draws from a distribution $H_{2}$. The two samples are independent. Each distribution is $d$-variate with $d\in\mathbb N$, and we write $X^i=(X_1^i,\ldots,X_d^i)$ and $Y^j=(Y_1^j,\ldots,Y_d^j)$ for $i=1,...,n$ and $j=1,...,m$. We denote the univariate margins of $H_1$ by $F_1,\ldots,F_d$ and the univariate margins of $H_2$ by $G_1,\ldots,G_d$. Let $N$ be the total number of observations (that is, $N=n+m$), and let $W$ be the stacked matrix of the size $N\times d$,

equation[equation omitted — 95 chars of source]

For the case of testing the hypothesis $\mathcal H_0:(H_{1},H_{2})\in \Theta_{00} $ with $\Theta_{00} =\{(H,H')\in\mathbb D^2|H=H'\}$, where $\mathbb D$ is the set of all $d$-variate probability distributions, we may easily construct an exact level $\alpha $ test using a randomization test based on permuting the rows of $W$. Complications arise when the null hypothesis of interest is strictly larger than $ \Theta_{00} $. If this is the case, it is well known that permutation tests generally cannot control the probability of Type I error even asymptotically. Hence, inferences based on a permutation test can be highly misleading (Romano, 1990; Chung and Romano, 2013).

Let $\Phi : \mathbb{D}\rightarrow \ell ^{\infty }([0,1]^{d})$ be the map that sends a cdf $\tilde{H}\in \mathbb{D}$ with margins $\tilde{F} _{1},...,\tilde{F}_{d}$ to $\tilde{H}(\tilde{F}_{1}^{-1},...,\tilde{F} _{d}^{-1})$. For each $k=1,...,d$ and $t\in \lbrack 0,1]$, $\tilde{F} _{k}^{-1}(t)$ is defined to be $\inf \{x|\tilde{F}_{k}(x)\geq t\}$. In words, given a joint distribution function, $\Phi $ is the map which provides the corresponding copula as an output. Now we can formulate our null hypothesis of copula equality as

equation[equation omitted — 142 chars of source]

Since we may have two different distributions $H_{1}$ and $H_{2}$ which share the same copula, our null set $\Theta _{0}$ is strictly larger than $ \Theta_{00} $. This implies that a randomization test based on the permutations of $W$ may possibly lead to a permutation distribution which does not agree with the correct limit of our test statistic.

To illustrate this point, let's consider the following example with two bivariate ($d=2$) cdfs $H_{1}$ and $H_{2}$. We let $H_{1}$ be the distribution of $X_{1}$ and $X_{2}$ where $X_{1}$ is uniformly distributed between zero and one, and $X_{2}$ is defined to be $X_{1}+1$. Now consider another distribution $H_{2}$, that is the distribution of $Y_{1}$ and $Y_{2}$ , where $Y_{2}$ is uniformly distributed between zero and one, and $Y_{1}$ is defined to be $Y_{2}+1$. The joint distributions $H_{1}$ and $H_{2}$ are given by $H_{1}(x_{1},x_{2})=\min (x_{1},x_{2}-1)$ on $(x_{1},x_{2})\in \lbrack 0,1]\times \lbrack 1,2]$ and $H_{2}(x_{1},x_{2})=$ $\min (x_{1}-1,x_{2})$ on $(x_{1},x_{2})\in \lbrack 1,2]\times \lbrack 0,1]$, respectively. See Figure 2.1 for a graphical illustration of the shapes of the distributions. For both $H_{1}$ and $H_{2}$,\ it is easy to verify that the associated copula is $C_{1}(u_{1},u_{2})=C_{2}(u_{1},u_{2})= \min (u_{1},u_{2})$ for $(u_{1},u_{2})\in \lbrack 0,1]^{2}$. We see that $ H_{1}$ and $H_{2}$ are different but their copulas $C_{1}$ and $C_{2}$ are identical. In other words, $(H_{1},H_{2})\in \Theta _{0}$ but $ (H_{1},H_{2})\notin \Theta_{00} $.

figure[figure omitted — 436 chars of source]

Suppose we will use a test statistic $T_{n,m}$ to test copula equality. The limit distribution of $T_{n,m}$ may generally depend on the underlying copulas and we hope to approximate it using the permutation distribution of $T_{n,m}$. Since the permuted samples behave as though drawn from a mixture of $H_1$ and $H_2$, the permutation distribution of $ T_{n,m}$ can be inferred from its unconditional distribution when the samples are drawn from that mixture distribution. This argument is well explained in Chung and Romano (2013), where the authors employed a contiguity argument and coupling construction assuming the linearity of $T_{n,m}$. Although our test statistic in the next section is not linear with respect to the original sample, the validity of our test procedure can be demonstrated by applying results in Chung and Romano (2013) to a simpler infeasible test statistic whose construction depends on the univariate margins being known. Viewing our test in this light, it is natural to ask how the copula of a mixture of $H_1$ and $H_2$ is determined.

Now return to our example and consider a mixture of $H_{1}$ and $H_{2}$ defined by $\bar{H}=\lambda H_{1}+(1-\lambda )H_{2}$ with some weight $\lambda \in \lbrack 0,1]$. The support of the mixture distribution $ \bar{H}$ is displayed in Panel (a) of Figure 2.2. When $\lambda =1/2$ for instance, $\bar{H}$ corresponds to the distribution of a pair of random variables uniformly distributed over the two diagonals in Panel (a) of Figure 2.2. The corresponding copula in this case is,

equation*[equation* omitted — 418 chars of source]

which has the uniform probability mass over the bold lines in Panel (b) of Figure 2.2. We observe from this example that the copula associated with the mixture distribution $\bar{H}$ can be different from $C_{1}$ (or $C_{2}$) even when $C_{1}=C_{2}$.

figure[figure omitted — 446 chars of source]

The discussion in the preceding paragraph should provide some insight into why a permutation test obtained by permuting the rows of $W$ can be invalid. Failure of such a permutation test may be attributed to the fact that (i) the limit distribution of $T_{n,m}$ is determined by the true underlying copulas in general, and (ii) unless $\lambda$ is zero or one, the copula of the mixture distribution is different from $C_{1}$ or $C_{2}$. In our example above, the limit of $ T_{n,m}$ is determined by $C_{1}$ which is equal to $C_{2}$, whereas the limit of the permutation distribution is determined by the copula associated with the mixture distribution with $\lambda$ being determined by the limiting value of $n/(n+m)$. This suggests that estimating the asymptotic null distribution of $T_{n,m}$ by permuting the samples from $H_{1}$ and $H_{2}$ may not be adequate.

A natural solution to this problem is to manipulate the test statistic in such a way that its limiting distribution does not depend on the underlying probability distributions or copulas. When applied to a properly studentized test statistic, a permutation test can deliver asymptotically valid inference in the sense that the rejection probability converges to the nominal level $\alpha$ when the null hypothesis is true. However, this can only be done in limited circumstances where the test statistic is of a simple linear form. For more general applications of the permutation method, it requires more technical development.

To adapt the permutation method for the inference of copula equality, we will instead commence from the observation that copulas can be regarded as the joint distribution of the probability integral transform of each univariate variable. For instance, the copula in (1) is the joint distribution of $d$ random variables, $F_{1}(X_{1}),...,F_{d}(X_{d})$. From this point of view, the problem of testing copula equality is closely related to the problem of testing the equality of probability distributions; assuming that the univariate margins $F_{1},...,F_{d}$ are known, the classical theory of randomization tests applies and an exact $ \alpha $ level test can be constructed by the permutation method. However, the marginal distributions are not known in practice and can only be estimated consistently, and this suggests that we may only rely on asymptotic group invariance conditions in our application of the permutation method. In this respect, our results in the next section complement those of Canay et al. (2017) or Beare and Seo (2017), who investigated the behavior of randomization tests under approximate symmetry conditions, albeit with different formalizations of approximate symmetry.

Test construction

In this section, we explain how to construct valid permutation tests of the hypothesis of equal dependence structure. Two asymptotically valid procedures are proposed. The first is similar to the multiplier technique of R\'emilard and Scaillet (2009). The second is a modification of the first that eliminates the need for partial derivative estimation. Our main results are summarized in Theorem 3.1 and Theorem 3.2. Although we confine our attention to the two-sample problem for simplicity, it is straightforward to extend our results to the more general $k$-sample problem.

We start out by defining the test statistic. As in the previous section, let $C_{1}$ and $ C_{2}$ be the copulas associated with $H_{1}\,$and $H_{2}$ respectively; that is, $ C_{1}=\Phi (H_{1})$ and $C_{2}=\Phi (H_{2})$. Since our null hypothesis is given as in (3), a proper test statistic can be constructed based on the discrepancy between $C_{1}$ and $C_{2}$. Let $\hat{F}_{1,n},...,\hat{F}_{d,n},\hat{G}_{1,m},..,\hat{G}_{d,m}$ be the empirical distributions corresponding to $F_{1},...,F_{d}$, $ G_{1},...,G_{d}$,

equation[equation omitted — 213 chars of source]

and $\hat{C}_{1,n}$, $\hat{C}_{2,m}$ be the empirical copulas computed from the two independent samples as below.

eqnarray*[eqnarray* omitted — 366 chars of source]

Our test statistic is provided by

equation*[equation* omitted — 108 chars of source]

where $\Vert \cdot \Vert _{p}$ is the $L_{p}$ norm with respect to the Lebesgue measure on $[0,1]^{d}$ given $p\in \lbrack 1,\infty ]$.

The empirical copula is the most widely used nonparametric estimator of copulas and its asymptotic results have been well established in the literature. The limit theory of the empirical copula process was firstly developed by Deheuvels (1981a, 1981b) under independence, and extended to nonindependent cases by Gaenssler and Stute (1987) in the Skorokhod space $D([0,1]^{2})$ and later, by Fermanian et al. (2004) in the space $\ell ^{\infty }\left( [0,1]^{2}\right) $. See also Segers (2012), van der Vaart and Wellner (1996, 2007) and Tsukahara (2005) for more results. As in those papers, we henceforth assume that $F_{1},...,F_{d},G_{1},...,G_{d}$ are continuous, and that $C_{1}$ and $C_{2}$ admit continuous partial derivatives on $ [0,1]^{d}$ to ensure the weak convergence of the empirical copula process. In what follows, let $\rightsquigarrow $ denote Hoffmann-J\o rgensen convergence in \(\ell^\infty([0,1]^d)\) as $\min (n,m)$ tends to infinity.

Lemma 3.1. Suppose that $\lambda _{n,m}\equiv n/(n+m)\rightarrow \lambda $ as $\min (n,m)\rightarrow \infty $. Then we have

equation*[equation* omitted — 174 chars of source]

where $\mathbb{C}_{C_{1}}$ and $\mathbb{C}_{C_{2}}$ are the weak limits of the empirical copula processes $\sqrt{n}(\hat{C}_{1,n}-C_{1})$ and $\sqrt{m}( \hat{C}_{2,m}-C_{2})$ respectively. Accordingly, when $C_1$ and $C_2$ are identical,

equation*[equation* omitted — 184 chars of source]

The specific forms $\mathbb{C}_{C_{1}}$ and $\mathbb{C}_{C_{2}}$ are determined by the $C_{1}$-Brownian bridge, $C_{2}$-Brownian bridge, and the partial derivatives of $C_{1}$ and $C_{2}$. Let $\mathbf{u}$ denote the vector of $d$ entries, $(u_{1},...,u_{d})\in \lbrack 0,1]^{d}$ and $\mathbf{u}_{(q)}=(1,1...,u_{q},...,1,1)\in \lbrack 0,1]^{d}$ denote the vector which has $u_{q}$ as its $q$-th entry and $1$ elsewhere. For a copula $C$, define $\mathbb{B}_{C}$ to be a Brownian bridge on $[0,1]^{d}$ with covariance kernel

equation[equation omitted — 204 chars of source]

where $\mathbf{u}\wedge \mathbf{u}^{\prime }$ be the minimum taken componentwise, i.e., $(\mathbf{u}\wedge \mathbf{u}^{\prime })=(u_{1}\wedge u_{1}^{\prime },...,u_{d}\wedge u_{d}^{\prime })$. Then, $\mathbb{C}_{C_{1}}$ and $\mathbb{C}_{C_{2}}$ can be expressed as

equation[equation omitted — 209 chars of source]

with $\partial _{q}C_{l}(\mathbf{u})$ denoting the partial derivative of $ C_{l}(\mathbf{u})$ with respect to the $q$-th argument, for $q=1,...,d$ and $l=1,2$. The asymptotic null distribution of our test statistic $T_{n,m}^{(p)}$ can be obtained by applying the continuous mapping theorem to the weak convergence established in Lemma 3.1 as,

equation[equation omitted — 187 chars of source]

Now we seek to use the permutation method to approximate the limit of our test statistic in (7), under copula equality. Before entering into the main analysis, it is worth noting that when the univariate margins are known, an exact level $\alpha$ test can be delivered by the permutation method. This is because $C_{1}$ is the joint distribution of $ F_{1}(X_{1}),...,F_{d}(X_{d})$ and $C_{2}$ is the joint distribution of $ G_{1}(Y_{1}),...,G_{d}(Y_{d})$. Therefore, given the univariate margins $ F_{1},...,F_{d}$ and $G_{1},...,G_{d}$, copula equality can be reformulated as distributional equality.\footnote{ See Remark 3.2 for conditions under which copula equality is necessary and sufficient for distributional equality.} In the application of the permutation method, however, permutations should be applied to the transformed data

equation[equation omitted — 84 chars of source]

as opposed to (2), properly accounting for the group invariance conditions. Here, the vectors $U^{i}=(F_{1}(X_{1}^{i}),...,F_{d}(X_{d}^{i}))$ and $ V^{j}=(G_{1}(Y_{1}^{j}),...,G_{d}(Y_{d}^{j}))$ are the probability integral transforms of the $i$-th and $j$-th observations of $X$ and $Y$ respectively, for $i=1,...,n$ and $j=1,...,m$.

When the univariate margins are known, $C_{1}$ and $C_{2}$ can be estimated by the empirical distributions computed from $\{U^{i}\}_{i=1}^{n}$ and $\{V^{j}\}_{j=1}^{m}$, respectively. Let $\tilde{C}_{1,n}(Z)(\mathbf{u})$ $=$ $\frac{1}{n} \sum_{i=1}^{n}1(U^{i}\leq \mathbf{u})$ and $\tilde{C}_{2,m}(Z)(\mathbf{u})= \frac{1}{m}\sum_{j=1}^{m}1(V^{j}\leq \mathbf{u})$ be the empirical distributions, and let $\tilde{T}_{n,m}(Z)$ be a scaled difference between the two empirical distributions, i.e.,

equation*[equation* omitted — 113 chars of source]

The permutation distribution of $\tilde{T}_{n,m}$ can be found by verifying the Hoeffding's condition (Hoeffding, 1952). This is done in the next Lemma 3.2, where the joint convergence in (9) implies that the permutation distribution of $\tilde{T}_{n,m}$ converges to $\mathbb{\tilde{T}}$ defined therein, as $n$ and $m$ tend to infinity. In Lemma A.1 in Appendix, we additionally show that this limit process $\mathbb{\tilde{T}}$ coincides with the limit of $\tilde{T}_{n,m}(Z)$ under the null hypothesis. As a direct consequence, the permutation test based on $\Vert \tilde{T}_{n,m}(Z)\Vert _{p}$, which can be regarded as a version of our test statistic $T_{n,m}^{(p)}$ for known margins, controls the size of the test asymptotically.\footnote{ We can also verify that this test is exact employing the usual proof techniques for the permutation method (Lemma 3.2 is only an asymptotic result).}

In the following, let $\mathbf{G}_{N}$ be the set of all permutations of $ \{1,...,N\}$ and $Z^{\pi }$ $=(Z^{\pi (1)};...;Z^{\pi (N)})$ be the permuted sample for a permutation $\pi =(\pi (1),\pi (2),...,\pi (N))$ drawn from the uniform distribution on $\mathbf{G}_{N}$ independently of the data. We further denote $\pi ^{\prime }$ to be a permutation drawn from the uniform distribution on $\mathbf{G}_{N}$ independently of $\pi $ and the data.

Lemma 3.2. Under the assumption that $\lambda _{n,m}\equiv n/(n+m)=\lambda +O((n+m)^{-1/2})$, we have the convergence

equation[equation omitted — 156 chars of source]

where $\mathbb{\tilde{T}}$ is a Gaussian process that can be written as

equation*[equation* omitted — 190 chars of source]

Here, $\mathbb{\tilde{T}} ^{\prime }$ is an independent copy of $\mathbb{\tilde{T}}$, and $\mathbb{B}_{ \bar{C}}^{\prime }$ is an independent copy of $\mathbb{B}_{\bar{C}}$.

In applying the result in Lemma 3.2, however, we may encounter a practical problem because the construction of $Z$ in (8) is actually not feasible when the marginal distributions $F_{1},...,F_{d}$ and $G_{1},...,G_{d}$ are unknown. In the next step, we shall drop the condition that the univariate margins are known and instead, consider the permutations based on $\hat{Z}$,

equation[equation omitted — 145 chars of source]

in which we estimate each univariate margin using its empirical distribution. Hence now, for $i=1,...,n$ and $j=1,...,m$, we have

equation*[equation* omitted — 144 chars of source]

where $\hat{U}_{q,n}^{i}=\hat{F}_{q,n}(X_{q}^{i})$ and $\hat{V}_{q,m}^{j}= \hat{G}_{q,m}(Y_{q}^{j})$ for each\ $q=1,...,d$. Note that under this specification, $\hat{C}_{1,n}$ and $\hat{C}_{2,m}$ can be written by

equation*[equation* omitted — 193 chars of source]

Also, $\hat{T}_{n,m}$ in Lemma 3.1 is equivalent to $\tilde{T}_{n,m}(\hat{Z})$.

Lemma 3.3. For $\pi\in \mathbf{G}_{N}$, let $\hat{Z}^{\pi }=(\hat{Z}^{\pi (1)};...;\hat{Z}^{\pi (N)})$ be the permuted sample of $\hat{Z}$. Define $\tilde{C}_{1,n}^{\pi }$ and $\tilde{C} _{2,m}^{\pi }$ by

equation*[equation* omitted — 222 chars of source]

so that $\tilde{C}_{1,n}^{\pi }$ and $\tilde{C}_{2,m}^{\pi }$ are shorthand for $\tilde{C}_{1,n}(\hat{Z}^{\pi})$ and $\tilde{C}_{2,m}(\hat{Z}^{\pi})$ resepectively, and

equation*[equation* omitted — 128 chars of source]

Under the assumption that $\lambda _{n,m}\equiv n/(n+m)=\lambda +O((n+m)^{-1/2})$, we have

equation[equation omitted — 168 chars of source]

where $\mathbb{\tilde{T}}$ and $\mathbb{\tilde{T}}^{\prime }$ are the Gaussian processes defined in Lemma 3.2.

Lemma 3.3 provides some important implications for the application of permutation method. Firstly, we find that the permutation distribution of $\tilde{T}_{n,m}$ based on the permutations of $\hat{Z}$ is asymptotically equivalent to the one based on the permutations of $Z$, suggesting that $\hat{Z}$ serves as a good proxy for $Z$. However, since the limit of $\tilde{T}_{n,m}(\hat{Z})$ (equivalently, the limit of $\hat{T}_{n,m}$ in Lemma 3.1) is different from the limit of $\tilde{T}_{n,m}(\hat{Z}^{\pi})$ in Lemma 3.3, we can not validly employ the standard randomization procedure to test copula equality based on $\tilde{T}_{n,m}$.

One solution to this problem is to combine the result in Lemma 3.3 with a proper estimation procedure for the copula derivatives. To understand how, note that when $C_1$ and $C_2$ are identical, $\mathbb{\tilde{T}}$ is equivalent to $\mathbb{\tilde{T}}_{0}\equiv \sqrt{1-\lambda }\mathbb{B} _{C_{1}}-\sqrt{\lambda }\mathbb{B}_{C_{2}}$ (see Lemma A.1 in Appendix) which appears in the limit of our test statistic under the null hypothesis. After approximating this term by simulating $ \tilde{T}_{n,m}(\hat{Z}^{\pi })$ over $\pi \in \mathbf{G}_{N}$, the remaining terms are the partial derivatives of copulas that can be consistently estimated by either smoothed or non-smoothed versions of estimators such as those in Scaillet (2005), R\'{e}millard and Scaillet (2009), or Segers (2012). We formally state this result as in Theorem 3.1.

Theorem 3.1. For $\pi\in \mathbf{G}_{N}$, define $ \mathfrak{T}_{n,m}^{\pi }$ as

equation*[equation* omitted — 215 chars of source]

where $\widehat{\partial _{q}C}$ is a consistent estimator of $\partial _{q}C$ for $q=1,...,d$. When $C_1$ and $C_2$ are identical, we have

equation*[equation* omitted — 127 chars of source]

for any continuity points $c\in (0,\infty )$.

In the related work, R\'emilard and Scaillet (2009) also propose a statistical procedure for testing copula equality. Their test statistic is based on the Cram\'{e}r-von Mises distance and the critical values are obtained through an approximation of each term appearing in (6). More specifically, $\mathbb{B} _{C_{l}}$ for $l=1,2$ is approximated by a multiplier bootstrap as in Scaillet (2005), and the derivatives are estimated individually. Although we may develop a valid permutation test scheme based on Theorem 3.1, the practical advantage is relatively small because we simply end up with replacing the multiplier bootstrap part by the randomization procedure. A shortcoming of such tests is that, as the dimension $d$ increases, the number of derivative terms to be estimated increases and hence, the test procedure becomes more complicated and less precise. The problem is aggravated in the $k$-sample problem as more copulas are involved.\footnote{Besides the multiplier bootstrap method in R\'{e}millard and Scaillet (2009), other bootstrap schemes may possibly be applicable to test copula equality. For instance, Fermanian et al.\ (2004) adopted the usual i.i.d. bootstrap based on the sampling with replacement while B\"{u}cher and Dette (2010) proposed a direct multiplier bootstrap which does not require the estimation procedure of the copula derivatives. However, we shall not dwell on them in this paper because these approximations have shown be less accurate in many simulation experiments than the bootstrap method of R\'{e}millard and Scaillet (2009) that involves estimation procedure of copula derivatives. For the reason, recent studies which use bootstrap procedures for the empirical copula process rely more on the multiplier bootstrap with derivative estimation, or its extension. See Kojadinovic and Yan (2011), Kojadinovic et al.\ (2011), Genest et al.\ (2012), and Genest and Ne\v{s}lehov\'{a} (2014).}

We will thus propose another method for approximating the asymptotic null distribution of our test statistic which, unlike the proposal in Theorem 3.1 or the test in R\'{e}millard and Scaillet (2009), does not require an estimation procedure for the copula derivatives. For technical purposes we first introduce the notion of the conditional weak convergence. Let $\mathbf{A}$ be a metric space and $\xi _{N}^{\pi }\in \mathbf{A}$ be an element which depends on the data and $\pi $, regarding $\pi $ as a random element uniformly distributed on $\mathbf{G}_{N}$. Definition 3.1 on the conditional weak convergence is adopted from Kosorok (2008).

Definition 3.1. Let $BL_{1}(\mathbf{A})$ be the set of real Lipschitz functions on $\mathbf{A}$ that are uniformly bounded by one with Lipschitz constant bounded by one, let $E_{\pi }$ be the expectation over $\pi $ holding the data fixed, and let $f(\xi _{N}^{\pi })^{\ast }$ and $f(\xi _{N}^{\pi })_{\ast }$ be the minimal measurable majorant and maximal measurable minorant of $f(\xi _{N}^{\pi })$ with respect to the data and random index jointly. Then we have $\xi _{N}^{\pi }\overset{\mathrm{P}}{\underset{\pi }{ \rightsquigarrow }}\xi $ if (i) $\sup_{f\in \mathrm{BL}_{1}(\mathbf{A})}|E_{\pi }f(\xi _{N}^{\pi })-Ef(\xi )|\rightarrow 0$ in outer probability and (ii) $E_{\pi }f(\xi _{N}^{\pi })^{\ast }-E_{\pi }f(\xi _{N}^{\pi })_{\ast }\rightarrow 0$ in probability for every $f\in BL_{1}(\mathbf{A})$.

The conditional convergence in Definition 3.1 is frequently employed to verify the validity of bootstrap techniques, and can be also used to verify the asymptotic validity of other resampling procedures. In our context, the Hoeffding's condition $(\xi _{N}^{\pi },\xi _{N}^{\pi^{\prime }})\rightsquigarrow (\xi ,\xi ^{\prime })$, for a random element $\xi _{N}^{\pi } \in \mathbf{A}$, can be replaced by the conditional convergence, $\xi _{N}^{\pi }\overset{\mathrm{P}}{\underset{\pi }{\rightsquigarrow }}\xi $. This property is also used in Beare and Seo (2017) and it turned out to be particularly useful when the test statistic of interest has a nonlinear form and the continuous mapping theorem and functional delta method for the conditional convergence are to be invoked. We provide the details of the continuous mapping theorem and the functional delta method for the conditional convergence in Lemma A.2 and Lemma A.3 in the Appendix.

Secondly, we also need to introduce a differentiability result of the operator $\Phi $ defined in Section 2. Let $\mathbb{C}$ be the space of continuous functions on $[0,1]^{d}$ and let $\mathbb{H}$ be the set of functions $f\in\mathbb C$ which are grounded and pinned at the end point $ f(1,1,...,1)=0$. Then, by B\"ucher and Volgushev (2013), $\Phi $ is Hadamard differentiable at any regular copula $C$ in $\mathbb{D}$ tangentially to $\mathbb{H}$, with the derivative given by

equation*[equation* omitted — 132 chars of source]

This expression appears in our next Lemma 3.4, where we apply the functional delta method to the conditional convergence of $\tilde{T}_{n,m}(\hat{Z}^{\pi })$ implied by Lemma 3.3. Since $\bar{C}=\lambda C_{1}+(1-\lambda )C_{2}$ is a regular copula, we have

equation[equation omitted — 230 chars of source]

for any $f\in \mathbb{H}$. Based on this expression, one can easily see that the limit in Lemma 3.4 is finally equivalent to the limit of $\hat{T}_{n,m}$ when the null hypothesis is true.

Lemma 3.4. Under the assumption that $\lambda _{n,m}\equiv n/(n+m)=\lambda+o((\min (n,m))^{-1/2})$, we have

equation*[equation* omitted — 232 chars of source]

where $\Phi _{\lambda C_{1}+(1-\lambda )C_{2}}^{\prime }\mathbb{\tilde{T}}$ is the Hadamard derivative of $\Phi $ at $\lambda C_{1}+(1-\lambda )C_{2}$ in direction $\mathbb{\tilde{T}}$. In addition, we have

equation*[equation* omitted — 234 chars of source]

when $C_{1}$ and $C_{2}$ are identical.

Hence we conclude that, to approximate the correct limit of our test statistic, the permutation distribution should be obtained through computing the process defined in Lemma 3.4. However, it may not be clear how to compute $\Phi ( \tilde{C}_{1,n}^{\pi }) $ and $\Phi ( \tilde{C}_{2,m}^{\pi }) $ practically. For the last step, our Theorem 3.2 facilitates the computation by showing that we can replace $\Phi ( \tilde{C}_{1,n}^{\pi }) $ and $\Phi ( \tilde{C}_{2,m}^{\pi }) $ with the two empirical copulas computed from the permuted sample of $\hat{Z}$, as the difference is asymptotically negligible.

Theorem 3.2. For $\pi\in \mathbf{G}_{N}$, define $\mathfrak{R}_{n,m}^{\pi }$ to be

equation*[equation* omitted — 109 chars of source]

where $\hat{C}_{1,n}^{\pi }$ and $\hat{C}_{2,m}^{\pi }$ are the empirical copulas\footnote{To be explicit, the empirical copulas computed from the samples $\{\hat{Z}^{\pi (i)}\}_{i=1}^{n}$ and $\{\hat{Z}^{\pi (i)}\}_{i=n+1}^{N}$ are,

eqnarray*[eqnarray* omitted — 458 chars of source]

respectively, where $\hat{Z}_{q}^{\pi (i)}$ is the $q$-th component of $\hat{Z}^{\pi (i)}$, i.e., $\hat{Z}^{\pi (i)}=(\hat{Z}_{1}^{\pi (i)},...,\hat{Z}_{d}^{\pi (i)})$ and

equation*[equation* omitted — 208 chars of source]

for $q=1,...,d$ and $x\in \mathbb{R}$.} computed from $\hat{Z}^{\pi}$, and $c_{\alpha }^{(p)}$ to be the empirical $(1-\alpha )$ quantile of its $L_p$ norm over $\pi \in \mathbf{G}_{N}$. Assuming that $\lambda _{n,m}\equiv n/(n+m)=\lambda +o((\min (n,m))^{-1/2})$, the following statements are true.

enumerate• When $C_{1}=C_{2}$, $P(T_{n,m}^{(p)}>c_{\alpha }^{(p)}|\hat{Z})\xrightarrow{p}\alpha $ as $\min (n,m)\rightarrow \infty $. • When $C_{1}\neq C_{2}$, $P(T_{n,m}^{(p)}>c|\hat{Z})\xrightarrow{p}1$ as $\min (n,m)\rightarrow \infty $, for any constant $c\in (0,\infty )$.

We close this section by providing a step-by-step guideline to implement the test suggested by Theorem 3.2. Our strategy is simple. First, compute the test statistic $T_{n,m}^{(p)}=\Vert\tilde{T}_{n,m}(\hat{Z})\Vert _{p}=\Vert \sqrt{\frac{nm}{n+m}}( \hat{C} _{1,n}-\hat{C}_{2,m}) \Vert _{p}$. Second, we recompute $\Vert \mathfrak{R}_{n,m}^\pi\Vert _{p}=\Vert \sqrt{\frac{nm }{n+m}}( \hat{C}_{1,n}^{\pi }-\hat{C}_{2,m}^{\pi }) \Vert _{p}$ for all permutations $\pi \in \mathbf{G}_{N}$ of $\hat{Z}$, and let their ordered values be

equation*[equation* omitted — 164 chars of source]

For a nominal level $\alpha $, let $h=N!-[\alpha N!]$ where $[\alpha N!]$ denotes the largest integer less than or equal to $\alpha N!$ and let the $h$ -th largest value of the permutation statistic be $c_{\alpha }^{(p)}$, that is, $c_{\alpha }^{(p)}=\Vert \mathfrak{R}_{n,m}\Vert _{p}^{(h)}$. Then inference is made through the randomization test function $\phi^{(p)} (\hat{Z})$ constructed as

equation*[equation* omitted — 251 chars of source]

Here, $a(\hat{Z})$ is defined by

equation*[equation* omitted — 75 chars of source]

where $M^{+}(\hat{Z})$ and $M^{0}(\hat{Z})$ are the number of values of $ \left\Vert \mathfrak{R}_{n,m}\right\Vert _{p}^{(i)}$ which are greater than $ \left\Vert \mathfrak{R}_{n,m}\right\Vert _{p}^{(h)}$ and equal to $ \left\Vert \mathfrak{R}_{n,m}\right\Vert _{p}^{(h)}$, respectively. By Theorem 3.2 (i), this procedure delivers a test with limiting rejection rate equal to nominal size whenever the null hypothesis is true. Theorem 3.2 (ii) establishes the consistency of the test procedure.

In the second stage of implementation, one should be cautious not to compute $\Vert\tilde{T}_{n,m}(\hat{Z}^{\pi})\Vert _{p}$ over $\pi \in \mathbf{G}_{N}$. Since we have $\tilde{T}_{n,m}(\hat{Z})=\sqrt{\frac{nm}{n+m}}( \hat{C}_{1,n}-\hat{C}_{2,m})$, the permutation statistic $\tilde{T}_{n,m}(\hat{Z}^{\pi})$ can be easily employed to approximate the limit of $\sqrt{\frac{nm}{n+m}}( \hat{C}_{1,n}-\hat{C}_{2,m})$ without much caution. However, this is not correct as we discussed in Lemma 3.3. Given the data, the empirical copula for each group is equivalent to the empirical distribution computed from $\hat{Z}$, while the same argument is not true with regard to the permuted samples. Due to the distortion that permutation causes, each univariate component of $\hat{Z}$ no longer approximates a uniform random variable in each group after the permutations, and this leads to a discrepancy between $\tilde{T}_{n,m}(\hat{Z}^{\pi})$ and $\sqrt{\frac{nm}{n+m}}(\hat{C}^{\pi}_{1,n}-\hat{C}^{\pi}_{2,m})$. The computation based on the latter permutation statistic provides a valid approximation of the limit of our test statistic whereas the former permutation statistic can only be used to approximate $\mathbb{\tilde{T}}$.

Remark 3.1. An important distinction has to be made between the copula associated with the mixture distribution and the mixture copula. In Section 2, we argued that the copula associated with the mixture distribution $\bar{H}=\lambda H_1+(1-\lambda)H_2$ is not necessarily equal to $C_{1}$ (or $C_{2}$ ) even when $C_{1}=C_{2}$. On the other hand, the mixture copula $\bar{C}=\lambda C_1+(1-\lambda)C_2$ is always equal to $C_{1}$ (or $C_{2}$) whenever $C_{1}=C_{2}$ for any $ \lambda \in \lbrack 0,1]$. The latter property suggests that the permutations of $Z$ in (8) properly reflect the group invariance conditions implied by copula equality.

Remark 3.2. Under the condition that $F_{i}=G_{i}$ with $F_{i}^{\prime}>0$ for all $i=1,...,d$, the permutation test based on the permutations of $W$ in (2) is also valid with the test statistic $T_{n,m}^{(p)}$. This is due to the invariance property of copula, which says that for any strictly increasing transformations $\beta _{i}$ for $i=1,...,d$, the copula of $(\beta _{1}(X_{1}),...,\beta _{d}(X_{d}))$ is equivalent to the the copula of $(X_{1},...,X_{d})$. In light of the discussion in Section 2, note that $H_{1}=H_{2}$ if and only if (1) $F_{i}=G_{i}$ for all $ i=1,...,d$ and (2) $C_{1}=C_{2}$. Therefore, under the condition that $ F_{i}=G_{i}$ for all $i=1,...,d$, the two sets $\Theta_{0} $ and $\Theta _{00}$ are equal.

Remark 3.3. In related literature, Canay et al. (2017) investigate randomization tests under approximate group invariance conditions, satisfied when there is a map $S_{N}$ from a sample space $ \mathcal{W}_{N}$ to $\mathcal{S}$ such that (i) $S_{N}(W)$ for $W\in \mathcal{W}_{N}$ converges in distribution to $S$ $\in \mathcal{S}$ as $N\rightarrow \infty $, and (ii) $gS\buildrel d\over= S$ for all $g$ in some finite group of transformations $\mathbf{G}$. Our setting is somewhat different. For us, $S_{N}$ can be defined by the map which transforms $W$ into $ \hat{Z}$, where $\hat{Z}$ is an approximation to $Z$. Then an approximate group invariance holds by the distributional equality $g(Z)\buildrel d\over=Z$ for any $N$ and $g\in \mathbf{G}_{N}$. Since $\mathcal{S}$ and $\mathbf{G}$ in this context depend on $N$, our approximate group invariance condition cannot be reformulated in the framework of Canay et al. (2017) and should be handled in a different manner.

Remark 3.4. Beare and Seo (2017) also explore an asymptotic group invariance condition to develop a quasi-randomization test of copula symmetry. While $g\in \mathbf{ G}_{N}$ in our setting is specified by a permutation $\pi $ which is uniformly distributed over $\mathbf{G}_{N}$, $g\in \mathbf{G}_{n}$ in Beare and Seo (2017) is defined to be a transformation from $([0,1]^{2})^{n}$ onto itself defined by $g((x_{1},y_{1}),...,(x_{n},y_{n}))=(\pi ^{\tau _{1}}(x_{1},y_{1}),...,\pi ^{\tau _{n}}(x_{n},y_{n}))$ where $(\pi ^{0}(x,y),\pi ^{1}(x,y))=((x,y),(y,x))$ or $(\pi ^{0}(x,y),\pi ^{1}(x,y))=((x,y),(1-x,1-y))$ with $\tau _{i}$ being $n$ i.i.d. draws from the Bernoulli distribution. Unlike our framework, the group invariance conditions in Beare and Seo (2017) hold between paired normalized observations (with $n=m$) which are completely dependent with each other. Hence, our proofs are considerably different from theirs though some results on the conditional convergence in Beare and Seo (2017) are also useful to us.

Remark 3.5. It should not be overlooked that the assumptions on $\lambda _{n,m}$ have been strengthened to obtain desired results. For Lemma 3.1, we only require $ \lambda _{n,m}$ to converge to $\lambda $ as $\min (n,m)\rightarrow \infty $. For Lemma 3.2, Lemma 3.3 and Theorem 3.1, $\lambda _{n,m}$ should satisfy a certain convergence rate $\lambda _{n,m}-\lambda =O((n+m)^{-1/2})$ for applications of contiguity argument and coupling construction in Chung and Romano (2013). The assumption in Lemma 3.4 and Theorem 3.2 is stronger, requiring that $\lambda _{n,m}-\lambda $ decays faster than $(\min (n,m))^{-1/2}$.

Simulations

Here we report some Monte Carlo simulation results to show the finite sample performance of our proposed test. We particularly examine the two cases, $p=2$ and $p=\infty $ for the choice of $p$, which lead to the Cram\'er von-Mises statistic and Kolmogorov-Smirnov statistic respectively. We distinguish them by using the different notations $T_{n,m}^{(2)}$ and $T_{n,m}^{(\infty )}$. $T_{n,m}^{(2)}$ can be computed using the formula

eqnarray*[eqnarray* omitted — 438 chars of source]

while $T_{n,m}^{(\infty )}$ can be calculated as

eqnarray*[eqnarray* omitted — 526 chars of source]

We found that the test proposed in Theorem 3.1 generally leads to similar results with slightly less power than the test proposed in Theorem 3.2. Hence, we only display the rejection frequencies of the test proposed in Theorem 3.2 in this section. Throughout the simulation, we employed $1000$ replications, and in each replication we conducted $1000$ random permutations to obtain the critical values. Random sampling of the permutation does not change the asymptotic properties established in Section 3 as the number of random permutations approaches infinity.

table[table omitted — 3,435 chars of source]

We first generate two groups of i.i.d. samples from the same copula to examine the null rejection rates of our tests. For the choice of copula, we employ bivariate Gaussian, Student-t, Clayton, Frank, Gumbel, symmetrized Joe-Clayton and Plackett copulas. The parameter of each copula is fixed so as to provide the same level of dependence, which yields the correlation coefficient $0.5$ in the Gaussian case. The dependence structures of these copulas are displayed in Figure 1 of Patton (2006). As we observe in Table 4.1, the computed rejection frequencies with $T_{n,m}^{(2)}$ and $T_{n,m}^{(\infty )}$ are very close to the nominal level with the sample size of $(n,m)=(5,10)$ or $(10,5)$, while the bootstrap test by R\'{e}millard and Scaillet (2009) tends to overreject in small samples. Although our test procedure has been justified asymptotically, the table shows that it has excellent size control in finite samples. On the other hand, both permutation and bootstrap tests control size well when the sample size is over $50$ in each group.

table[table omitted — 5,568 chars of source]

Next, we provide Table 4.2 to illustrate the power of the tests for copula equality. We follow the simulation design of R\'{e}millard and Scaillet (2009) and select the copulas $C_{1}$ and $C_{2}$ from the same copula family, but with possibly different copula parameters. More specifically, we fix the Kendall's $\tau $ of $C_{1}$ to be $\tau_1=0.2$, while we let the Kendall's $\tau $ of $C_{2}$ vary over the range $\tau_{2}\in \{0.3, 0.4,0.5,0.6,0.7,0.8, 0.9\}$. Note that the results for our test with the statistic $T_{n,m}^{(2)}$ can be directly comparable to the results for the boostrap test in R\'{e}millard and Scaillet (2009) because the two tests use the same test statistic but different critical values. Since the null rejection rates of the bootstrap test are found to be between $0.12$ and $0.15$ with the sample size of $(n,m)=(10,50)$, we report size-adjusted power of the bootstrap test for this sample size. When the sample size is small, we find that the power of permutation test dominates that of bootstrap test, while the rejection frequencies are similar when the sample size is large. This phenomenon can be understood in the same context as Romano (1989).

table[table omitted — 5,635 chars of source]

Lastly, we conduct a simulation study using different marginal distributions from the uniform distribution. The two margins of $H_{1}$, namely $F_{1}$ and $F_{2}$, are taken from $N(0,1)$ while those of $H_{2}$, $G_{1}$ and $G_{2}$, are taken from $N(5,1)$. In Table 4.3, we find that changing the margins leads to no appreciable difference in rejection rates. This reflects the fact that the asymptotic property of our tests is not affected by such changes in the margins, as it should be. On the other hand, permutation tests based on the permutations of the original samples (i.e., permutations of the rows of $W$) do not yield the correct size, as we discussed in Section 2. Under the same simulation setup, we found that these tests over-reject the null hypothesis with the rejection rates reaching almost $0.3$, with the sample size $(n,m)=(100,100)$. Increasing the sample size does not improve the size of the tests. We also performed a simulation study based on the different setting of $F_{1}=G_{1} \sim N(0,1)$ and $F_{2}=G_{2}\sim N(5,1)$. Under this framework, both of the randomization tests by permuting the rows of $W$ and the rows of $\hat{Z}$ are valid by Remark 3.2. We do not report these additional results separately because the rejection rates are very similar to those in Table 4.2 (or Table 4.3).

Empirical applications

Dependence of income and consumption

Household consumption decisions are of central concern in economics. In the short run, fluctuations in consumption induce business cycles while in the long run, consumption behavior serves as a primary determinant of economic growth. In the classical model of Keynes (1936), consumption is represented as a function of income. Although there can be other determinants of consumption, substantial empirical evidence shows that disposable income plays the most important role in explaining consumer behavior. See Hall (1978), Flavin (1981), Hall and Mishkin (1982), Campbell and Deaton (1989), Shapiro and Slemrod (1995), Shea (1995), Parker (1999), Souleles (1999), Johnson et al.\ (2006), Parker et al.\ (2013) and Kaplan and Violante (2014) for research along these lines.\footnote{There has been a controversial debate on the relationship between income and consumption. According to the Permanent Income Hypothesis (Friedman, 1957), individual income consists of transitory income and permanent income, and consumption is determined by the permanent income component rather than the transitory income component. In a similar context, the Life Cycle Hypothesis (Ando and Modigliani, 1963) assumes that the utility of an individual consumer depends on his own total consumption in current and future periods, and utility maximization under intertemporal budget constraints yields the solution of current consumption expressed in terms of the total resources over the life time and the rate of capital return. From the perspective of the Permanent Income Hypothesis or Life Cycle Hypothesis, current consumption should not be affected by a change in transitory income or anticipated income. See Sargent (1978), Browning and Collado (2001) and Hsieh (2003) for related empirical evidence.} In this section, we estimate copulas to study how consumption relates to disposable income with a focus on their dependence structure. By applying the proposed test of copula equality, we examine whether the dependence structure is identical in several different circumstances.

Since our goal is to study the structure of dependence, we explore micro-level household data of income and consumption in several countries at different stages of economic development. For the cross-country analysis, we investigate household annual income and annual consumption in 2010 collected from the U.S., Mexico and South Africa. In each survey, we compute household income by summing over the annual income from work, transfers, rental, and other miscellaneous income (including public assistance), after deducting the taxes. For the household consumption, we use the total sum of the expenses that households made on the food, clothing, housing, health care, education, transportation, trip, furniture and equipment, entertainment and other miscellaneous expenditure for a year. The sample sizes of the data for the U.S., Mexico, and South Africa are 16,803, 27,614 and 25,243, respectively.\footnote{Data sources are \href{https://www.bls.gov/cex/pumd_data.htm}{the U.S. Consumer Expenditure Survey} (https://www.bls.gov/cex/pumddata.htm), \href{http://www.beta.inegi.org.mx/proyectos/enchogares/regulares/enigh/tradicional/2010/default.html}{Mexico Household Income and Expenditure Survey} (https://www.beta.inegi.org.mx/proyectos/enchogares /regulares/enigh/tradicional/2010/default.html) and \href{https://www.datafirst.uct.ac.za/dataportal/index.php/catalog/316}{Income and Expenditure Survey in South Africa} (https://www.datafirst.uct.ac.za/dataportal/index.php/catalog/316).}

To provide an overview of the dependence structure between income and consumption, we apply the probability integral transforms to the data of the three countries and display the scatterplots in Figure 5.1. We observe that the dependence of income and consumption is stronger in South Africa and Mexico than in the U.S. To be precise, the correlation coefficients of income and consumption in the U.S., Mexico, and South Africa, are 0.54, 0.74 and 0.76, while the Kendall's $\tau$ are $0.62$, $0.78$ and $0.76$ respectively. In general, the ratio of consumption to income is relatively higher in developing countries, which leads to stronger positive dependence in the relation of income and consumption. In addition, consumers face lack of credit and insurance in developing countries, and this results in less consumption smoothing over the life time. As a consequence, the dependence of current income and current consumption may tend to be stronger in developing countries (See Jappelli and Pagano; 1989, Campbell and Mankiw; 1991, Rosenzweig and Wolpin; 1993, Zimmerman and Carter; 2003, Gin\'{e} and Yang; 2009 and Karlan et al.\ (2014)).

figure[figure omitted — 317 chars of source]

Here, we apply our test to the income and consumption data of the three countries. By applying our test procedure, we may detect difference in the degree of dependence between different group as well as a discrepancy in the structure of the dependence. In the comparison of the U.S., and Mexico, the test statistics are computed as $(T_{n,m}^{(2)}, T_{n,m}^{(\infty )})=(1.49, 3.32)$ and the p-values for both statistics turn out to be zero, suggesting that the dependence structures of income and consumption in the two countries are significantly different. In the comparison of Mexico and South Africa, the test statistics are computed as $(T_{n,m}^{(2)}, T_{n,m}^{(\infty )})=(0.43, 1.30)$. Surprisingly, our randomization test leads to zero p-values for both $T_{n,m}^{(2)}$ and $T_{n,m}^{(\infty )}$ despite that the correlation coefficients and Kendall's $\tau$ are very similar in South Africa and Mexico. This indicates that a strong dissimilarity detected by our test procedure may not be detected by correlation coefficients or Kendall's $\tau$.

table[table omitted — 892 chars of source]

Table 5.1 provides more detailed information on the structure of income and consumption dependence in the U.S., Mexico and South Africa in 2010. In the table, we report the conditional Kendall's $\tau$ of income and consumption at several different exceedance levels. Recall that for a pair of random variables $(X,Y)$, Kendall's $\tau$ is defined by $\tau (X,Y) =P((X_{i}-X_{j})(Y_{i}-Y_{j})>0)-P((X_{i}-X_{j})(Y_{i}-Y_{j})<0)$, where $(X_{i},Y_{i})$ and $(X_{j},Y_{j})$ are two random draws of $(X,Y)$. Then, the conditional Kendall's $\tau$ at the exceedance level $c\in (0,1)$ is provided by\footnote{The exceedance Kendall's $\tau$ is an analogue to the exceedance correlation in Longin and Solnik (2001), Ang and Chen (2002), and Hong and Zhou (2007). While the exceedance correlation only captures linear conditional dependence, the exceedance Kendall's $\tau$ can also capture nonlinear features of conditional dependence. See also, Manner (2010).}

eqnarray*[eqnarray* omitted — 164 chars of source]

We observe that in all three countries, the upper conditional dependence is stronger than the lower conditional dependence at any fixed exceedance level. It implies that within a country, the dependence patterns are different depending on relative income and consumption levels. In particular, we found that when both income and consumption levels are relatively low, a greater portion of consumption is on necessities, and the consumption of necessities does not increase much by an increase in the income. On the other hand, when both income and consumption levels are relatively high, consumers spend a greater portion of their income on luxury goods. Hence, the dependence between income and consumption tends to be stronger in the upper tail than in the lower tail.\footnote{Our analysis in Table 5.1 is based on the comparisons between the dependence in the left lower quadrant and right upper quadrant, and the result does not imply that consumption change is more sensitive to income change at higher income level. In fact, the estimated marginal propensity to consumption (MPC) for our data tends to increase when income is low, but it starts to decrease at around the 60th percentile of income in each country. As income increases, the estimated MPC of the U.S. increases up to 0.32 and decreases to 0.17, while that of Mexico increases up to 0.63, and decrease to 0.40. In South Africa, the estimated MPC increases from 0.52 to 0.71, but then decrease again to 0.50.}

In addition, Table 5.1 shows that both the upper conditional dependence and the lower conditional dependence are weaker in the U.S. than in Mexico and South Africa at all exceedance levels. This is consistent with our finding that the U.S. has the lowest correlation coefficient and the lowest Kendall's $\tau$ among the three countries. On the other hand, compared to Mexico, South Africa has smaller lower conditional dependence but larger upper conditional dependence at any fixed exceedance level. As the surplus in upside dependence is offset by the deficit in downside dependence, the value of the correlation coefficients or Kendall's $\tau$ may be similar to that of Mexico. Nevertheless, the patterns of asymmetric dependence in the two countries are very different, and our test procedure can distinguish the dissimilarity of such nonlinear patterns.

figure[figure omitted — 423 chars of source]

Since the U.S. Consumer Expenditure Survey provides richer information than the Mexican and South African surveys, we may explore in more depth the recent changes in the income and consumption dependence in the U.S. Hence, we provide additional analysis on the dependence of income and total consumption in the U.S. in 2006, 2010 and 2014, as well as that of income and non-durable consumption in 2006, 2010 and 2014. Following the definition in Attanasio et al.\ (2012), we define the consumption of non-durable goods as the household expenditure on food, clothing, footwear and non-durable entertainment. Based on this definition, we find that the correlation coefficients of income and consumption are very similar in the three years; the correlation coefficients of income and total consumption are 0.5, 0.53 and 0.5 in 2006, 2010 and 2014 respectively, while those of income and non-durable consumption are 0.44, 0.41 and 0.37.

Figure 5.2 provides the graphs of the conditional Kendall's $\tau$ of income and consumption in the U.S. in 2006, 2010 and 2014. For both total consumption and non-durable consumption, the upper conditional Kendall's $\tau$ remains almost the same in the three years, while the lower conditional Kendall's $\tau$ tends to decrease slightly over time. The decrease in the lower conditional dependence is more noticeable between 2006 and 2010 than between 2010 and 2014. Our inference results reflect this finding. In testing for equal dependence of total consumption and income between 2006 and 2010, our p-values for $T_{n,m}^{(2)}$ and $T_{n,m}^{(\infty )}$ are computed as 0.04 and 0.08, while they are 0.30 and 0.99 between 2010 and 2014. Using non-durable consumption, we obtain slightly smaller p-values. The p-values for $T_{n,m}^{(2)}$ and $T_{n,m}^{(\infty )}$ are 0.03 and 0.01 between 2006 and 2010, while they are 0.22 and 0.94 between 2010 and 2014. The results suggest that for both specifications of consumption, the dependence structure of income and consumption is significantly different between 2006 and 2010 at the $5\%$ and $10\%$ significance levels. However, the difference is less significant between 2010 and 2014.

Brexit effect on financial integration in Europe

Since established, the European Union (EU) has served as a coalition which fostered the integration of economic development among European countries. The economic integration has forced the economic decisions of the countries in the EU to be intimately dependent on each other. The countries have benefited from the upside gains but also suffered from the downside losses. Among others, the U.K. has played a central role in the decision making process of EU, providing a large portion of its budget over the past years. However, in pursuit of financial and political independence, the U.K. eventually announced `Brexit' (Britain's exit from the EU) in June of 2016. The term `Brexit', which had been previously used for the potential exit of the U.K. from the EU, does not refer to a hypothetical situation anymore.

The announcement of Brexit created a huge external shock to the international economy. Following the announcement, there have been massive debates on what would happen to the U.K. and EU in the future. One of the major concerns is the impact of the U.K.'s decision on the economic integration in Europe. In this section, we aim to provide statistical evidence of the `Brexit effect' on European financial market integration by investigating stock returns in the U.K., France, Germany and Switzerland. Although the process of Brexit is still ongoing and no definitive answer may be given to the question of the long run effect of Brexit, our findings in this section will provide evidence of the Brexit effect in its early and transitional stage.

figure[figure omitted — 344 chars of source]

For our analysis, we collect the four stock indices of FTSE (U.K.), CAC (France), DAX (Germany) and SMI (Switzerland) from February 2013 to July 2017, at a daily frequency. Using the announcement of the Brexit as a cutoff point, we divide our observations into two groups, the data of `before' the Brexit announcement and `after' the Brexit announcement. To be more precise, we define the `pre-Brexit' period as February 2013 through January 2016, and `Brexit' period as July 2016 through December 2017. Although the referendum for Brexit was held in June 2016, the financial market structure in Europe has been potentially influenced by the announcement in February 2016 of making such a poll. Thus, we leave out the observations from February 2016 to June 2016 in defining the `pre-Brexit' period. \footnote{ We have also examined the Brexit effect on the European financial market based on different definitions of pre-Brexit period. Neither of using one year nor two years of observations before February 2016 as pre-Brexit data changes the overall inference results presented in this section.} After excluding weekends and public holidays, the sample sizes are 734 for the pre-Brexit period and 338 for the Brexit period. These indices, as displayed in Figure 5.3, present a high degree of comovement in both pre-Brexit and Brexit periods. To remove temporal dependence, we apply the AR(1)-GJR-GARCH $(1,1)$ filter to the stock returns obtained by taking the log-difference of the indices.\footnote{ We examined several different GARCH specifications with student $t$, skewed $t$ and Gaussian innovations, and found that the AR(1)-GJR-GARCH(1,1) model with student $t$ innovations provides the best fit. After applying a proper GARCH filter, the filtered stock returns are conventionally regarded as i.i.d. in the empirical finance literature. See Abhyankar et al. (1997), Manner(2001), Hu (2006), Roch and Alegre (2006), Cotter (2007), Aas et al. (2009), Giacomini et al. (2009), Hu and Kercheval (2010), Aloui et al. (2011), Nikoloulopoulos et al. (2012), St\ove et al. (2014) and many others. Specifically, Chan et al. (2009) provided a theoretical justification for the GARCH residuals based estimation of copulas in the semi-parametric setting.}

figure[figure omitted — 479 chars of source]

Figure 5.4 displays scatterplots of the FTSE against CAC, DAX, and SMI in the pre-Brexit period and the Brexit period. The stock returns have been normalized by applying the probability integral transforms. Casual inspection reveals that the observations are less concentrated on the diagonal line in the Brexit period, suggesting that the cross market comovement of stock returns has become weaker after the U.K. decided to leave the EU. This can also be observed from the changes in the correlation coefficient. The correlation coefficient of FTSE and CAC is 0.86 in the pre-Brexit period and 0.71 in the Brexit period. Similarly, the correlation coefficients of FTSE and DAX, and of FTSE and SMI also have decreased from 0.82 to 0.65 and from 0.71 to 0.66, respectively.

table[table omitted — 1,013 chars of source]

In Table 5.2, we report our randomization test results concerning the Brexit effect. Firstly, our results show that the U.K.'s financial market dependence on the French and German financial markets has completely changed after the Brexit announcement. Applying our randomization test to the pairs of stock returns (FTSE, CAC) and (FTSE, DAX) yields p-values close to zero, which indicates overwhelming rejection of equality in the pairwise dependence. This change is due to the overall decrease in dependency (a change in the degree of dependence), and unequal decrease in dependency--more decrease in the downside than in the upside-- that leads to asymmetric dependence (a change in the structure of dependence). For a fixed level $c=0.5$ for instance, the upper conditional Kendall's $\tau$ of FTSE and CAC slightly decreased from $0.76$ to $0.65$ while the lower conditional Kendall's $\tau$ substantially decreased from $0.77$ to $0.26$. In the (FTSE, DAX) pair, the upper conditional Kendall's $\tau$ at level $c=0.5$ decreased from $0.44$ to $0.35$, while the lower conditional Kendall's $\tau$ decreased from $0.43$ to $0.14$, and again, the decrease is more prominent in the downside than in the upside. On the other hand, the change of the dependence structure in the pair of (FTSE, SMI) turns out to be less significant: we fail to reject the null hypothesis of copula equality at both $5\%$ and $10\%$ significance levels. We also found that in terms of distributional symmetry, no obvious change in the pair (FTSE, SMI) is observed after the Brexit announcement.

While there is no universal consensus on the extension of the correlation coefficient or Kendall's $\tau$ to more than two variables, our statistical procedure can be naturally applied to study higher dimensional dependence structures. For instance, by comparing the four dimensional copulas of stock returns to the FTSE, CAC, DAX, and SMI in the pre-Brexit and Brexit period, we can perform a statistical test to examine the Brexit effect on the dependence structure of the four returns. It turns out that our randomization tests with the statistics $T_{n,m}^{(2)}$ and $T_{n,m}^{(\infty )}$ yield p-values of $0.01$ and $0.02$ respectively, indicating that there has been a significance change in mutual dependence of the stock returns of FTSE, CAC, DAX, and SMI after the announcement of Brexit. Similarly, we can also examine the Brexit effect on the dependence of CAC, DAX and SMI, excluding FTSE. The question arises when our concern is to test for the Brexit effect on the financial market dependence among remaining EU countries. Here, our inference provides somewhat mixed results: we detect a significant change in the dependence structure of CAC, DAX and SMI with the test statistic $T_{n,m}^{(2)}$ at both $5\%$ and $10\%$ levels, but the change is less significant with $T_{n,m}^{(\infty )}$.

To summarize, our empirical findings reveal that the European financial market has experienced a substantial change after the Brexit announcement. Firstly, the dependence of financial market returns between the U.K. and other European countries has decreased after the Brexit announcement. In particular, we observe a greater decrease in dependence during market downturns than market upturns, which implies that the U.K.'s financial market is now less likely to crash together with other financial markets. Secondly, the financial integration among the U.K., France, Germany and Switzerland has decreased in reaction to the announcement of Brexit. Our test results indicate that there has been a significant change in the dependence structure of the stock markets in the four countries. We also found some evidence that the announcement of Brexit not only changed the financial market dependence of the U.K., but also that of the remaining EU countries. Lastly, the change in the dependence structure seems to be more recognizable among the larger economies. In our analysis, Brexit has caused more significant influence on financial market dependence among the U.K., France and Germany, while Switzerland has been less affected by the Brexit announcement.

Final remarks

Randomization tests provide useful tools for inference in situations where the null hypothesis implies that the distribution of the data is invariant to a group of transformations. In particular, permutation tests have been widely applied for testing the equality of distribution functions. Since copulas constitute a class of distribution functions, they may be naturally employed for testing the equality of copulas. However, the problem has not been examined to date and this is an important missing point in the literature.

This paper examines the use of permutation methods for testing copula equality. Unfortunately, the classical permutation method is not applicable because we do not observe samples directly from the copulas. Although we may instead explore the pseudo samples which consist of normalized ranks, the application of the permutation method still requires caution due to the distortion that permutation causes to the univariate margins of the pseudo samples. Asymptotically valid inference can be achieved by either modifying the permutation method considering the additional terms of the limit distribution induced by the uncertainty in margins (Theorem 3.1), or by correcting the distortion through taking one more step of normalizing the margins after each permutation (Theorem 3.2).