EconBase
← Back to paper

Distributional Granger Causality: Identification, Sequential Inference, and Adaptive Testing

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

54,854 characters · 24 sections · 26 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Distributional Granger Causality: Identification, Sequential Inference, and Adaptive Testing

abstractPredictive dependence in time series need not be confined to the conditional mean. Outside the Gaussian setting, causal content may arise through conditional scale, tail behavior, asymmetry, or other distributional features, implying that no single Granger-type test provides a complete characterization of predictive dependence. This paper develops a framework for distributional Granger causality based on a finite collection of channel-specific restrictions. Under suitable determinacy conditions, the channel menu is shown to be complete, yielding an identification result that links distributional Granger non-causality to a finite set of testable hypotheses. Building on this representation, we develop an adaptive sequential testing procedure that allocates inferential resources across channels while maintaining familywise error control through an alpha-investing mechanism. A policy-invariant validity theorem establishes finite-sample size control under arbitrary admissible selection rules, while an asymptotic efficiency theorem shows that a confidence-bound allocation rule achieves power equivalent to that of an infeasible oracle benchmark. The theoretical guarantees are derived from primitive mixing and moment conditions together with a circular-block permutation scheme.
keywordsDistributional Granger causality; Sequential testing; Multiple testing; Adaptive inference.
jelC12; C14; C22; C52.

\onehalfspacing

Introduction

Granger causality is among the most widely used tools for studying predictive relationships in time series. In its classical form, the concept is operationalized through incremental forecasting ability: a process $X_t$ is said to Granger-cause a process $Y_t$ if past values of $X_t$ improve forecasts of $Y_t$ beyond the information contained in the past of $Y_t$ alone Granger1969,Sims1972. Under linear dynamics and Gaussian innovations, this notion admits a particularly simple characterization through lag-exclusion restrictions in vector autoregressions Geweke1982,Geweke1984. In such environments, the conditional distribution is fully summarized by its first two moments, and causal inference reduces to a problem of conditional-mean predictability.

Outside the Gaussian setting, however, predictive dependence need not be confined to the conditional mean. A predictor may influence the conditional variance of future outcomes, alter tail probabilities, modify higher-order moments, or affect other features of the conditional distribution without changing expected values. This observation has motivated a large literature on alternative notions of Granger causality, including causality in variance CheungNg1996, causality in risk and tail events HongLiuWang2009,WhiteKimManganelli2015, causality in quantiles JeongHardleSong2012,SongTaamouti2021, nonlinear and nonparametric causality HiemstraJones1994,DiksPanchenko2006,NishiyamaHitomiKawasakiJeong2011, copula-based dependence measures BouezmarniRomboutsTaamouti2012, and information-theoretic approaches such as transfer entropy BarnettSethBossomaier2009.

The resulting literature provides a rich collection of channel-specific tests but leaves open a fundamental question. If predictive dependence may arise through multiple dimensions of the conditional distribution, how should inference be conducted when the relevant channel is unknown ex ante? In empirical practice, researchers often evaluate several causality tests and interpret the resulting collection of $p$-values informally. This approach faces two difficulties. First, simultaneous consideration of multiple channels generates a multiple-testing problem, potentially leading to substantial distortions in familywise error rates. Second, there is generally no principled rule for allocating finite inferential resources across competing tests or for determining which channel should receive the greatest attention.

This paper develops a unified framework for distributional Granger causality that addresses both issues. The starting point is an identification result. Rather than viewing existing causality tests as competing methodologies, the paper interprets them as measuring distinct coordinates of a common object: the conditional distribution of future outcomes. Distributional Granger non-causality is characterized through a finite collection of channel-specific restrictions corresponding to conditional location, scale, tail behavior, and higher-order distributional features. Under suitable determinacy conditions, these restrictions are shown to be complete in the sense that distributional Granger non-causality holds if and only if every channel-specific restriction is satisfied. This representation transforms an infinite-dimensional hypothesis concerning conditional distributions into a finite collection of testable components while preserving identification of the underlying causal object.

Building on this representation, the paper develops an adaptive sequential testing procedure for channel selection. The procedure allocates a finite testing budget across channels, updates allocation decisions using previously observed outcomes, and terminates once sufficient evidence against the null has been accumulated or the testing budget has been exhausted. The inferential framework combines channel-specific hypothesis tests with an alpha-investing mechanism drawn from the sequential multiple-testing literature FosterStine2008,AharoniRosset2014. This construction permits data-dependent channel selection while maintaining rigorous control of familywise error rates.

The theoretical contribution consists of three results. First, a completeness theorem establishes that the proposed channel menu fully characterizes distributional Granger non-causality. Second, a policy-invariant familywise error theorem shows that inferential validity is preserved under arbitrary admissible channel-selection rules. The result separates validity from allocation: any adaptive policy satisfying the filtration-adaptedness requirements inherits the same finite-sample error guarantee. Third, an asymptotic efficiency theorem demonstrates that a confidence-bound allocation rule achieves power equivalent to that of an infeasible oracle benchmark in the limit. Consequently, the power loss associated with uncertainty regarding the active channel vanishes asymptotically.

An important feature of the analysis is that the theoretical guarantees are derived from primitive conditions on the data-generating process rather than imposed as abstract assumptions. Conditional super-uniformity of the channel-specific $p$-values follows from a circular-block permutation procedure under strict stationarity and suitable mixing conditions. Likewise, the concentration properties required for the efficiency analysis are obtained from Bernstein-type inequalities for dependent processes. These results provide an explicit link between the high-level assumptions underlying the sequential testing framework and standard regularity conditions for weakly dependent time series.

The framework has implications for several areas of applied econometrics. In financial applications, predictive dependence frequently appears through volatility, downside risk, or tail exposure rather than through expected returns. Similar considerations arise in macroeconomics, network analysis, and systemic-risk measurement, where distributional spillovers are often more informative than conditional-mean effects. By treating causality as a property of the entire conditional distribution rather than a single moment condition, the proposed framework accommodates such settings while preserving formal inferential guarantees.

The remainder of the paper is organized as follows. Section (ref) reviews the related literature. Section (ref) develops the distributional-causality framework and establishes the completeness of the channel representation. Section (ref) introduces the adaptive sequential testing procedure and presents the validity and efficiency results. Section (ref) reports Monte Carlo evidence on finite-sample performance. Section (ref) concludes. Proofs and additional technical results are collected in the Online Appendix.

Related Literature

This paper contributes to three strands of the econometrics literature: distributional approaches to Granger causality, sequential multiple-testing procedures with error-control guarantees, and adaptive allocation methods for statistical experimentation.

Distributional Granger causality

The classical formulation of Granger causality characterizes predictive dependence through incremental linear forecasting ability Granger1969,Sims1972,Geweke1982,Geweke1984. Under Gaussianity and linear dynamics, this characterization is sufficient because the conditional distribution is fully summarized by its first two moments. Outside that setting, however, predictive content may arise through conditional scale, tail behavior, asymmetry, or other higher-order distributional features that are not captured by conditional-mean restrictions alone.

A substantial literature has therefore developed tests targeting specific dimensions of predictive dependence. Examples include causality in variance CheungNg1996, causality in risk and tail events HongLiuWang2009,WhiteKimManganelli2015, causality in distribution CandelonTokpavi2016, causality in quantiles JeongHardleSong2012,SongTaamouti2021, nonlinear and nonparametric causality HiemstraJones1994,DiksPanchenko2006,NishiyamaHitomiKawasakiJeong2011, and copula-based approaches to conditional dependence BouezmarniRomboutsTaamouti2012. Information-theoretic measures such as transfer entropy provide additional tools for detecting departures from linear predictability and coincide with classical Granger causality under Gaussianity BarnettSethBossomaier2009.

Related developments arise in the literature on identification under non-Gaussianity, where higher-order moments and distributional features play a central role in recovering structural relationships ShimizuHoyerHyvarinenKerminen2006,LanneMeitzSaikkonen2017,GourierouxMonfortRenne2017,MontielOleaPlagborgMollerQian2022. Higher-order cumulants likewise provide a natural representation of departures from Gaussianity and characterize important dimensions of distributional shape ArevalilloNavarro2026. Recent work has also continued to refine inference for predictive regressions and Granger-causality tests in nonstandard environments, including settings involving boundary parameters and weak identification CavaliereGeorgiev2020,CavaliereGeorgievZanelli2025.

The present paper differs from this literature in its objective. Rather than proposing a new channel-specific test, it studies how existing tests can be combined within a unified framework. The contribution is an identification result showing that, under suitable determinacy conditions, a sufficiently rich collection of channel-specific restrictions provides a complete characterization of distributional Granger non-causality.

Sequential multiple testing and online inference

Testing multiple dimensions of predictive dependence naturally raises a multiple-testing problem. Classical approaches control familywise error rates through fixed corrections such as Bonferroni procedures, but these methods are static and often conservative. A complementary literature studies sequential and online testing procedures that allocate significance levels adaptively as hypotheses are examined FosterStine2008,AharoniRosset2014,JavanmardMontanari2018,RamdasZrnicWainwrightJordan2018.

Closely related developments arise in the literature on anytime-valid inference, test martingales, and e-processes, which provide inferential guarantees that remain valid under optional stopping and adaptive continuation rules RamdasGrunwaldVovkShafer2023,GrunwaldHeideKoolen2024. These methods emphasize the construction of inferential procedures whose validity is preserved under data-dependent experimentation.

The present paper applies these ideas to distributional causality testing. The distinguishing feature of the setting considered here is that the sequence of hypotheses is not exogenously given but is selected adaptively from a menu of candidate channels. The resulting policy-invariant familywise error theorem establishes that valid inference is preserved under arbitrary admissible channel-selection rules. In this sense, the paper extends sequential testing methods to an environment in which both testing and selection are data dependent.

Adaptive allocation and statistical efficiency

The allocation of finite testing resources across competing hypotheses is closely related to the literature on sequential experimentation and adaptive allocation. Multi-armed bandit methods provide a canonical framework for balancing exploration and exploitation and have generated a large body of results concerning regret minimization and efficient learning AuerCesaBianchiFischer2002,LattimoreSzepesvari2020,KaufmannCappeGarivier2016.

The problem considered here differs from the standard bandit setting in both objective and interpretation. The goal is not reward maximization per se, but efficient allocation of inferential effort across competing dimensions of predictive dependence. The objects being learned are channel-specific measures of statistical informativeness, represented through the non-centrality parameters of the underlying tests.

This perspective connects with a broader movement toward adaptive and data-driven procedures in econometrics ChernozhukovEtAl2018,FarrellLiangMisra2021. Recent methodological discussions have emphasized the growing interaction between econometric theory, machine learning, and artificial intelligence GuggenbergerSuSun2026. The contribution of the present paper is to embed adaptive allocation within a formal inferential framework. The resulting procedure combines three features that are typically studied separately: identification of the causal object, finite-sample error control, and asymptotically efficient allocation of testing resources. The asymptotic oracle-efficiency result demonstrates that adaptation can improve power without compromising the validity guarantees established by the sequential testing framework.

Framework: Channel Decomposition and Menu Completeness

Notation and information sets

Let $\{(X_t,Y_t)\}_{t\in\mathbb{Z}}$ be a strictly stationary bivariate stochastic process defined on the probability space $(\Omega,\mathcal{F},\mathbb{P})$. Define the information sets \[ \mathcal{I}_{t-1} := \sigma\big(\{X_{t-j},Y_{t-j}\}_{j\ge 1}\big), \qquad \mathcal{I}^{(-X)}_{t-1} := \sigma\big(\{Y_{t-j}\}_{j\ge 1}\big). \] Let $F_{Y_t\mid\mathcal{J}}(\cdot)$ denote a regular conditional distribution of $Y_t$ given a sub-$\sigma$-field $\mathcal{J}$, and let $Q_{Y_t}(\tau\mid\mathcal{J})$ denote the associated conditional $\tau$-quantile. Unless otherwise stated, all equalities and inequalities involving conditional objects hold $\mathbb{P}$-almost surely, with continuity-point qualifications imposed where required.

The subsequent asymptotic analysis relies on a collection of primitive conditions governing dependence, moments, and conditional smoothness. These assumptions are standard in the nonparametric and resampling literature for weakly dependent stochastic processes and are sufficient to establish the higher-level regularity conditions employed in the finite-sample size and power results developed in Section (ref). Formal derivations are provided in Appendix (ref).

assumption[Data-generating primitives] The process $\{(X_t,Y_t)\}_{t\in\mathbb{Z}}$ is strictly stationary and absolutely regular ($\beta$-mixing) with coefficients $\beta(m)\le c_0\,\rho^{m}$ for some $c_0<\infty$, $\rho\in(0,1)$ (geometric mixing; polynomial decay $\beta(m)=O(m^{-b})$ with $b>2$ suffices for all results at the stated rates). The variables admit $4+\epsilon$ finite moments, $\mathbb{E}\,|Y_t|^{4+\epsilon}+\mathbb{E}\,|X_t|^{4+\epsilon}<\infty$ for some $\epsilon>0$. The conditional distribution $F_{Y_t\mid\mathcal{I}_{t-1}}$ has a density that is bounded and, near each tail quantile $Q_{Y_t}(\tau_L\mid\cdot)$, $Q_{Y_t}(\tau_U\mid\cdot)$, bounded away from zero and Lipschitz in $y$ uniformly in the conditioning history.
definition[Circular-block permutation scheme] For a channel statistic $S_k$ computed from $\{(X_t,Y_t)\}_{t=1}^T$, define the permutation $p$-value \[ P_k = \frac{ 1+\#\{b:S_k(X^{(b)},Y)\ge S_k(X,Y)\} }{B+1}, \] where each permuted sequence is generated by the circular shift \[ X^{(b)}_t=X_{((t+s_b-1)\bmod T)+1}, \] with $s_b$ drawn independently and uniformly from $\{1,\ldots,T-1\}$ for $b=1,\ldots,B$. The circular shift preserves the marginal distribution and serial dependence structure of $\{X_t\}$. Under the channel null hypothesis $H_{0,k}$, the resulting statistic is invariant in distribution to such shifts, providing the basis for permutation validity; see Appendix (ref).

Distributional Granger non-causality

The object of interest is the conditional distribution of $Y_t$ given the available information. The strongest form of predictive irrelevance is defined through equality of conditional laws.

definition[Granger non-causality in distribution] The process $X$ does not Granger-cause $Y$ in distribution if \[ F_{Y_t\mid\mathcal{I}_{t-1}}(y) = F_{Y_t\mid\mathcal{I}^{(-X)}_{t-1}}(y) \quad \forall y\in\mathbb{R},\ \forall t. \]

Definition (ref) characterizes predictive irrelevance at the level of the entire conditional distribution. It therefore excludes any incremental predictive contribution of the history of $X$ once the history of $Y$ has been conditioned upon. Since this condition is infinite-dimensional, direct implementation is generally infeasible. The framework developed below replaces the distributional null by a collection of lower-dimensional restrictions defined through measurable functionals of the conditional law.

Channel functionals and channel nulls

Let $\mathcal{M}=\{1,\ldots,K\}$ index a finite collection of measurable functionals of the conditional distribution. For each channel $k\in\mathcal{M}$, let $\phi_k$ denote a functional defined on the space of conditional laws. The corresponding channel null hypothesis is \[ H_{0,k}: \quad \phi_k\!\left(F_{Y_t\mid\mathcal{I}_{t-1}}\right) = \phi_k\!\left(F_{Y_t\mid\mathcal{I}^{(-X)}_{t-1}}\right) \quad \mathbb{P}\text{-a.s.}, \qquad \forall t. \]

The canonical menu considered throughout the paper is given by

align[align omitted — 427 chars of source]

The first four channels correspond to conditional location, scale, and tail behavior. The fifth channel captures higher-order distributional features through conditional skewness- and kurtosis-related cumulants. Each channel induces a testable restriction and may be associated with an established inferential procedure. Throughout, the testing methodology is treated as given; the emphasis is on the logical relationship between the collection of channel restrictions and distributional Granger non-causality.

The nesting property

The channel hypotheses are nested within the distributional null.

proposition[Distributional non-causality implies every channel null] Suppose the functionals $\phi_k$ in (ref) are well defined. If $X$ does not Granger-cause $Y$ in distribution according to Definition (ref), then $H_{0,k}$ holds for every $k\in\mathcal{M}$.
proofUnder Definition (ref), \[ F_{Y_t\mid\mathcal{I}_{t-1}} = F_{Y_t\mid\mathcal{I}^{(-X)}_{t-1}} \qquad \mathbb{P}\text{-a.s.} \] for every $t$. Since each $\phi_k$ is a measurable functional of the conditional distribution, application of $\phi_k$ to both sides yields \[ \phi_k\!\left(F_{Y_t\mid\mathcal{I}_{t-1}}\right) = \phi_k\!\left(F_{Y_t\mid\mathcal{I}^{(-X)}_{t-1}}\right), \] which establishes $H_{0,k}$.

Proposition (ref) establishes that distributional Granger non-causality implies the validity of every channel restriction. Consequently, rejection of any channel null is sufficient to reject distributional non-causality. The converse implication, however, requires additional structure. Specifically, the collection of channel functionals must be sufficiently informative to identify the conditional law.

Menu completeness

To establish the converse implication, the collection of channel functionals must uniquely characterize the conditional law within an admissible class of distributions. The following assumption formalizes this requirement.

assumption[Determinacy class] For each $t$, the conditional law $F_{Y_t\mid\mathcal{I}_{t-1}}$ belongs $\mathbb{P}$-a.s.\ to a family $\mathcal{D}$ satisfying the following properties: \begin{enumerate} • Every distribution in $\mathcal{D}$ is uniquely determined by its cumulant sequence together with its lower- and upper-tail quantile functions; • $\mathcal{D}$ is closed under conditioning. \end{enumerate} The skewed scale-mixture families commonly employed in financial econometrics satisfy these requirements over the parameter regions considered in empirical applications.

Assumption (ref) is an identification condition. It guarantees that equality of the functionals defining the menu implies equality of the underlying conditional distributions. Consequently, the finite collection of channel restrictions can be used to recover an infinite-dimensional statement concerning conditional laws.

theorem[Menu completeness] Under Assumption (ref), suppose that the canonical menu (ref) is augmented so that $\phi_5$ indexes the full conditional cumulant sequence $\{\kappa_m\}_{m\ge3}$. Then the following statements are equivalent: \begin{enumerate} • $X$ does not Granger-cause $Y$ in distribution; • $H_{0,k}$ holds for every channel $k\in\mathcal{M}$. \end{enumerate} Consequently, \[ X \text{ Granger-causes } Y \text{ in distribution} \quad\Longleftrightarrow\quad \exists\, k\in\mathcal{M} \text{ such that } H_{0,k} \text{ fails}. \]
proofThe implication $(a)\Rightarrow(b)$ follows directly from Proposition (ref). To establish $(b)\Rightarrow(a)$, suppose that $H_{0,k}$ holds for every channel $k\in\mathcal{M}$. Then the conditional mean, conditional variance, all higher-order conditional cumulants, and the lower- and upper-tail conditional quantiles are invariant to the inclusion of the history of $X$ once the history of $Y$ has been conditioned upon. By Assumption (ref)(i), members of $\mathcal{D}$ are uniquely characterized by precisely these objects. Therefore the conditional laws \[ F_{Y_t\mid\mathcal{I}_{t-1}} \quad\text{and}\quad F_{Y_t\mid\mathcal{I}^{(-X)}_{t-1}} \] coincide $\mathbb{P}$-a.s. Since this equality holds for every $t$, Definition (ref) follows. The final statement is obtained by contraposition.

Theorem (ref) establishes that the collection of channel restrictions constitutes a complete representation of distributional Granger non-causality within the determinacy class $\mathcal{D}$. The result is fundamentally one of identification: the conditional law can be recovered from the collection of channel coordinates, and therefore testing the complete menu is equivalent to testing the distributional null itself.

remark[Finite truncation] The practical implementation of the procedure employs a truncated menu consisting of conditional scale, lower- and upper-tail quantiles, and the third and fourth conditional cumulants. The resulting collection of restrictions no longer provides a complete characterization of the conditional law. Accordingly, failure to reject all channel nulls need not imply distributional non-causality. The finite menu instead defines a restricted null hypothesis corresponding to invariance of the selected coordinates. All finite-sample and asymptotic guarantees developed in Section (ref) are stated relative to this operational null. Theorem (ref) characterizes the limiting case in which the menu is sufficiently rich to recover the full conditional distribution.

The Gaussian boundary as a degenerate menu

The classical theory of linear Granger causality emerges as a special case in which the distributional menu collapses to a single informative coordinate.

proposition[Gaussian collapse of the menu] Suppose that the conditional law of $Y_t$ given $\mathcal{I}_{t-1}$ is Gaussian, with conditional mean affine in the conditioning history and conditional variance invariant to the conditioning history. Then: \begin{enumerate} • Channels $2$--$5$ are non-informative with respect to Granger causality; • The only potentially informative channel is the conditional mean channel $\phi_1$; • The channel null $H_{0,1}$ is equivalent to the linear lag-exclusion restrictions defining classical Granger non-causality. \end{enumerate} Consequently, distributional Granger non-causality, joint validity of the channel nulls, and linear Granger non-causality coincide.
proofUnder the stated assumptions, \[ Y_t = \mu_t + \varepsilon_t, \] where $\varepsilon_t$ is conditionally Gaussian with variance independent of the conditioning history. Hence the conditional variance is invariant by construction and all conditional cumulants of order $m\ge3$ vanish identically. It follows immediately that channels corresponding to conditional scale, tail asymmetry, and higher-order cumulants cannot convey additional predictive information. Therefore $H_{0,2},\ldots,H_{0,5}$ hold automatically. The conditional mean channel remains the sole source through which the history of $X$ may affect the conditional law. Under the affine specification, dependence on the lagged values of $X$ is governed by the coefficients $a_{yx,j}$, so that \[ H_{0,1} \quad\Longleftrightarrow\quad a_{yx,1} = \cdots = a_{yx,p} = 0. \] These are precisely the linear Granger non-causality restrictions. The conclusion then follows from Theorem (ref).

Proposition (ref) shows that the conventional linear Granger framework corresponds to a degenerate setting in which the conditional distribution is fully characterized by its first moment. In such environments, testing distributional causality reduces to testing conditional mean predictability. Outside the Gaussian setting, however, predictive content may enter through conditional scale, tail behavior, or higher-order distributional characteristics, necessitating a broader collection of channel restrictions.

Adaptive Sequential Testing for Distributional Granger Causality

This section introduces the adaptive testing procedure. The objective is to test the composite null of distributional Granger non-causality using the finite channel menu developed in Section (ref). The procedure allocates a finite testing budget across channels, updates inference based on previously observed outcomes, and terminates once sufficient evidence against the null has been accumulated or the available budget has been exhausted.

The analysis proceeds in two stages. First, a familywise error guarantee is established for arbitrary adaptive channel-selection rules. Second, a particular selection policy is shown to attain asymptotically optimal power relative to an infeasible oracle benchmark.

Adaptive testing environment

Fix a sample $\{(X_t,Y_t)\}_{t=1}^{T}$ and the truncated menu $\mathcal{M}=\{1,\dots,K\}$ introduced in Remark (ref). Associated with each channel $k\in\mathcal{M}$ is a test statistic $S_k$ and a corresponding $p$-value $P_k$.

The testing procedure operates sequentially. At each stage a channel is selected, its associated hypothesis is evaluated, and the resulting information is incorporated into subsequent selection decisions. Let \[ \mathcal{R}_r\subseteq\mathcal{M} \] denote the set of channels examined by stage $r$.

The information available after stage $r$ is summarized by the filtration \[ \mathcal{F}_r = \sigma \Big( \mathcal{R}_r, \{P_k:k\in\mathcal{R}_r\}, W_r, G_r \Big), \] where $W_r$ denotes the remaining testing wealth and $G_r$ is an auxiliary diagnostic statistic defined below.

\paragraph{Diagnostic signal.}

To guide channel selection, the procedure computes

equation[equation omitted — 115 chars of source]

where $\widehat{\kappa}_3$ and $\widehat{\kappa}_4$ denote the empirical third and fourth cumulants of the fitted VAR residuals.

The quantity $G_r$ provides a low-dimensional summary of departures from conditional Gaussianity. Under Proposition (ref), values of $G_r$ close to zero indicate that predictive content is likely concentrated in the conditional-mean channel, whereas larger values suggest the potential relevance of scale, tail, or higher-order channels.

\paragraph{Adaptive selection rule.}

At each stage the procedure selects an action \[ A_r \in \bigl(\mathcal{M}\setminus\mathcal{R}_r\bigr) \cup \{\textsc{stop}\}. \]

A selection policy is a measurable mapping \[ \pi: \mathcal{F}_r \longrightarrow \bigl(\mathcal{M}\setminus\mathcal{R}_r\bigr) \cup \{\textsc{stop}\}. \]

The policy may be deterministic or randomized and may depend arbitrarily on previously observed outcomes.

\paragraph{Testing wealth.}

Inference is conducted using an alpha-investing mechanism. Let the initial wealth satisfy \[ W_0=\alpha. \]

Before testing channel $k$, the procedure commits a testing level $\alpha_k\le W_r$. The null hypothesis $H_{0,k}$ is rejected whenever \[ P_k\le\alpha_k. \]

Testing wealth evolves according to

equation[equation omitted — 112 chars of source]

where $\psi\le\alpha$ denotes a fixed reward parameter FosterStine2008.

The procedure rejects the composite null of distributional Granger non-causality whenever at least one channel null is rejected before termination.

Familywise error control

The first result establishes that inferential validity is invariant to the adaptive channel-selection mechanism.

assumption[Valid channel $p$-values] For each channel $k$, under its channel null $H_{0,k}$ the $p$-value $P_k$ is conditionally super-uniform: \[ \mathbb{P}(P_k\le u \mid \mathcal{F}_{r^-}) \le u, \qquad u\in[0,1], \] where $\mathcal{F}_{r^-}$ denotes the information available immediately prior to selection of channel $k$.

Assumption (ref) is a high-level inferential condition. Theorem (ref) in Appendix (ref) derives this property from the primitive assumptions on the data-generating process together with the circular-block permutation mechanism of Definition (ref).

theorem[Policy-invariant familywise error control] Let the global null be distributional Granger non-causality (Definition (ref)). Under Assumption (ref), the alpha-investing procedure (ref) with $W_0=\alpha$ and $\psi\le\alpha$ satisfies \[ \mathbb{P} \Big( \text{reject the global null} \Big) \le \alpha \] for every admissible selection policy $\pi$.
proofThe proof is based on a test-supermartingale construction. Under the global null, every channel null holds by Proposition (ref). Conditional super-uniformity therefore applies to every selected channel. A nonnegative supermartingale can be constructed from the sequence of committed testing levels, and this process exceeds the threshold $1/\alpha$ whenever the global null is rejected. Application of Ville's inequality at the stopping time induced by the testing procedure yields the desired bound. Since conditional super-uniformity is imposed with respect to the filtration generated by the adaptive selection mechanism, the argument remains valid irrespective of the policy used to select channels. Complete details are provided in Appendix (ref).

Theorem (ref) establishes that familywise error control is invariant to the adaptive channel-selection mechanism. Inferential validity is therefore separated from the design of the selection rule: any admissible policy satisfying the filtration-adaptedness requirement inherits the same finite-sample error guarantee.

Asymptotic efficiency of adaptive channel selection

The familywise error guarantee established in Theorem (ref) holds uniformly over all admissible selection policies. The role of the selection rule is therefore not inferential validity but statistical efficiency. This subsection studies the power properties of an adaptive allocation rule and compares its performance with an infeasible oracle benchmark.

\paragraph{Local alternatives and channel informativeness.}

Consider a sequence of local alternatives indexed by a signal-strength parameter $\delta\ge0$. Let \[ \lambda_k(\delta,T) \] denote the non-centrality parameter associated with channel $k$ at sample size $T$. The collection \[ \mathcal{A} = \{k\in\mathcal{M}:\lambda_k(\delta,T)>0\} \] defines the active channel set.

By Theorem (ref), whenever distributional Granger causality is present, at least one channel belongs to $\mathcal{A}$. The objective of the adaptive procedure is therefore to allocate testing resources toward the most informative active channel.

\paragraph{Oracle allocation.}

Define \[ k^\star = \arg\max_{k\in\mathcal{M}} \lambda_k(\delta,T), \] and let \[ \beta^\star(\delta,T) = \Pr\!\big( P_{k^\star}\le\alpha \big) \] denote the power of an infeasible oracle that knows the identity of the most informative channel ex ante.

The quantity $\beta^\star(\delta,T)$ serves as an upper benchmark for all procedures restricted to the channel menu. Efficiency is evaluated through the power gap \[ \mathcal{E}_\pi(\delta,T) = \beta^\star(\delta,T) - \beta_\pi(\delta,T), \] where $\beta_\pi(\delta,T)$ denotes the power attained under selection policy $\pi$.

assumption[Identification and separation] \begin{enumerate} • The diagnostic statistic $G_r$ is informative in the sense that \[ \mathbb{E}[G_r] \] is weakly increasing in the non-centralities associated with the scale, tail, and cumulant channels. • There exists a unique maximizer \[ k^\star = \arg\max_{k\in\mathcal{M}} \lambda_k, \] satisfying \[ \Delta = \lambda_{k^\star} - \max_{k\neq k^\star}\lambda_k > 0. \] \end{enumerate} Furthermore, the channel statistics satisfy the concentration bounds established in Theorem (ref).
theorem[Asymptotic efficiency relative to the oracle allocation] Consider the adaptive upper-confidence-bound allocation rule that selects, at stage $r$, the channel maximizing \[ \widehat{\lambda}_{k,r} + c\,U_{k,r}(G_r), \] where $\widehat{\lambda}_{k,r}$ is the current estimate of the channel non-centrality, $U_{k,r}(G_r)$ is a confidence radius, and $c>0$ is a fixed exploration constant. Under Assumptions (ref) and (ref), \[ \beta^\star(\delta,T) - \beta_\pi(\delta,T) \le \frac{C\log B} {\Delta^2\sqrt{T}} + o(1), \] where $C<\infty$ is a constant independent of $B$ and $T$. Consequently, \[ \lim_{T\rightarrow\infty} \Big( \beta^\star(\delta,T) - \beta_\pi(\delta,T) \Big) = 0. \] That is, the adaptive procedure achieves asymptotically the same power as the infeasible oracle allocation.
proofThe adaptive allocation problem may be represented as a finite-armed stochastic experimentation problem in which the channel non-centralities $\{\lambda_k\}_{k=1}^{K}$ constitute the arm means. By Theorem (ref), the associated statistics satisfy exponential concentration inequalities. Standard upper-confidence-bound arguments imply that the expected number of allocations to any suboptimal channel satisfies \[ \mathbb{E}[n_k(B)] = O\!\left( \frac{\log B}{\Delta^2} \right). \] Hence the probability of allocating to the oracle channel converges to unity at rate \[ 1 - O\!\left( \frac{\log B}{\Delta^2 B} \right). \] Combining this allocation result with the monotonicity of the power function in the non-centrality parameter and the $\sqrt{T}$-consistency established in Theorem (ref) yields the stated bound.

Theorem (ref) establishes an asymptotic efficiency result. Although the adaptive procedure does not observe the active channel, the power loss relative to the infeasible oracle allocation vanishes asymptotically. Combined with Theorem (ref), this yields a separation between validity and efficiency: familywise error control holds uniformly over admissible policies, while suitably designed policies achieve asymptotically optimal allocation of testing resources across channels.

Algorithmic implementation

The implementation proceeds as follows.

enumerate• Compute the VAR residuals and evaluate the diagnostic statistic $G_r$ in (ref). • Initialize testing wealth at $W_0=\alpha$. • Sequentially select channels according to the allocation rule, commit testing levels $\alpha_k$, evaluate channel-specific $p$-values, and update wealth using (ref). • Terminate when a channel null is rejected, testing wealth is exhausted, or all admissible channels have been examined. • Report the global decision together with the channel(s) responsible for rejection.

Section (ref) evaluates finite-sample performance through Monte Carlo experiments under a range of data-generating mechanisms corresponding to mean, scale, tail, nonlinear, and state-dependent alternatives. Particular attention is devoted to the finite-sample validity predicted by Theorem (ref) and the efficiency properties characterized in Theorem (ref).

Finite-Sample Performance of Adaptive Channel Selection

This section examines the finite-sample properties of the adaptive channel-selection procedure developed in Section (ref). The simulations are designed to evaluate the two principal theoretical results established earlier. First, Theorem (ref) predicts familywise error control under arbitrary adaptive selection policies. Second, Theorem (ref) implies that the power of the adaptive procedure should converge toward that of an infeasible oracle allocation as sample size increases.

The experimental design isolates distinct forms of predictive dependence corresponding to individual coordinates of the channel menu. This permits direct evaluation of the procedure's ability to identify the relevant source of distributional dependence and allocate testing resources accordingly.

Data-generating processes

Four channel-specific data-generating mechanisms are considered. Each design is indexed by a signal-strength parameter $s\ge0$, where $s=0$ corresponds to the global null of distributional Granger non-causality.

For all designs, \[ X_t = 0.5X_{t-1} + \varepsilon_t^x, \] and \[ Y_t = 0.3Y_{t-1} + u_t. \]

The specifications differ only in the mechanism through which lagged values of $X_t$ influence the conditional distribution of $Y_t$.

enumerate• Conditional-mean alternative \[ Y_t = 0.3Y_{t-1} + sX_{t-1} + \varepsilon_t^y. \] Predictive content enters exclusively through the conditional mean. The mean channel is therefore the unique active coordinate. • Conditional-scale alternative \[ Y_t = 0.3Y_{t-1} + \sqrt{1+sX_{t-1}^{2}}\,\varepsilon_t^y. \] The conditional mean remains unchanged, while dependence enters through conditional scale and tail behavior. • Nonlinear alternative \[ Y_t = 0.3Y_{t-1} + s(X_{t-1}^{2}-1) + \varepsilon_t^y. \] Dependence is nonlinear and cannot be recovered through linear projection under symmetry. • \textbf{State-dependent scale alternative} The innovation variance is multiplied by \[ 1+s\mathbf{1}\{X_{t-1}>0\}, \] creating a regime-dependent scale effect without a structural conditional-mean component.

Innovations are generated from either a Gaussian distribution or a standardized skew-$t$ distribution. Sample sizes are $T\in\{250,500\}$ and each design is replicated $R$ times.

For every channel, inference is calibrated through the circular permutation procedure of Definition (ref). The permutation destroys predictive dependence between $X_t$ and $Y_t$ while preserving the serial dependence structure of $X_t$. Consequently, the simulation design directly evaluates the conditional super-uniformity property underlying Theorem (ref).

The adaptive procedure is compared with four benchmark methods:

enumerate• channel-specific fixed tests; • a naive multiple-testing procedure that rejects whenever any channel-specific $p$-value falls below $\alpha$; • a Bonferroni-adjusted multiple-testing procedure; • an infeasible oracle allocation that concentrates all testing resources on the truly active channel.

Familywise error control

Figure (ref) reports empirical rejection frequencies under the global null ($s=0$). Since all four designs coincide under the null hypothesis, rejection frequencies should be invariant across scenarios up to simulation uncertainty.

The results strongly support Theorem (ref). Across all designs and sample sizes, the adaptive procedure maintains rejection frequencies at or below the nominal significance level $\alpha=0.05$. In contrast, the naive multiple-testing procedure exhibits substantial size distortions, with rejection frequencies increasing from approximately $0.13$ at $T=250$ to approximately $0.16$ at $T=500$.

The comparison illustrates the distinction between adaptive allocation and unrestricted multiple testing. The adaptive procedure preserves familywise error control through the alpha-investing mechanism, whereas the naive procedure accumulates rejection probability across channels without accounting for multiplicity. The Bonferroni procedure also controls familywise error, although at the cost of reduced power due to its conservative adjustment.

figure[figure omitted — 486 chars of source]

Power and oracle efficiency

Figure (ref) reports rejection frequencies as a function of signal strength $s$ under skew-$t$ innovations.

Several patterns emerge consistently across designs. First, power increases monotonically with both signal strength and sample size. Second, the adaptive procedure closely tracks the oracle benchmark throughout the parameter space. Third, procedures restricted to a single channel perform well only when the active source of dependence coincides with their maintained specification.

Under the conditional-mean alternative (S1), the adaptive procedure allocates testing effort primarily toward the location channel and achieves power nearly identical to that of the oracle benchmark. Under the conditional-scale and state-dependent alternatives (S2 and S4), predictive content resides principally in scale and tail coordinates. In these environments, mean-based procedures exhibit little power, whereas the adaptive procedure reallocates testing effort toward the relevant channels and recovers most of the oracle benchmark.

The nonlinear alternative (S3) produces a similar pattern. The adaptive procedure successfully identifies the informative nonlinear coordinate and achieves rejection frequencies nearly indistinguishable from those of the oracle allocation.

Overall, the results indicate that adaptive allocation substantially mitigates the power losses associated with channel misspecification while maintaining valid familywise error control.

figure[figure omitted — 470 chars of source]

Table (ref) summarizes familywise error rates under the null and rejection frequencies under the strongest alternative considered. The results reinforce the graphical evidence. The adaptive procedure maintains familywise error control throughout while achieving power levels close to those of the oracle allocation.

table[table omitted — 1,479 chars of source]

Convergence to the oracle benchmark

Theorem (ref) predicts that the power gap between the adaptive procedure and the oracle allocation should diminish with sample size. Figure (ref) evaluates this prediction directly.

For each design, the figure reports \[ \beta^\star(\delta,T) - \beta_\pi(\delta,T), \] the difference between oracle power and the power of the adaptive procedure.

Two features are apparent. First, the magnitude of the gap decreases as sample size increases from $T=250$ to $T=500$. This pattern is consistent with the asymptotic efficiency result of Theorem (ref): larger samples improve estimation of channel informativeness, allowing testing resources to be allocated more effectively.

Second, the gap is occasionally negative under the state-dependent scale design (S4). In this environment predictive content is distributed across multiple coordinates of the channel menu. Since the adaptive procedure may accumulate evidence across several informative channels, it can outperform a benchmark restricted to a single channel. This phenomenon highlights a limitation of the oracle benchmark rather than a violation of the theorem.

figure[figure omitted — 441 chars of source]

Summary

The Monte Carlo results provide strong support for the theoretical analysis. Under the global null, the adaptive procedure maintains familywise error control across all designs, in accordance with Theorem (ref). Under channel-specific alternatives, rejection frequencies closely track those of the oracle allocation, and the efficiency gap decreases with sample size, consistent with Theorem (ref).

Taken together, the results indicate that adaptive channel selection provides a practical mechanism for detecting distributional Granger causality when the relevant source of dependence is unknown ex ante. The procedure preserves valid inference while allocating testing resources toward the most informative coordinates of the conditional distribution.

Conclusion

This paper develops a framework for testing distributional Granger causality when predictive dependence may arise through multiple features of the conditional distribution. Outside the Gaussian setting, predictive content need not be confined to the conditional mean; it may instead appear through conditional scale, tail behavior, asymmetry, or higher-order distributional characteristics. Consequently, no single Granger-type test can provide a complete characterization of predictive dependence.

The analysis proceeds by decomposing distributional causality into a collection of channel-specific restrictions defined on functionals of the conditional distribution. Under a determinacy condition, the resulting channel menu is complete in the sense that distributional Granger non-causality is equivalent to the joint validity of the channel restrictions. This characterization converts an infinite-dimensional hypothesis concerning conditional laws into a finite collection of testable restrictions while preserving identification of the underlying causal object.

Building on this representation, the paper develops an adaptive sequential testing procedure for channel selection. The procedure combines channel-specific inference with an alpha-investing mechanism that permits data-dependent allocation of testing resources while maintaining familywise error control. The theoretical analysis establishes two principal results. First, familywise error control is invariant to the channel-selection policy, yielding finite-sample validity under arbitrary admissible adaptive rules. Second, a confidence-bound allocation rule achieves asymptotic efficiency relative to an infeasible oracle benchmark, implying that the power loss attributable to uncertainty regarding the active channel vanishes asymptotically.

The Monte Carlo evidence is consistent with these theoretical predictions. Across a broad class of data-generating processes involving location, scale, nonlinear, and regime-dependent forms of dependence, the proposed procedure maintains familywise error control while attaining power levels that closely track those of the oracle allocation. The results indicate that adaptive channel selection can substantially reduce the efficiency losses associated with channel misspecification without sacrificing inferential validity.

Several extensions merit further investigation. One direction is the development of network-based versions of the procedure in which testing resources are allocated jointly across channels and graph structures to detect distributional spillovers in high-dimensional systems. A second direction is the incorporation of more general adaptive allocation mechanisms, including reinforcement-learning-based policies, within the same error-control framework. A third direction is the enrichment of the channel menu through frequency-domain, option-implied, or other distribution-sensitive functionals that capture additional dimensions of predictive dependence. Because the identification and error-control results are formulated at the level of the channel menu itself, these extensions can be accommodated without altering the fundamental structure of the framework.

More broadly, the results suggest that distributional causality is most naturally viewed as a problem of adaptive inference over a collection of complementary predictive channels. The framework developed here provides a unified approach to identification, inference, and efficient testing in such environments, extending the classical theory of Granger causality beyond the conditional-mean paradigm while preserving rigorous finite-sample and asymptotic guarantees.