Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
54,854 characters · 24 sections · 26 citation commands
Distributional Granger Causality: Identification, Sequential Inference, and Adaptive Testing
\onehalfspacing
Granger causality is among the most widely used tools for studying predictive relationships in time series. In its classical form, the concept is operationalized through incremental forecasting ability: a process $X_t$ is said to Granger-cause a process $Y_t$ if past values of $X_t$ improve forecasts of $Y_t$ beyond the information contained in the past of $Y_t$ alone Granger1969,Sims1972. Under linear dynamics and Gaussian innovations, this notion admits a particularly simple characterization through lag-exclusion restrictions in vector autoregressions Geweke1982,Geweke1984. In such environments, the conditional distribution is fully summarized by its first two moments, and causal inference reduces to a problem of conditional-mean predictability.
Outside the Gaussian setting, however, predictive dependence need not be confined to the conditional mean. A predictor may influence the conditional variance of future outcomes, alter tail probabilities, modify higher-order moments, or affect other features of the conditional distribution without changing expected values. This observation has motivated a large literature on alternative notions of Granger causality, including causality in variance CheungNg1996, causality in risk and tail events HongLiuWang2009,WhiteKimManganelli2015, causality in quantiles JeongHardleSong2012,SongTaamouti2021, nonlinear and nonparametric causality HiemstraJones1994,DiksPanchenko2006,NishiyamaHitomiKawasakiJeong2011, copula-based dependence measures BouezmarniRomboutsTaamouti2012, and information-theoretic approaches such as transfer entropy BarnettSethBossomaier2009.
The resulting literature provides a rich collection of channel-specific tests but leaves open a fundamental question. If predictive dependence may arise through multiple dimensions of the conditional distribution, how should inference be conducted when the relevant channel is unknown ex ante? In empirical practice, researchers often evaluate several causality tests and interpret the resulting collection of $p$-values informally. This approach faces two difficulties. First, simultaneous consideration of multiple channels generates a multiple-testing problem, potentially leading to substantial distortions in familywise error rates. Second, there is generally no principled rule for allocating finite inferential resources across competing tests or for determining which channel should receive the greatest attention.
This paper develops a unified framework for distributional Granger causality that addresses both issues. The starting point is an identification result. Rather than viewing existing causality tests as competing methodologies, the paper interprets them as measuring distinct coordinates of a common object: the conditional distribution of future outcomes. Distributional Granger non-causality is characterized through a finite collection of channel-specific restrictions corresponding to conditional location, scale, tail behavior, and higher-order distributional features. Under suitable determinacy conditions, these restrictions are shown to be complete in the sense that distributional Granger non-causality holds if and only if every channel-specific restriction is satisfied. This representation transforms an infinite-dimensional hypothesis concerning conditional distributions into a finite collection of testable components while preserving identification of the underlying causal object.
Building on this representation, the paper develops an adaptive sequential testing procedure for channel selection. The procedure allocates a finite testing budget across channels, updates allocation decisions using previously observed outcomes, and terminates once sufficient evidence against the null has been accumulated or the testing budget has been exhausted. The inferential framework combines channel-specific hypothesis tests with an alpha-investing mechanism drawn from the sequential multiple-testing literature FosterStine2008,AharoniRosset2014. This construction permits data-dependent channel selection while maintaining rigorous control of familywise error rates.
The theoretical contribution consists of three results. First, a completeness theorem establishes that the proposed channel menu fully characterizes distributional Granger non-causality. Second, a policy-invariant familywise error theorem shows that inferential validity is preserved under arbitrary admissible channel-selection rules. The result separates validity from allocation: any adaptive policy satisfying the filtration-adaptedness requirements inherits the same finite-sample error guarantee. Third, an asymptotic efficiency theorem demonstrates that a confidence-bound allocation rule achieves power equivalent to that of an infeasible oracle benchmark in the limit. Consequently, the power loss associated with uncertainty regarding the active channel vanishes asymptotically.
An important feature of the analysis is that the theoretical guarantees are derived from primitive conditions on the data-generating process rather than imposed as abstract assumptions. Conditional super-uniformity of the channel-specific $p$-values follows from a circular-block permutation procedure under strict stationarity and suitable mixing conditions. Likewise, the concentration properties required for the efficiency analysis are obtained from Bernstein-type inequalities for dependent processes. These results provide an explicit link between the high-level assumptions underlying the sequential testing framework and standard regularity conditions for weakly dependent time series.
The framework has implications for several areas of applied econometrics. In financial applications, predictive dependence frequently appears through volatility, downside risk, or tail exposure rather than through expected returns. Similar considerations arise in macroeconomics, network analysis, and systemic-risk measurement, where distributional spillovers are often more informative than conditional-mean effects. By treating causality as a property of the entire conditional distribution rather than a single moment condition, the proposed framework accommodates such settings while preserving formal inferential guarantees.
The remainder of the paper is organized as follows. Section (ref) reviews the related literature. Section (ref) develops the distributional-causality framework and establishes the completeness of the channel representation. Section (ref) introduces the adaptive sequential testing procedure and presents the validity and efficiency results. Section (ref) reports Monte Carlo evidence on finite-sample performance. Section (ref) concludes. Proofs and additional technical results are collected in the Online Appendix.
This paper contributes to three strands of the econometrics literature: distributional approaches to Granger causality, sequential multiple-testing procedures with error-control guarantees, and adaptive allocation methods for statistical experimentation.
The classical formulation of Granger causality characterizes predictive dependence through incremental linear forecasting ability Granger1969,Sims1972,Geweke1982,Geweke1984. Under Gaussianity and linear dynamics, this characterization is sufficient because the conditional distribution is fully summarized by its first two moments. Outside that setting, however, predictive content may arise through conditional scale, tail behavior, asymmetry, or other higher-order distributional features that are not captured by conditional-mean restrictions alone.
A substantial literature has therefore developed tests targeting specific dimensions of predictive dependence. Examples include causality in variance CheungNg1996, causality in risk and tail events HongLiuWang2009,WhiteKimManganelli2015, causality in distribution CandelonTokpavi2016, causality in quantiles JeongHardleSong2012,SongTaamouti2021, nonlinear and nonparametric causality HiemstraJones1994,DiksPanchenko2006,NishiyamaHitomiKawasakiJeong2011, and copula-based approaches to conditional dependence BouezmarniRomboutsTaamouti2012. Information-theoretic measures such as transfer entropy provide additional tools for detecting departures from linear predictability and coincide with classical Granger causality under Gaussianity BarnettSethBossomaier2009.
Related developments arise in the literature on identification under non-Gaussianity, where higher-order moments and distributional features play a central role in recovering structural relationships ShimizuHoyerHyvarinenKerminen2006,LanneMeitzSaikkonen2017,GourierouxMonfortRenne2017,MontielOleaPlagborgMollerQian2022. Higher-order cumulants likewise provide a natural representation of departures from Gaussianity and characterize important dimensions of distributional shape ArevalilloNavarro2026. Recent work has also continued to refine inference for predictive regressions and Granger-causality tests in nonstandard environments, including settings involving boundary parameters and weak identification CavaliereGeorgiev2020,CavaliereGeorgievZanelli2025.
The present paper differs from this literature in its objective. Rather than proposing a new channel-specific test, it studies how existing tests can be combined within a unified framework. The contribution is an identification result showing that, under suitable determinacy conditions, a sufficiently rich collection of channel-specific restrictions provides a complete characterization of distributional Granger non-causality.
Testing multiple dimensions of predictive dependence naturally raises a multiple-testing problem. Classical approaches control familywise error rates through fixed corrections such as Bonferroni procedures, but these methods are static and often conservative. A complementary literature studies sequential and online testing procedures that allocate significance levels adaptively as hypotheses are examined FosterStine2008,AharoniRosset2014,JavanmardMontanari2018,RamdasZrnicWainwrightJordan2018.
Closely related developments arise in the literature on anytime-valid inference, test martingales, and e-processes, which provide inferential guarantees that remain valid under optional stopping and adaptive continuation rules RamdasGrunwaldVovkShafer2023,GrunwaldHeideKoolen2024. These methods emphasize the construction of inferential procedures whose validity is preserved under data-dependent experimentation.
The present paper applies these ideas to distributional causality testing. The distinguishing feature of the setting considered here is that the sequence of hypotheses is not exogenously given but is selected adaptively from a menu of candidate channels. The resulting policy-invariant familywise error theorem establishes that valid inference is preserved under arbitrary admissible channel-selection rules. In this sense, the paper extends sequential testing methods to an environment in which both testing and selection are data dependent.
The allocation of finite testing resources across competing hypotheses is closely related to the literature on sequential experimentation and adaptive allocation. Multi-armed bandit methods provide a canonical framework for balancing exploration and exploitation and have generated a large body of results concerning regret minimization and efficient learning AuerCesaBianchiFischer2002,LattimoreSzepesvari2020,KaufmannCappeGarivier2016.
The problem considered here differs from the standard bandit setting in both objective and interpretation. The goal is not reward maximization per se, but efficient allocation of inferential effort across competing dimensions of predictive dependence. The objects being learned are channel-specific measures of statistical informativeness, represented through the non-centrality parameters of the underlying tests.
This perspective connects with a broader movement toward adaptive and data-driven procedures in econometrics ChernozhukovEtAl2018,FarrellLiangMisra2021. Recent methodological discussions have emphasized the growing interaction between econometric theory, machine learning, and artificial intelligence GuggenbergerSuSun2026. The contribution of the present paper is to embed adaptive allocation within a formal inferential framework. The resulting procedure combines three features that are typically studied separately: identification of the causal object, finite-sample error control, and asymptotically efficient allocation of testing resources. The asymptotic oracle-efficiency result demonstrates that adaptation can improve power without compromising the validity guarantees established by the sequential testing framework.
Let $\{(X_t,Y_t)\}_{t\in\mathbb{Z}}$ be a strictly stationary bivariate stochastic process defined on the probability space $(\Omega,\mathcal{F},\mathbb{P})$. Define the information sets \[ \mathcal{I}_{t-1} := \sigma\big(\{X_{t-j},Y_{t-j}\}_{j\ge 1}\big), \qquad \mathcal{I}^{(-X)}_{t-1} := \sigma\big(\{Y_{t-j}\}_{j\ge 1}\big). \] Let $F_{Y_t\mid\mathcal{J}}(\cdot)$ denote a regular conditional distribution of $Y_t$ given a sub-$\sigma$-field $\mathcal{J}$, and let $Q_{Y_t}(\tau\mid\mathcal{J})$ denote the associated conditional $\tau$-quantile. Unless otherwise stated, all equalities and inequalities involving conditional objects hold $\mathbb{P}$-almost surely, with continuity-point qualifications imposed where required.
The subsequent asymptotic analysis relies on a collection of primitive conditions governing dependence, moments, and conditional smoothness. These assumptions are standard in the nonparametric and resampling literature for weakly dependent stochastic processes and are sufficient to establish the higher-level regularity conditions employed in the finite-sample size and power results developed in Section (ref). Formal derivations are provided in Appendix (ref).
The object of interest is the conditional distribution of $Y_t$ given the available information. The strongest form of predictive irrelevance is defined through equality of conditional laws.
Definition (ref) characterizes predictive irrelevance at the level of the entire conditional distribution. It therefore excludes any incremental predictive contribution of the history of $X$ once the history of $Y$ has been conditioned upon. Since this condition is infinite-dimensional, direct implementation is generally infeasible. The framework developed below replaces the distributional null by a collection of lower-dimensional restrictions defined through measurable functionals of the conditional law.
Let $\mathcal{M}=\{1,\ldots,K\}$ index a finite collection of measurable functionals of the conditional distribution. For each channel $k\in\mathcal{M}$, let $\phi_k$ denote a functional defined on the space of conditional laws. The corresponding channel null hypothesis is \[ H_{0,k}: \quad \phi_k\!\left(F_{Y_t\mid\mathcal{I}_{t-1}}\right) = \phi_k\!\left(F_{Y_t\mid\mathcal{I}^{(-X)}_{t-1}}\right) \quad \mathbb{P}\text{-a.s.}, \qquad \forall t. \]
The canonical menu considered throughout the paper is given by
The first four channels correspond to conditional location, scale, and tail behavior. The fifth channel captures higher-order distributional features through conditional skewness- and kurtosis-related cumulants. Each channel induces a testable restriction and may be associated with an established inferential procedure. Throughout, the testing methodology is treated as given; the emphasis is on the logical relationship between the collection of channel restrictions and distributional Granger non-causality.
The channel hypotheses are nested within the distributional null.
Proposition (ref) establishes that distributional Granger non-causality implies the validity of every channel restriction. Consequently, rejection of any channel null is sufficient to reject distributional non-causality. The converse implication, however, requires additional structure. Specifically, the collection of channel functionals must be sufficiently informative to identify the conditional law.
To establish the converse implication, the collection of channel functionals must uniquely characterize the conditional law within an admissible class of distributions. The following assumption formalizes this requirement.
Assumption (ref) is an identification condition. It guarantees that equality of the functionals defining the menu implies equality of the underlying conditional distributions. Consequently, the finite collection of channel restrictions can be used to recover an infinite-dimensional statement concerning conditional laws.
Theorem (ref) establishes that the collection of channel restrictions constitutes a complete representation of distributional Granger non-causality within the determinacy class $\mathcal{D}$. The result is fundamentally one of identification: the conditional law can be recovered from the collection of channel coordinates, and therefore testing the complete menu is equivalent to testing the distributional null itself.
The classical theory of linear Granger causality emerges as a special case in which the distributional menu collapses to a single informative coordinate.
Proposition (ref) shows that the conventional linear Granger framework corresponds to a degenerate setting in which the conditional distribution is fully characterized by its first moment. In such environments, testing distributional causality reduces to testing conditional mean predictability. Outside the Gaussian setting, however, predictive content may enter through conditional scale, tail behavior, or higher-order distributional characteristics, necessitating a broader collection of channel restrictions.
This section introduces the adaptive testing procedure. The objective is to test the composite null of distributional Granger non-causality using the finite channel menu developed in Section (ref). The procedure allocates a finite testing budget across channels, updates inference based on previously observed outcomes, and terminates once sufficient evidence against the null has been accumulated or the available budget has been exhausted.
The analysis proceeds in two stages. First, a familywise error guarantee is established for arbitrary adaptive channel-selection rules. Second, a particular selection policy is shown to attain asymptotically optimal power relative to an infeasible oracle benchmark.
Fix a sample $\{(X_t,Y_t)\}_{t=1}^{T}$ and the truncated menu $\mathcal{M}=\{1,\dots,K\}$ introduced in Remark (ref). Associated with each channel $k\in\mathcal{M}$ is a test statistic $S_k$ and a corresponding $p$-value $P_k$.
The testing procedure operates sequentially. At each stage a channel is selected, its associated hypothesis is evaluated, and the resulting information is incorporated into subsequent selection decisions. Let \[ \mathcal{R}_r\subseteq\mathcal{M} \] denote the set of channels examined by stage $r$.
The information available after stage $r$ is summarized by the filtration \[ \mathcal{F}_r = \sigma \Big( \mathcal{R}_r, \{P_k:k\in\mathcal{R}_r\}, W_r, G_r \Big), \] where $W_r$ denotes the remaining testing wealth and $G_r$ is an auxiliary diagnostic statistic defined below.
\paragraph{Diagnostic signal.}
To guide channel selection, the procedure computes
where $\widehat{\kappa}_3$ and $\widehat{\kappa}_4$ denote the empirical third and fourth cumulants of the fitted VAR residuals.
The quantity $G_r$ provides a low-dimensional summary of departures from conditional Gaussianity. Under Proposition (ref), values of $G_r$ close to zero indicate that predictive content is likely concentrated in the conditional-mean channel, whereas larger values suggest the potential relevance of scale, tail, or higher-order channels.
\paragraph{Adaptive selection rule.}
At each stage the procedure selects an action \[ A_r \in \bigl(\mathcal{M}\setminus\mathcal{R}_r\bigr) \cup \{\textsc{stop}\}. \]
A selection policy is a measurable mapping \[ \pi: \mathcal{F}_r \longrightarrow \bigl(\mathcal{M}\setminus\mathcal{R}_r\bigr) \cup \{\textsc{stop}\}. \]
The policy may be deterministic or randomized and may depend arbitrarily on previously observed outcomes.
\paragraph{Testing wealth.}
Inference is conducted using an alpha-investing mechanism. Let the initial wealth satisfy \[ W_0=\alpha. \]
Before testing channel $k$, the procedure commits a testing level $\alpha_k\le W_r$. The null hypothesis $H_{0,k}$ is rejected whenever \[ P_k\le\alpha_k. \]
Testing wealth evolves according to
where $\psi\le\alpha$ denotes a fixed reward parameter FosterStine2008.
The procedure rejects the composite null of distributional Granger non-causality whenever at least one channel null is rejected before termination.
The first result establishes that inferential validity is invariant to the adaptive channel-selection mechanism.
Assumption (ref) is a high-level inferential condition. Theorem (ref) in Appendix (ref) derives this property from the primitive assumptions on the data-generating process together with the circular-block permutation mechanism of Definition (ref).
Theorem (ref) establishes that familywise error control is invariant to the adaptive channel-selection mechanism. Inferential validity is therefore separated from the design of the selection rule: any admissible policy satisfying the filtration-adaptedness requirement inherits the same finite-sample error guarantee.
The familywise error guarantee established in Theorem (ref) holds uniformly over all admissible selection policies. The role of the selection rule is therefore not inferential validity but statistical efficiency. This subsection studies the power properties of an adaptive allocation rule and compares its performance with an infeasible oracle benchmark.
\paragraph{Local alternatives and channel informativeness.}
Consider a sequence of local alternatives indexed by a signal-strength parameter $\delta\ge0$. Let \[ \lambda_k(\delta,T) \] denote the non-centrality parameter associated with channel $k$ at sample size $T$. The collection \[ \mathcal{A} = \{k\in\mathcal{M}:\lambda_k(\delta,T)>0\} \] defines the active channel set.
By Theorem (ref), whenever distributional Granger causality is present, at least one channel belongs to $\mathcal{A}$. The objective of the adaptive procedure is therefore to allocate testing resources toward the most informative active channel.
\paragraph{Oracle allocation.}
Define \[ k^\star = \arg\max_{k\in\mathcal{M}} \lambda_k(\delta,T), \] and let \[ \beta^\star(\delta,T) = \Pr\!\big( P_{k^\star}\le\alpha \big) \] denote the power of an infeasible oracle that knows the identity of the most informative channel ex ante.
The quantity $\beta^\star(\delta,T)$ serves as an upper benchmark for all procedures restricted to the channel menu. Efficiency is evaluated through the power gap \[ \mathcal{E}_\pi(\delta,T) = \beta^\star(\delta,T) - \beta_\pi(\delta,T), \] where $\beta_\pi(\delta,T)$ denotes the power attained under selection policy $\pi$.
Theorem (ref) establishes an asymptotic efficiency result. Although the adaptive procedure does not observe the active channel, the power loss relative to the infeasible oracle allocation vanishes asymptotically. Combined with Theorem (ref), this yields a separation between validity and efficiency: familywise error control holds uniformly over admissible policies, while suitably designed policies achieve asymptotically optimal allocation of testing resources across channels.
The implementation proceeds as follows.
Section (ref) evaluates finite-sample performance through Monte Carlo experiments under a range of data-generating mechanisms corresponding to mean, scale, tail, nonlinear, and state-dependent alternatives. Particular attention is devoted to the finite-sample validity predicted by Theorem (ref) and the efficiency properties characterized in Theorem (ref).
This section examines the finite-sample properties of the adaptive channel-selection procedure developed in Section (ref). The simulations are designed to evaluate the two principal theoretical results established earlier. First, Theorem (ref) predicts familywise error control under arbitrary adaptive selection policies. Second, Theorem (ref) implies that the power of the adaptive procedure should converge toward that of an infeasible oracle allocation as sample size increases.
The experimental design isolates distinct forms of predictive dependence corresponding to individual coordinates of the channel menu. This permits direct evaluation of the procedure's ability to identify the relevant source of distributional dependence and allocate testing resources accordingly.
Four channel-specific data-generating mechanisms are considered. Each design is indexed by a signal-strength parameter $s\ge0$, where $s=0$ corresponds to the global null of distributional Granger non-causality.
For all designs, \[ X_t = 0.5X_{t-1} + \varepsilon_t^x, \] and \[ Y_t = 0.3Y_{t-1} + u_t. \]
The specifications differ only in the mechanism through which lagged values of $X_t$ influence the conditional distribution of $Y_t$.
Innovations are generated from either a Gaussian distribution or a standardized skew-$t$ distribution. Sample sizes are $T\in\{250,500\}$ and each design is replicated $R$ times.
For every channel, inference is calibrated through the circular permutation procedure of Definition (ref). The permutation destroys predictive dependence between $X_t$ and $Y_t$ while preserving the serial dependence structure of $X_t$. Consequently, the simulation design directly evaluates the conditional super-uniformity property underlying Theorem (ref).
The adaptive procedure is compared with four benchmark methods:
Figure (ref) reports empirical rejection frequencies under the global null ($s=0$). Since all four designs coincide under the null hypothesis, rejection frequencies should be invariant across scenarios up to simulation uncertainty.
The results strongly support Theorem (ref). Across all designs and sample sizes, the adaptive procedure maintains rejection frequencies at or below the nominal significance level $\alpha=0.05$. In contrast, the naive multiple-testing procedure exhibits substantial size distortions, with rejection frequencies increasing from approximately $0.13$ at $T=250$ to approximately $0.16$ at $T=500$.
The comparison illustrates the distinction between adaptive allocation and unrestricted multiple testing. The adaptive procedure preserves familywise error control through the alpha-investing mechanism, whereas the naive procedure accumulates rejection probability across channels without accounting for multiplicity. The Bonferroni procedure also controls familywise error, although at the cost of reduced power due to its conservative adjustment.
Figure (ref) reports rejection frequencies as a function of signal strength $s$ under skew-$t$ innovations.
Several patterns emerge consistently across designs. First, power increases monotonically with both signal strength and sample size. Second, the adaptive procedure closely tracks the oracle benchmark throughout the parameter space. Third, procedures restricted to a single channel perform well only when the active source of dependence coincides with their maintained specification.
Under the conditional-mean alternative (S1), the adaptive procedure allocates testing effort primarily toward the location channel and achieves power nearly identical to that of the oracle benchmark. Under the conditional-scale and state-dependent alternatives (S2 and S4), predictive content resides principally in scale and tail coordinates. In these environments, mean-based procedures exhibit little power, whereas the adaptive procedure reallocates testing effort toward the relevant channels and recovers most of the oracle benchmark.
The nonlinear alternative (S3) produces a similar pattern. The adaptive procedure successfully identifies the informative nonlinear coordinate and achieves rejection frequencies nearly indistinguishable from those of the oracle allocation.
Overall, the results indicate that adaptive allocation substantially mitigates the power losses associated with channel misspecification while maintaining valid familywise error control.
Table (ref) summarizes familywise error rates under the null and rejection frequencies under the strongest alternative considered. The results reinforce the graphical evidence. The adaptive procedure maintains familywise error control throughout while achieving power levels close to those of the oracle allocation.
Theorem (ref) predicts that the power gap between the adaptive procedure and the oracle allocation should diminish with sample size. Figure (ref) evaluates this prediction directly.
For each design, the figure reports \[ \beta^\star(\delta,T) - \beta_\pi(\delta,T), \] the difference between oracle power and the power of the adaptive procedure.
Two features are apparent. First, the magnitude of the gap decreases as sample size increases from $T=250$ to $T=500$. This pattern is consistent with the asymptotic efficiency result of Theorem (ref): larger samples improve estimation of channel informativeness, allowing testing resources to be allocated more effectively.
Second, the gap is occasionally negative under the state-dependent scale design (S4). In this environment predictive content is distributed across multiple coordinates of the channel menu. Since the adaptive procedure may accumulate evidence across several informative channels, it can outperform a benchmark restricted to a single channel. This phenomenon highlights a limitation of the oracle benchmark rather than a violation of the theorem.
The Monte Carlo results provide strong support for the theoretical analysis. Under the global null, the adaptive procedure maintains familywise error control across all designs, in accordance with Theorem (ref). Under channel-specific alternatives, rejection frequencies closely track those of the oracle allocation, and the efficiency gap decreases with sample size, consistent with Theorem (ref).
Taken together, the results indicate that adaptive channel selection provides a practical mechanism for detecting distributional Granger causality when the relevant source of dependence is unknown ex ante. The procedure preserves valid inference while allocating testing resources toward the most informative coordinates of the conditional distribution.
This paper develops a framework for testing distributional Granger causality when predictive dependence may arise through multiple features of the conditional distribution. Outside the Gaussian setting, predictive content need not be confined to the conditional mean; it may instead appear through conditional scale, tail behavior, asymmetry, or higher-order distributional characteristics. Consequently, no single Granger-type test can provide a complete characterization of predictive dependence.
The analysis proceeds by decomposing distributional causality into a collection of channel-specific restrictions defined on functionals of the conditional distribution. Under a determinacy condition, the resulting channel menu is complete in the sense that distributional Granger non-causality is equivalent to the joint validity of the channel restrictions. This characterization converts an infinite-dimensional hypothesis concerning conditional laws into a finite collection of testable restrictions while preserving identification of the underlying causal object.
Building on this representation, the paper develops an adaptive sequential testing procedure for channel selection. The procedure combines channel-specific inference with an alpha-investing mechanism that permits data-dependent allocation of testing resources while maintaining familywise error control. The theoretical analysis establishes two principal results. First, familywise error control is invariant to the channel-selection policy, yielding finite-sample validity under arbitrary admissible adaptive rules. Second, a confidence-bound allocation rule achieves asymptotic efficiency relative to an infeasible oracle benchmark, implying that the power loss attributable to uncertainty regarding the active channel vanishes asymptotically.
The Monte Carlo evidence is consistent with these theoretical predictions. Across a broad class of data-generating processes involving location, scale, nonlinear, and regime-dependent forms of dependence, the proposed procedure maintains familywise error control while attaining power levels that closely track those of the oracle allocation. The results indicate that adaptive channel selection can substantially reduce the efficiency losses associated with channel misspecification without sacrificing inferential validity.
Several extensions merit further investigation. One direction is the development of network-based versions of the procedure in which testing resources are allocated jointly across channels and graph structures to detect distributional spillovers in high-dimensional systems. A second direction is the incorporation of more general adaptive allocation mechanisms, including reinforcement-learning-based policies, within the same error-control framework. A third direction is the enrichment of the channel menu through frequency-domain, option-implied, or other distribution-sensitive functionals that capture additional dimensions of predictive dependence. Because the identification and error-control results are formulated at the level of the channel menu itself, these extensions can be accommodated without altering the fundamental structure of the framework.
More broadly, the results suggest that distributional causality is most naturally viewed as a problem of adaptive inference over a collection of complementary predictive channels. The framework developed here provides a unified approach to identification, inference, and efficient testing in such environments, extending the classical theory of Granger causality beyond the conditional-mean paradigm while preserving rigorous finite-sample and asymptotic guarantees.