Every paper in arXiv econ.EM, its bibliography parsed from the LaTeX source, and a weighted citation graph over the whole corpus — linked to author profiles and publication records.
When comparison units may also respond to treatment, panel comparisons reflect both the treatment effect and spillovers. If the interference pattern is unknown, observed outcomes alone do not separate the two. I characterize what can nevertheless be learned from panel outcomes under general restrictions, without requiring an exposure mapping or prior classification of affected donors. The framework scales validity bounds for every convex donor weight by…
Many questions across the sciences take the same form: several coupled series are observed together, and the analyst wants to know not merely that they move together but which one moves first, and how strongly. This paper sets out a complete method built on one organising idea: the direction of a coupled system is exactly the part of its behaviour that changes when the record is played backwards. Tools built on contemporaneous covariance alone (correlati…
I consider an AR($p$) process that is observed every $q$ periods, either as a snapshot (stock variable) or as a sum over the sampling interval (flow variable). Under fairly mild assumptions, I derive the identified set for general lag lengths $p \in \mathbb{N}$ and sampling frequencies $q \in \mathbb{N}$, I bound its cardinality, and I provide a recipe to compute all candidate points and determine their membership in the identified set. My analysis suppo…
We consider the classical additive measurement-error model $X=Y+Z$, where the latent random variable $Y$ has unknown distribution $F_Y$ and the error $Z$ has a known distribution. We develop direct estimators for three functionals of $F_Y$: (i) $F_Y(x)$ at continuity points; (ii) interval probabilities $F_Y(y)-F_Y(x)$ when $x<y$ are continuity points; and (iii) the size of a jump at a prespecified discontinuity. We derive non-asymptotic bias and variance…
Factor-MIDAS regressions forecast a low-frequency target by extracting common factors from a large panel of high-frequency predictors via principal component analysis (PCA). While PCA mitigates the curse of dimensionality, it relies on factor pervasiveness, an assumption often violated when factors are weak, as is common in macro-financial forecasting. We propose SsPCA-MIDAS, which integrates supervised scaled PCA (SsPCA) into the mixed-data sampling fra…
Outcomes are increasingly regressed on a calibrated probability vector for unobserved class membership, and that vector is often coarsened to a hard label first. Under a constant-coefficient structural mean and conditional calibration, the observed-data problem is a partially linear regression of the outcome on the probability vector; we take this reduction as the starting point and ask what coarsening costs. For any coarsening, the plug-in estimator con…
We provide an estimator for the perturbed utility route choice (PURC) model that works with data at the level of individual trips. The estimator is a nested fixed-point algorithm that combines an upper bias-corrected linear regression problem with a lower individual-level perturbed utility maximization problem. We establish the statistical properties of the microPURC estimator and confirm these results with an experiment using simulated data. Finally, we…
Standard multinomial probit (MNP) models specify symmetric latent utility distributions, implying that choice probabilities respond symmetrically to positive and negative covariate shifts of the same magnitude. This restriction is often implausible in empirical choice settings and can lead to misleading elasticity and substitution predictions. We propose a skewed multinomial probit (SMNP) model that captures asymmetric choice responses by specifying a mu…
This article considers the problem of testing sign agreement among a finite number of parameters. This problem arises in empirical settings such as detecting treatment effects with opposite signs across subgroups, outcomes, or time periods, and testing instrument validity for local average treatment effects. For the null hypothesis that the parameters are either all non-negative or all non-positive, I propose two novel tests: a least favorable test and a…
This paper considers design-based inference on the average treatment effect in finely stratified experiments, where uncertainty arises only from the randomized treatment assignment. We focus on settings in which units are first stratified into groups of fixed size according to baseline covariates and, then within each group, exactly one unit is assigned to treatment. In this setting, we introduce a class of graph-Laplacian variance estimators in which st…
We develop a bias-robust causal inference method for observational panel data settings. Such methods typically impute untreated outcomes, so counterfactual error passes straight into the estimated treatment effect while conventional standard errors ignore it. We adapt bias-aware minimax methods, developed for estimating regression coefficients in factor-model panels, to a causal target: the average effect on the treated, which has to be imputed and may v…
I study the optimal design and analysis of randomized experiments for estimating finite-population average treatment effects when potential outcomes are known to be bounded, as with binary outcomes. Among all assignment mechanisms and a broad class of affine estimators, worst-case mean-squared error (MSE) is minimized by independent random assignment and an unconventional regression of the support-midpoint-centered outcome on the recentered treatment, wi…
Francesco Del Prato, Yaroslav Korobka, Paolo Zacchia
10 Aug 2026 · Econometrics
How much wage dispersion is attributed to workers, firms, and their sorting depends on how wages are adjusted for observed characteristics. Standard AKM decompositions impose a known linear adjustment. We develop Generalized AKM, a framework that permits an unknown smooth covariate function and group-specific nonlinear interactions while preserving the original variance components. We prove consistency and asymptotic normality with heteroskedastic errors…
Standard CATE estimators become inadequate under strong treatment-effect heterogeneity: confidence intervals for conditional means need not cover individual counterfactual effects. We propose an Individualized Causal Prediction (ICP) framework that constructs finite-sample valid conformal prediction intervals for the individual causal effect of a specific query unit. The method localizes calibration to a causally relevant neighborhood using cosine simila…
Road safety mechanisms operate within seconds, minutes and trips, whereas motor insurance observes liability claims aggregated over policy years. An annual rating coefficient can therefore predict claims accurately while leaving the crash-generating process unresolved. We propose a multiscale causal DAG framework with three parts: a proposed crash-occurrence graph constructed from a structured, non-exhaustive map of 72 study--edge records; a separate obs…
Individuals are often influenced by their peers because deviating from prevailing behavior entails social costs. However, existing peer effects models typically assume that individuals respond similarly to peers who perform better or worse than they do. This paper introduces a novel structural model of asymmetric peer effects in which conformity incentives depend on whether individuals perform below or above each of their peers. We establish that the mod…
The paper considers the problem of variable selection for forecasting electricity spot prices. High-dimensional methods such as LASSO and Elastic Net are widely used for this purpose, and while they exhibit strong predictive performance, their tendency to select over-parameterized models raises questions about interpretability. We evaluate the performance of six variable selection procedures, includingthe recently proposed Boosting Multiple Testing (BMT)…
We provide a new asymptotic framework to derive approximately optimal treatment assignments when sampling noise from data is compounded by fundamental uncertainty due to partial identification. We recenter the reduced-form parameter around its least-favorable configuration and consider drifting parameter sequences that yield both diminishing levels of sampling uncertainty and of partial identification. We characterize the limiting decision problem as a n…
This paper studies a linear panel model with an unrestricted individual effect and a time- stationary idiosyncratic disturbance. We first show that stationarity is a strong restriction in a quantile model. In a linear conditional quantile specification with quantile-dependent slopes, equality of the conditional residual distributions across periods generically forces the slope coefficient to be constant over the quantile index. Thus, a stationary-error m…
Interconnection queues, not electricity prices, now govern where data centers can be built, and the standard levelized-cost comparison answers a question no developer faces: it assumes a load profile, freezes the grid price while modeling the demand that moves it, and quotes busbar costs a facility cannot buy. This paper evaluates nine on-site supply technologies against a delivered grid whose price is endogenous to projected data-center demand, on a com…
In staggered difference-in-differences (DiD) designs, units enter treatment at different calendar times, so the treatment effect is not a single number but a set of Cohort-Average Treatment effects on the Treated (CATTs), one per cohort-time cell. Estimating every CATT as its own parameter, as the standard fully flexible estimator does, is unbiased but inefficient when some of these effects are in fact equal, whereas pooling them all into a single two-wa…
Many economic and financial relationships may change gradually rather than abruptly. We study panel data models in which the coefficient vector is continuous and piecewise linear in calendar time, with a finite number of unknown kink dates at which its slope changes. We propose a penalised least squares estimator that applies adaptive weighted group penalties to the second differences of the coefficient path, and develop asymptotic theory showing that it…
Introduction: We present a framework to assess the economic value of healthcare interventions by disaggregating value and examining heterogeneity. We applied it to an early health-economic model of population screening in England with a multi-cancer early detection (MCED) test. Value for such technologies often includes benefits, such as those associated with earlier detection, alongside potential harms from, for example, false positives or overdiagnosis…
The extended-onion and C-vine constructions of Lewandowski, Kurowicka and Joe (2009) are standard methods for sampling from the $LKJ_n(η)$ distribution on correlation matrices. We show that both arise from the simpler row-normalized Bartlett construction associated with the restricted-Wishart representation of Wang, Wu and Chu (2018), which reuses random quantities that the classical samplers regenerate. Two exact row-wise couplings establish this: the s…
Fixed-effect saturation alone is not weak identification: in the baseline model, fixed-effect--residualized OLS is unbiased and conventional inference is asymptotically exact for every residual treatment variance $τ^2=nQ_K>0$. Classical measurement error in the treatment restores it, and we derive Stock--Yogo-style critical values for $τ^2$. Under the local drift $σ_ν^2 = c^2/n$, attenuation produces a non-central limit whose non-centrality $η$ decreases…
We characterize asymmetric tail risk across over one hundred U.S. macroeconomic and financial variables using a dynamic factor model with stochastic volatility. A single mechanism unifies growth-at-risk, inflation-at-risk, and sectoral risk heterogeneity: common factors and their volatilities move together, while heterogeneous loadings transmit the resulting asymmetry unevenly across variables. We find that asymmetric tail risk is pervasive but heterogen…
Generative models can reproduce an observational distribution while encoding an incorrect causal structure. We study a sequential game in which a structural causal generator proposes observational and interventional distributions, while an adversarial experimentalist selects interventions intended to maximally falsify the generator. The discriminator is therefore not merely a real-versus-synthetic classifier: it is indexed by an intervention and tests wh…
Every SVAR result is conditional on two choices: the restrictions that identify the shock and the variables on which they operate. The literature disciplines the first; the second is chosen by hand. We develop a Bayesian methodology that constructs information sets, uses an out-of-sample criterion, and retains the largest system it admits. Under recursive identification, output rises with housing production rather than household credit alone. For monetar…
I develop an estimation and inference framework for distribution regression in dyadic network settings with two-way fixed effects that vary across thresholds of the outcome. I show that identification of the structural parameters is achieved through binarization of the outcome at each threshold, and estimate the model by conditional maximum likelihood, which "differences out" the fixed effects and circumvents the incidental parameter problem. The estimat…
In saturated fixed-effects regressions, Gaussian inference depends not on total identifying variation but on its concentration, measured by the self-normalized leverage $λ_n$ of the residualized treatment. When finitely many score weights remain persistent, the $t$-statistic converges to a convolution of raw errors and a Gaussian component. At full concentration, its null distribution varies across symmetric error laws with equal variance, so no fixed cr…
We develop a hierarchical Bayesian panel quantile regression model in which unit-specific coefficient paths are smoothed across quantiles by Gaussian processes, while a common time effect absorbs aggregate shocks. Componentwise-monotone Bernstein polynomials, perturbed by unit-specific deviations, deliver soft noncrossing, and we provide identification conditions together with a bound on the crossing probability. Applying the model to 33 countries over 1…
Estimating the dynamic effects of economic shocks in short and very short samples is impeded by a lack of degrees of freedom. We offer a solution based on a Bayesian hierarchical framework for estimating local projection (LP) impulse response functions across a panel of related time series. The framework explicitly accommodates unbalanced panels in which some series are substantially shorter than others, allowing the short series to borrow information fr…
Formula One outcomes reflect the joint contributions of drivers and constructors, but these contributions are unobserved and vary over time. We propose a Bayesian state-space model that disentangles dynamic driver and constructor abilities using two observed outcomes: fastest qualifying lap times and race rankings. Both outcomes depend jointly on latent driver and constructor states that evolve at the Grand Prix level, while the race equation additionall…
Credit stress testing requires impulse responses of portfolio default probabilities, not only macro-financial drivers. We derive closed-form generalized impulse responses for the mean, quantiles (PD-at-Risk), and expected shortfall in a modular framework combining a Bayesian VAR, a Gaussian satellite, and the Merton-Vasicek model underlying Basel IRB regulation. Results extend to any probit-Gaussian mapping of a latent factor. Nonlinearity makes response…
Objectives: While causal analysis of travel behavior is an emerging field, estimating heterogeneity in mode choice through causal modeling remains unexplored. This study demonstrates the application of a novel causal method, causal forest, to quantify the heterogeneity in travel mode choice shifts caused by the COVID-19 pandemic. Methods: We applied causal forests, a non-parametric causal machine learning method, to 802,935 trip records from the 2017 and…
The set of $n\times n$ correlation matrices, known as the elliptope, has volume decaying at the super-exponential rate $\exp{-\tfrac14 n^2\log n}$. We characterize where this vanishing volume concentrates. A uniform draw is entrywise close to the identity yet globally far from it and nearly singular: its maximum absolute correlation is of order $\sqrt{\log n/n}$, its Frobenius distance is asymptotic to $\sqrt n$, its empirical spectral distribution conve…
This paper considers difference-in-differences identification strategies when the parallel trends assumption holds after conditioning on covariates that may themselves be affected by the treatment (often referred to as "bad controls"). We show that common approaches such as simply dropping bad controls are often ill-advised and develop two alternative approaches that allow bad controls to function as genuine controls despite being affected by treatment.…
Empirical work often removes fixed effects, latent factors, or high-dimensional controls before estimating structural relationships. These transformations reduce confounding but may also remove identifying variation. We study linear panel IV after one equation-compatible nuisance projection under two-way dependence. The projected Jacobian determines which structural directions remain visible; the projected-score law determines their precision; and, on Ga…
This paper develops an econometric framework for analysing smooth structural change in cointegrated systems following a known intervention time. We consider a vector error-correction model in which the cointegration rank and the pre-intervention cointegrating structure are identified from a stable pre-intervention subsample. After the intervention, both the adjustment coefficients and the cointegrating vectors are allowed to evolve smoothly as functions…
We develop a fully nonlinear structural vector autoregressive framework in which the contemporaneous structural mapping may be nonlinear and non-additive. Identification is achieved by exploiting variation in the conditional distributions of the mutually independent structural shocks induced by an observed exogenous variable. Specifically, a general contrastive learning framework that makes use of this variation together with the assumed exponential-fami…