Every paper in arXiv econ.EM, its bibliography parsed from the LaTeX source, and a weighted citation graph over the whole corpus — linked to author profiles and publication records.
This study investigates the methodological and theoretical properties of session handover in applications that use large language models. A task may continue in a new session when the context reaches the model's input limit, when the application restarts, or when another agent is asked to finish the task. The application must then decide which information from the earlier session to pass on. We formulate handover as the transfer of a task-relative in-con…
This paper employs a Threshold Bayesian Vector Autoregression (TBVAR) to estimate the regime-dependent macroeconomic effects of capital regulation in Hungary. Using the Factor-based Index of Systemic Stress (FISS) as the threshold variable, the model identifies normal and stress regimes consistent with the occasionally binding constraints literature. The TBVAR offers a practical multivariate alternative to Growth-at-Risk for data-constrained economies. G…
We develop a method for estimating and testing a single block of a macroeconomic model with heterogeneous agents, without placing assumptions on the structure of the rest of the economy. In a large class of models, individual agents' decisions depend on the macroeconomy only through their expectations of the evolution of a finite-dimensional vector of "sufficient statistics" (e.g., asset returns or aggregate earnings). Our estimator selects the structura…
Limited dependent variable models are central to empirical economics, but likelihood-based inference is infeasible when likelihoods involve high-dimensional integration over latent variables. This paper proposes Stochastically Estimated Gradient Ascent (SEGA), a scalable estimation approach for limited dependent variable models. Using Fisher's identity, SEGA replaces the intractable likelihood score with an unbiased augmented-data score evaluated at a si…
We study minimax-optimal designs and estimators for estimating the sample average treatment effect in finite population randomized experiments, where both design and estimator are unrestricted. For binary potential outcomes, we show this minimax risk is equivalent to the minimax risk $ρ_n^*$ of an estimation problem with $2$ unknown parameters. We leverage this reduction to establish a second-order risk expansion $ρ_n^* = n^{-1} - Cn^{-4/3} + o_n(n^{-4/3…
When comparison units may also respond to treatment, panel comparisons reflect both the treatment effect and spillovers. If the interference pattern is unknown, observed outcomes alone do not separate the two. I characterize what can nevertheless be learned from panel outcomes under general restrictions, without requiring an exposure mapping or prior classification of affected donors. The framework scales validity bounds for every convex donor weight by…
Many questions across the sciences take the same form: several coupled series are observed together, and the analyst wants to know not merely that they move together but which one moves first, and how strongly. This paper sets out a complete method built on one organising idea: the direction of a coupled system is exactly the part of its behaviour that changes when the record is played backwards. Tools built on contemporaneous covariance alone (correlati…
I consider an AR($p$) process that is observed every $q$ periods, either as a snapshot (stock variable) or as a sum over the sampling interval (flow variable). Under fairly mild assumptions, I derive the identified set for general lag lengths $p \in \mathbb{N}$ and sampling frequencies $q \in \mathbb{N}$, I bound its cardinality, and I provide a recipe to compute all candidate points and determine their membership in the identified set. My analysis suppo…
We consider the classical additive measurement-error model $X=Y+Z$, where the latent random variable $Y$ has unknown distribution $F_Y$ and the error $Z$ has a known distribution. We develop direct estimators for three functionals of $F_Y$: (i) $F_Y(x)$ at continuity points; (ii) interval probabilities $F_Y(y)-F_Y(x)$ when $x<y$ are continuity points; and (iii) the size of a jump at a prespecified discontinuity. We derive non-asymptotic bias and variance…
Factor-MIDAS regressions forecast a low-frequency target by extracting common factors from a large panel of high-frequency predictors via principal component analysis (PCA). While PCA mitigates the curse of dimensionality, it relies on factor pervasiveness, an assumption often violated when factors are weak, as is common in macro-financial forecasting. We propose SsPCA-MIDAS, which integrates supervised scaled PCA (SsPCA) into the mixed-data sampling fra…
Outcomes are increasingly regressed on a calibrated probability vector for unobserved class membership, and that vector is often coarsened to a hard label first. Under a constant-coefficient structural mean and conditional calibration, the observed-data problem is a partially linear regression of the outcome on the probability vector; we take this reduction as the starting point and ask what coarsening costs. For any coarsening, the plug-in estimator con…
We provide an estimator for the perturbed utility route choice (PURC) model that works with data at the level of individual trips. The estimator is a nested fixed-point algorithm that combines an upper bias-corrected linear regression problem with a lower individual-level perturbed utility maximization problem. We establish the statistical properties of the microPURC estimator and confirm these results with an experiment using simulated data. Finally, we…
Standard multinomial probit (MNP) models specify symmetric latent utility distributions, implying that choice probabilities respond symmetrically to positive and negative covariate shifts of the same magnitude. This restriction is often implausible in empirical choice settings and can lead to misleading elasticity and substitution predictions. We propose a skewed multinomial probit (SMNP) model that captures asymmetric choice responses by specifying a mu…
This article considers the problem of testing sign agreement among a finite number of parameters. This problem arises in empirical settings such as detecting treatment effects with opposite signs across subgroups, outcomes, or time periods, and testing instrument validity for local average treatment effects. For the null hypothesis that the parameters are either all non-negative or all non-positive, I propose two novel tests: a least favorable test and a…
This paper considers design-based inference on the average treatment effect in finely stratified experiments, where uncertainty arises only from the randomized treatment assignment. We focus on settings in which units are first stratified into groups of fixed size according to baseline covariates and, then within each group, exactly one unit is assigned to treatment. In this setting, we introduce a class of graph-Laplacian variance estimators in which st…
We develop a bias-robust causal inference method for observational panel data settings. Such methods typically impute untreated outcomes, so counterfactual error passes straight into the estimated treatment effect while conventional standard errors ignore it. We adapt bias-aware minimax methods, developed for estimating regression coefficients in factor-model panels, to a causal target: the average effect on the treated, which has to be imputed and may v…
I study the optimal design and analysis of randomized experiments for estimating finite-population average treatment effects when potential outcomes are known to be bounded, as with binary outcomes. Among all assignment mechanisms and a broad class of affine estimators, worst-case mean-squared error (MSE) is minimized by independent random assignment and an unconventional regression of the support-midpoint-centered outcome on the recentered treatment, wi…
Francesco Del Prato, Yaroslav Korobka, Paolo Zacchia
10 Aug 2026 · Econometrics
How much wage dispersion is attributed to workers, firms, and their sorting depends on how wages are adjusted for observed characteristics. Standard AKM decompositions impose a known linear adjustment. We develop Generalized AKM, a framework that permits an unknown smooth covariate function and group-specific nonlinear interactions while preserving the original variance components. We prove consistency and asymptotic normality with heteroskedastic errors…
Standard CATE estimators become inadequate under strong treatment-effect heterogeneity: confidence intervals for conditional means need not cover individual counterfactual effects. We propose an Individualized Causal Prediction (ICP) framework that constructs finite-sample valid conformal prediction intervals for the individual causal effect of a specific query unit. The method localizes calibration to a causally relevant neighborhood using cosine simila…
Road safety mechanisms operate within seconds, minutes and trips, whereas motor insurance observes liability claims aggregated over policy years. An annual rating coefficient can therefore predict claims accurately while leaving the crash-generating process unresolved. We propose a multiscale causal DAG framework with three parts: a proposed crash-occurrence graph constructed from a structured, non-exhaustive map of 72 study--edge records; a separate obs…
Individuals are often influenced by their peers because deviating from prevailing behavior entails social costs. However, existing peer effects models typically assume that individuals respond similarly to peers who perform better or worse than they do. This paper introduces a novel structural model of asymmetric peer effects in which conformity incentives depend on whether individuals perform below or above each of their peers. We establish that the mod…
The paper considers the problem of variable selection for forecasting electricity spot prices. High-dimensional methods such as LASSO and Elastic Net are widely used for this purpose, and while they exhibit strong predictive performance, their tendency to select over-parameterized models raises questions about interpretability. We evaluate the performance of six variable selection procedures, includingthe recently proposed Boosting Multiple Testing (BMT)…
We provide a new asymptotic framework to derive approximately optimal treatment assignments when sampling noise from data is compounded by fundamental uncertainty due to partial identification. We recenter the reduced-form parameter around its least-favorable configuration and consider drifting parameter sequences that yield both diminishing levels of sampling uncertainty and of partial identification. We characterize the limiting decision problem as a n…
This paper studies a linear panel model with an unrestricted individual effect and a time- stationary idiosyncratic disturbance. We first show that stationarity is a strong restriction in a quantile model. In a linear conditional quantile specification with quantile-dependent slopes, equality of the conditional residual distributions across periods generically forces the slope coefficient to be constant over the quantile index. Thus, a stationary-error m…
Interconnection queues, not electricity prices, now govern where data centers can be built, and the standard levelized-cost comparison answers a question no developer faces: it assumes a load profile, freezes the grid price while modeling the demand that moves it, and quotes busbar costs a facility cannot buy. This paper evaluates nine on-site supply technologies against a delivered grid whose price is endogenous to projected data-center demand, on a com…
In staggered difference-in-differences (DiD) designs, units enter treatment at different calendar times, so the treatment effect is not a single number but a set of Cohort-Average Treatment effects on the Treated (CATTs), one per cohort-time cell. Estimating every CATT as its own parameter, as the standard fully flexible estimator does, is unbiased but inefficient when some of these effects are in fact equal, whereas pooling them all into a single two-wa…
Many economic and financial relationships may change gradually rather than abruptly. We study panel data models in which the coefficient vector is continuous and piecewise linear in calendar time, with a finite number of unknown kink dates at which its slope changes. We propose a penalised least squares estimator that applies adaptive weighted group penalties to the second differences of the coefficient path, and develop asymptotic theory showing that it…
Introduction: We present a framework to assess the economic value of healthcare interventions by disaggregating value and examining heterogeneity. We applied it to an early health-economic model of population screening in England with a multi-cancer early detection (MCED) test. Value for such technologies often includes benefits, such as those associated with earlier detection, alongside potential harms from, for example, false positives or overdiagnosis…
The extended-onion and C-vine constructions of Lewandowski, Kurowicka and Joe (2009) are standard methods for sampling from the $LKJ_n(η)$ distribution on correlation matrices. We show that both arise from the simpler row-normalized Bartlett construction associated with the restricted-Wishart representation of Wang, Wu and Chu (2018), which reuses random quantities that the classical samplers regenerate. Two exact row-wise couplings establish this: the s…
Fixed-effect saturation alone is not weak identification: in the baseline model, fixed-effect--residualized OLS is unbiased and conventional inference is asymptotically exact for every residual treatment variance $τ^2=nQ_K>0$. Classical measurement error in the treatment restores it, and we derive Stock--Yogo-style critical values for $τ^2$. Under the local drift $σ_ν^2 = c^2/n$, attenuation produces a non-central limit whose non-centrality $η$ decreases…
We characterize asymmetric tail risk across over one hundred U.S. macroeconomic and financial variables using a dynamic factor model with stochastic volatility. A single mechanism unifies growth-at-risk, inflation-at-risk, and sectoral risk heterogeneity: common factors and their volatilities move together, while heterogeneous loadings transmit the resulting asymmetry unevenly across variables. We find that asymmetric tail risk is pervasive but heterogen…
Generative models can reproduce an observational distribution while encoding an incorrect causal structure. We study a sequential game in which a structural causal generator proposes observational and interventional distributions, while an adversarial experimentalist selects interventions intended to maximally falsify the generator. The discriminator is therefore not merely a real-versus-synthetic classifier: it is indexed by an intervention and tests wh…
Every SVAR result is conditional on two choices: the restrictions that identify the shock and the variables on which they operate. The literature disciplines the first; the second is chosen by hand. We develop a Bayesian methodology that constructs information sets, uses an out-of-sample criterion, and retains the largest system it admits. Under recursive identification, output rises with housing production rather than household credit alone. For monetar…
I develop an estimation and inference framework for distribution regression in dyadic network settings with two-way fixed effects that vary across thresholds of the outcome. I show that identification of the structural parameters is achieved through binarization of the outcome at each threshold, and estimate the model by conditional maximum likelihood, which "differences out" the fixed effects and circumvents the incidental parameter problem. The estimat…
In saturated fixed-effects regressions, Gaussian inference depends not on total identifying variation but on its concentration, measured by the self-normalized leverage $λ_n$ of the residualized treatment. When finitely many score weights remain persistent, the $t$-statistic converges to a convolution of raw errors and a Gaussian component. At full concentration, its null distribution varies across symmetric error laws with equal variance, so no fixed cr…
We develop a hierarchical Bayesian panel quantile regression model in which unit-specific coefficient paths are smoothed across quantiles by Gaussian processes, while a common time effect absorbs aggregate shocks. Componentwise-monotone Bernstein polynomials, perturbed by unit-specific deviations, deliver soft noncrossing, and we provide identification conditions together with a bound on the crossing probability. Applying the model to 33 countries over 1…
Estimating the dynamic effects of economic shocks in short and very short samples is impeded by a lack of degrees of freedom. We offer a solution based on a Bayesian hierarchical framework for estimating local projection (LP) impulse response functions across a panel of related time series. The framework explicitly accommodates unbalanced panels in which some series are substantially shorter than others, allowing the short series to borrow information fr…
Formula One outcomes reflect the joint contributions of drivers and constructors, but these contributions are unobserved and vary over time. We propose a Bayesian state-space model that disentangles dynamic driver and constructor abilities using two observed outcomes: fastest qualifying lap times and race rankings. Both outcomes depend jointly on latent driver and constructor states that evolve at the Grand Prix level, while the race equation additionall…
Credit stress testing requires impulse responses of portfolio default probabilities, not only macro-financial drivers. We derive closed-form generalized impulse responses for the mean, quantiles (PD-at-Risk), and expected shortfall in a modular framework combining a Bayesian VAR, a Gaussian satellite, and the Merton-Vasicek model underlying Basel IRB regulation. Results extend to any probit-Gaussian mapping of a latent factor. Nonlinearity makes response…
Objectives: While causal analysis of travel behavior is an emerging field, estimating heterogeneity in mode choice through causal modeling remains unexplored. This study demonstrates the application of a novel causal method, causal forest, to quantify the heterogeneity in travel mode choice shifts caused by the COVID-19 pandemic. Methods: We applied causal forests, a non-parametric causal machine learning method, to 802,935 trip records from the 2017 and…