Every paper in arXiv econ.EM, its bibliography parsed from the LaTeX source, and a weighted citation graph over the whole corpus — linked to author profiles and publication records.
When parallel trends fails for some treated cohorts but not others, the average treatment effect on the treated (ATT), an average over all of them, is exactly the target that becomes hard to recover. We propose changing the estimand rather than defending it. The credible-subpopulation local ATT (LATT) is the effect for the subpopulation of cohorts whose parallel trends is credible, and it is point-identified under parallel trends for the selected cohorts…
This paper analyzes when choice probabilities reveal rankings of deterministic utility indices in semiparametric discrete choice models. It begins with binary choice, where quantile thresholds guarantee ranking recovery, and shows that such thresholds can arise either from behavioral departures from utility maximization (e.g., limited attention) under exchangeable unobservables, or from non-exchangeable unobservables under standard utility maximization.…
Empirical studies of peer effects often exploit conditional random assignment to peer groups within urns. We develop a GMM framework for estimation and inference in this setting. The framework separately identifies endogenous and contextual peer effects and nests tests of random peer-group assignment as a special case. It permits unknown heteroskedasticity and corrects finite-urn bias in variance estimation. Its asymptotic theory allows the number of pee…
We ask whether COVID-19 lockdown stringency altered national Olympic performance between Rio 2016 and Tokyo 2020, using the Oxford Stringency Index and the 99 countries that won a medal in either edition. As in \citet{liu2024}, mean performance is unaffected: stringency is insignificant in every OLS and ANOVA specification. The distribution is not. Among the 84 non-traditionally dominant nations, medal changes are three to six times more dispersed in hig…
We study the existence of a Regional Differential in rugby sevens: whether, in tournaments where no competing team enjoys formal home status, some national sides systematically over- or under-perform depending on where the event is staged. Using the universe of 2672 men's and women's matches from international rugby sevens tournaments played between 2016 and 2025, principally the World Rugby Sevens Series, but also the World Cup Sevens and the Olympic Ga…
We study how many observations are needed to determine the causal direction between two linearly related variables. Classical LiNGAM theory shows that independent non-Gaussian disturbances identify the direction, but does not quantify the difficulty when the causal effect is weak or the disturbances are nearly Gaussian. Let $β$ bound the absolute structural coefficient from below, let $ν$ measure each standardized disturbance's distance from Gaussianity,…
Route and activity choice are connected levels of a common sequential mobility decision problem: activity choice determines what people do, where, and when, while route choice governs how they move between activities. This review develops a unified framework connecting transportation choice modeling with inverse reinforcement learning (IRL) and imitation learning (IL). Under explicit assumptions, recursive logit, logit dynamic discrete choice, and maximu…
We propose a family of control variate estimators for variance reduction in design-based survey sampling and causal inference, with and without interference. In these settings, inverse probability weighting (IPW) estimators are widely used, but may have large variance when sampling, treatment, or exposure probabilities are small. Building on the observation that several common estimators, including the Hajek, normalized, and augmented inverse probability…
This study investigates the methodological and theoretical properties of session handover in applications that use large language models. A task may continue in a new session when the context reaches the model's input limit, when the application restarts, or when another agent is asked to finish the task. The application must then decide which information from the earlier session to pass on. We formulate handover as the transfer of a task-relative in-con…
This paper employs a Threshold Bayesian Vector Autoregression (TBVAR) to estimate the regime-dependent macroeconomic effects of capital regulation in Hungary. Using the Factor-based Index of Systemic Stress (FISS) as the threshold variable, the model identifies normal and stress regimes consistent with the occasionally binding constraints literature. The TBVAR offers a practical multivariate alternative to Growth-at-Risk for data-constrained economies. G…
We develop a method for estimating and testing a single block of a macroeconomic model with heterogeneous agents, without placing assumptions on the structure of the rest of the economy. In a large class of models, individual agents' decisions depend on the macroeconomy only through their expectations of the evolution of a finite-dimensional vector of "sufficient statistics" (e.g., asset returns or aggregate earnings). Our estimator selects the structura…
Limited dependent variable models are central to empirical economics, but likelihood-based inference is infeasible when likelihoods involve high-dimensional integration over latent variables. This paper proposes Stochastically Estimated Gradient Ascent (SEGA), a scalable estimation approach for limited dependent variable models. Using Fisher's identity, SEGA replaces the intractable likelihood score with an unbiased augmented-data score evaluated at a si…
We study minimax-optimal designs and estimators for estimating the sample average treatment effect in finite population randomized experiments, where both design and estimator are unrestricted. For binary potential outcomes, we show this minimax risk is equivalent to the minimax risk $ρ_n^*$ of an estimation problem with $2$ unknown parameters. We leverage this reduction to establish a second-order risk expansion $ρ_n^* = n^{-1} - Cn^{-4/3} + o_n(n^{-4/3…
When comparison units may also respond to treatment, panel comparisons reflect both the treatment effect and spillovers. If the interference pattern is unknown, observed outcomes alone do not separate the two. I characterize what can nevertheless be learned from panel outcomes under general restrictions, without requiring an exposure mapping or prior classification of affected donors. The framework scales validity bounds for every convex donor weight by…
Many questions across the sciences take the same form: several coupled series are observed together, and the analyst wants to know not merely that they move together but which one moves first, and how strongly. This paper sets out a complete method built on one organising idea: the direction of a coupled system is exactly the part of its behaviour that changes when the record is played backwards. Tools built on contemporaneous covariance alone (correlati…
I consider an AR($p$) process that is observed every $q$ periods, either as a snapshot (stock variable) or as a sum over the sampling interval (flow variable). Under fairly mild assumptions, I derive the identified set for general lag lengths $p \in \mathbb{N}$ and sampling frequencies $q \in \mathbb{N}$, I bound its cardinality, and I provide a recipe to compute all candidate points and determine their membership in the identified set. My analysis suppo…
We consider the classical additive measurement-error model $X=Y+Z$, where the latent random variable $Y$ has unknown distribution $F_Y$ and the error $Z$ has a known distribution. We develop direct estimators for three functionals of $F_Y$: (i) $F_Y(x)$ at continuity points; (ii) interval probabilities $F_Y(y)-F_Y(x)$ when $x<y$ are continuity points; and (iii) the size of a jump at a prespecified discontinuity. We derive non-asymptotic bias and variance…
Factor-MIDAS regressions forecast a low-frequency target by extracting common factors from a large panel of high-frequency predictors via principal component analysis (PCA). While PCA mitigates the curse of dimensionality, it relies on factor pervasiveness, an assumption often violated when factors are weak, as is common in macro-financial forecasting. We propose SsPCA-MIDAS, which integrates supervised scaled PCA (SsPCA) into the mixed-data sampling fra…
Outcomes are increasingly regressed on a calibrated probability vector for unobserved class membership, and that vector is often coarsened to a hard label first. Under a constant-coefficient structural mean and conditional calibration, the observed-data problem is a partially linear regression of the outcome on the probability vector; we take this reduction as the starting point and ask what coarsening costs. For any coarsening, the plug-in estimator con…
We provide an estimator for the perturbed utility route choice (PURC) model that works with data at the level of individual trips. The estimator is a nested fixed-point algorithm that combines an upper bias-corrected linear regression problem with a lower individual-level perturbed utility maximization problem. We establish the statistical properties of the microPURC estimator and confirm these results with an experiment using simulated data. Finally, we…
Standard multinomial probit (MNP) models specify symmetric latent utility distributions, implying that choice probabilities respond symmetrically to positive and negative covariate shifts of the same magnitude. This restriction is often implausible in empirical choice settings and can lead to misleading elasticity and substitution predictions. We propose a skewed multinomial probit (SMNP) model that captures asymmetric choice responses by specifying a mu…
This article considers the problem of testing sign agreement among a finite number of parameters. This problem arises in empirical settings such as detecting treatment effects with opposite signs across subgroups, outcomes, or time periods, and testing instrument validity for local average treatment effects. For the null hypothesis that the parameters are either all non-negative or all non-positive, I propose two novel tests: a least favorable test and a…
This paper considers design-based inference on the average treatment effect in finely stratified experiments, where uncertainty arises only from the randomized treatment assignment. We focus on settings in which units are first stratified into groups of fixed size according to baseline covariates and, then within each group, exactly one unit is assigned to treatment. In this setting, we introduce a class of graph-Laplacian variance estimators in which st…
We develop a bias-robust causal inference method for observational panel data settings. Such methods typically impute untreated outcomes, so counterfactual error passes straight into the estimated treatment effect while conventional standard errors ignore it. We adapt bias-aware minimax methods, developed for estimating regression coefficients in factor-model panels, to a causal target: the average effect on the treated, which has to be imputed and may v…
I study the optimal design and analysis of randomized experiments for estimating finite-population average treatment effects when potential outcomes are known to be bounded, as with binary outcomes. Among all assignment mechanisms and a broad class of affine estimators, worst-case mean-squared error (MSE) is minimized by independent random assignment and an unconventional regression of the support-midpoint-centered outcome on the recentered treatment, wi…
Francesco Del Prato, Yaroslav Korobka, Paolo Zacchia
10 Aug 2026 · Econometrics
How much wage dispersion is attributed to workers, firms, and their sorting depends on how wages are adjusted for observed characteristics. Standard AKM decompositions impose a known linear adjustment. We develop Generalized AKM, a framework that permits an unknown smooth covariate function and group-specific nonlinear interactions while preserving the original variance components. We prove consistency and asymptotic normality with heteroskedastic errors…
Standard CATE estimators become inadequate under strong treatment-effect heterogeneity: confidence intervals for conditional means need not cover individual counterfactual effects. We propose an Individualized Causal Prediction (ICP) framework that constructs finite-sample valid conformal prediction intervals for the individual causal effect of a specific query unit. The method localizes calibration to a causally relevant neighborhood using cosine simila…
Road safety mechanisms operate within seconds, minutes and trips, whereas motor insurance observes liability claims aggregated over policy years. An annual rating coefficient can therefore predict claims accurately while leaving the crash-generating process unresolved. We propose a multiscale causal DAG framework with three parts: a proposed crash-occurrence graph constructed from a structured, non-exhaustive map of 72 study--edge records; a separate obs…
Individuals are often influenced by their peers because deviating from prevailing behavior entails social costs. However, existing peer effects models typically assume that individuals respond similarly to peers who perform better or worse than they do. This paper introduces a novel structural model of asymmetric peer effects in which conformity incentives depend on whether individuals perform below or above each of their peers. We establish that the mod…
The paper considers the problem of variable selection for forecasting electricity spot prices. High-dimensional methods such as LASSO and Elastic Net are widely used for this purpose, and while they exhibit strong predictive performance, their tendency to select over-parameterized models raises questions about interpretability. We evaluate the performance of six variable selection procedures, includingthe recently proposed Boosting Multiple Testing (BMT)…
We provide a new asymptotic framework to derive approximately optimal treatment assignments when sampling noise from data is compounded by fundamental uncertainty due to partial identification. We recenter the reduced-form parameter around its least-favorable configuration and consider drifting parameter sequences that yield both diminishing levels of sampling uncertainty and of partial identification. We characterize the limiting decision problem as a n…
This paper studies a linear panel model with an unrestricted individual effect and a time- stationary idiosyncratic disturbance. We first show that stationarity is a strong restriction in a quantile model. In a linear conditional quantile specification with quantile-dependent slopes, equality of the conditional residual distributions across periods generically forces the slope coefficient to be constant over the quantile index. Thus, a stationary-error m…
Interconnection queues, not electricity prices, now govern where data centers can be built, and the standard levelized-cost comparison answers a question no developer faces: it assumes a load profile, freezes the grid price while modeling the demand that moves it, and quotes busbar costs a facility cannot buy. This paper evaluates nine on-site supply technologies against a delivered grid whose price is endogenous to projected data-center demand, on a com…
In staggered difference-in-differences (DiD) designs, units enter treatment at different calendar times, so the treatment effect is not a single number but a set of Cohort-Average Treatment effects on the Treated (CATTs), one per cohort-time cell. Estimating every CATT as its own parameter, as the standard fully flexible estimator does, is unbiased but inefficient when some of these effects are in fact equal, whereas pooling them all into a single two-wa…
Many economic and financial relationships may change gradually rather than abruptly. We study panel data models in which the coefficient vector is continuous and piecewise linear in calendar time, with a finite number of unknown kink dates at which its slope changes. We propose a penalised least squares estimator that applies adaptive weighted group penalties to the second differences of the coefficient path, and develop asymptotic theory showing that it…
Introduction: We present a framework to assess the economic value of healthcare interventions by disaggregating value and examining heterogeneity. We applied it to an early health-economic model of population screening in England with a multi-cancer early detection (MCED) test. Value for such technologies often includes benefits, such as those associated with earlier detection, alongside potential harms from, for example, false positives or overdiagnosis…
The extended-onion and C-vine constructions of Lewandowski, Kurowicka and Joe (2009) are standard methods for sampling from the $LKJ_n(η)$ distribution on correlation matrices. We show that both arise from the simpler row-normalized Bartlett construction associated with the restricted-Wishart representation of Wang, Wu and Chu (2018), which reuses random quantities that the classical samplers regenerate. Two exact row-wise couplings establish this: the s…
Fixed-effect saturation alone is not weak identification: in the baseline model, fixed-effect--residualized OLS is unbiased and conventional inference is asymptotically exact for every residual treatment variance $τ^2=nQ_K>0$. Classical measurement error in the treatment restores it, and we derive Stock--Yogo-style critical values for $τ^2$. Under the local drift $σ_ν^2 = c^2/n$, attenuation produces a non-central limit whose non-centrality $η$ decreases…
We characterize asymmetric tail risk across over one hundred U.S. macroeconomic and financial variables using a dynamic factor model with stochastic volatility. A single mechanism unifies growth-at-risk, inflation-at-risk, and sectoral risk heterogeneity: common factors and their volatilities move together, while heterogeneous loadings transmit the resulting asymmetry unevenly across variables. We find that asymmetric tail risk is pervasive but heterogen…
Generative models can reproduce an observational distribution while encoding an incorrect causal structure. We study a sequential game in which a structural causal generator proposes observational and interventional distributions, while an adversarial experimentalist selects interventions intended to maximally falsify the generator. The discriminator is therefore not merely a real-versus-synthetic classifier: it is indexed by an intervention and tests wh…