EconBase

The econometrics arXiv, with citations

Every paper in arXiv econ.EM, its bibliography parsed from the LaTeX source, and a weighted citation graph over the whole corpus — linked to author profiles and publication records.

5,657
papers
5,170
with citations extracted
227,394
references parsed
451,465
in-text mentions
14,756
in-corpus edges

Browse or search all papers · Author profiles · Journals

Latest papers

Qihui Chen, Ka Yan Cheng, Zheng Fang
27 Jul 2026 · Econometrics
We develop a general framework of identification and estimation for automatic debiased machine learning (DML) where the parameter of interest $θ_0$ is identified by a moment condition involving a nuisance $γ_0$ that may be high dimensional. DML leverages machine learning to estimate $γ_0$ while correcting for regularization and overfitting biases that may otherwise transmit to biased estimation of $θ_0$. We establish conditions under which the Riesz repr…
Gregor Steiner, Mark Steel
27 Jul 2026 · Statistics — Methodology
Causal inference is often focused on average effects, which can hide important aspects of the effect distributions. Here we consider the entire posterior effects distribution by estimating full counterfactual outcome distributions. We propose a methodology for inference on counterfactual distributions which builds upon the martingale posterior framework of Fong et al. (2023). This provides a highly flexible approach to estimating densities, distribution…
A. Montañés, E. Ruiz
26 Jul 2026 · Econometrics
It is obvious to say that an adequate estimation of the autocorrelation function is central in time series analysis. In this paper, we propose three new robust estimators based on ratios of observations, which offer strong resistance against outliers. While the first estimator, which is based on the median, is not efficient, the second is a Quasi Maximum Likelihood (QML) estimator with better efficiency properties. The third estimator is a plug-in estima…
26 Jul 2026 · Econometrics
I critique a set of entrenched methodological conventions that collectively create systemic dysfunction in statistical research for clinical decisions. These include: (1) the prevalent use of hypothesis tests to compare treatments, (2) remoteness from patient care of the methods used to evaluate the accuracy of predictions of patient outcomes, (3) poor practice of meta-analysis to combine findings across studies, and (4) widespread research with incredib…
Rabee Tourky
26 Jul 2026 · Mathematics — Probability
Let $V$ and $U$ be independent standard normal random variables. For any Borel map $φ\colon\mathbb{R}\to\mathbb{R}$, set $Y_φ=φ(V)+U$, and define $P_φ(y)=\mathbb{E}[V\mid Y_φ=y]$ and $F_φ(x)=\mathbb{E}[P_φ(x+U)]$. We prove that, if for every $v\in\mathbb{R}$, the quantity $φ(v)$ maximises $x(v-F_φ(x))$ over $x\in\mathbb{R}$, then $φ$ is the identity function. This is the normalised one-period Kyle (1985) model of insider trading. It follows that Kyle's c…
Muhammad Abdullah Haroon
25 Jul 2026 · Machine Learning
Bitcoin price prediction on sub-daily timescales is a hard open problem in computational finance. Bitcoin exhibits fat-tailed returns, non-stationary dynamics, and a price discovery process influenced by social discourse on Reddit and Twitter. Conventional approaches fuse OHLCV technical features with sentiment via static concatenation, applying identical fusion weights regardless of market state. This is inconsistent with the behavioural finance literat…
Harsh Parikh, Gabriel Levin-Konigsberg, Nilesh Tripuraneni, Dhruv Madeka, Michael I. Jordan, Dean Foster and 2 more
25 Jul 2026 · Statistics — Applications
Randomized controlled trials (RCTs) are fundamental tools for causal inference across technology companies, pharmaceutical research, and federal agencies. While the standard difference-in-means estimator provides unbiased treatment effect estimates, it often lacks precision, particularly when treatment effects are heterogeneous or outcomes exhibit heavy-tailed distributions. Although numerous precision-enhancing methods exist---from covariate adjustment…
Tomas Havranek, Zuzana Irsova, Martina Luskova, T. D. Stanley
25 Jul 2026 · Econometrics
Meta-analysts routinely face estimates that look too large or extreme. Yet, how to handle them is left to the reviewer's judgment. The methods for detecting such estimates are well known. What is missing is an informed assessment of how much alternative handling choices might change a meta-analysis' conclusions. We fill this gap by analyzing the effects of four pre-registered handling treatments across 358 behavioral science meta-analyses with at least t…
24 Jul 2026 · Econometrics
Causal forests that estimate conditional average treatment effects by averaging honest leaf-level effects across trees are widely used in fixed-effects panel settings. We show that this averaging systematically attenuates the estimated heterogeneity: the raw prediction behaves like a + b*tau(x) with slope b < 1, so the spread of the CATEs is compressed toward the average effect, and the additive recentering used to report an unbiased average treatment ef…
Gurkirat Wadhwa, Veeraruna Kavitha
24 Jul 2026 · Econometrics
Suppliers often encroach downstream by operating in-house production-units while continuing to supply independent production-units. We study the optimal configuration, including optimal pricing, for an encroaching supplier that balances these dual roles through a Stackelberg game. The integrated supplier determines the wholesale price charged to the outsourced production unit and the retail price of its own product, while the outsourced unit responds opt…
Jiyuan Tan, Vasilis Syrgkanis
24 Jul 2026 · Statistics — Machine Learning
Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A common approach is to close the research loop with a large language model (LLM) reviewer. However, such reviewers remain empirically unreliable: they may accept fabricated papers and detect them at rates close to chance (Bad Scientist, 2025). We present CausalForge, a framework for automated theoretical research in caus…
24 Jul 2026 · Econometrics
This paper proposes a nonlinear boosting with multiple testing (BMT) approach to variable selection in high-dimensional generalised linear models with binary responses. At each stage of the BMT procedure, the model is updated by adding only the most significant covariate, conditional on those already selected in previous stages, while taking into account the multiple testing nature of the problem. It is shown that, under the stated conditions, the BMT pr…
Dalia Ghanem, Felix Pretis, Daniel Schuurman
24 Jul 2026 · Econometrics
Understanding the degree to which we are able to adapt to climate change is central to economic assessments of future climate damages. Economists increasingly use comparisons between long differences and fixed effects estimators to measure climate adaptation. We show that such comparisons can be misleading. Neither estimator is consistent for its intended parameter, as both the long-difference (LD) and fixed effects (FE) estimands are weighted averages o…
Mojtaba Eslami
24 Jul 2026 · Statistics — Methodology
Let $p(x)$ be the joint density of variables $X$, and let $ψ(x)=\nabla_x\log p(x)$ be its score field. Geometry constructed from $p$ and $ψ$ alone cannot identify causal direction: structural models with the same observational distribution have the same score geometry. I develop an interventional analogue. A hard intervention $\operatorname{do}(X_k=ξ)$ does not merely reweight the joint law; it restricts the distribution to the submanifold ${x_k=ξ}$. Its…
Behrooz Moosavi Ramezanzadeh, Arie Beresteanu
23 Jul 2026 · Econometrics
Partial identification is often set aside in practice because the identification regions it delivers are too wide to be useful, pushing researchers toward strong assumptions that buy point identification at the cost of credibility. We show that a source of information already sitting in most interval-valued datasets can fix this without adding any assumption at all. When an outcome is reported only as an interval---because a data custodian bracketed, top…
23 Jul 2026 · Econometrics
Applied econometricians typically model each individual as having fixed outcomes under treatment and control and, in instrumental-variables (IV) settings, fixed treatment decisions under each value of the instrument. This paper asks what changes when outcomes and treatment allocations or choices are stochastic at the individual level. In the model, each individual has a stable (but possibly stochastic) response type consisting of two objects: a treatment…
Richard Grigorian
23 Jul 2026 · Econometrics
This paper develops a profiled sieve minimum-distance estimator for a semi-nonparametric differentiated-products demand model with micro-level choice data. Building on Berry and Haile (2024), the estimator uses within-market variation in consumer covariates to recover a flexible consumer-heterogeneity function and market-specific composite intercepts. Excluded price instruments then separate these intercepts into a flexible price-side function and struct…
23 Jul 2026 · Econometrics
Difference-in-differences (DID) are sometimes estimated with many pre-treatment periods. In such settings, the observed pre-treatment outcome evolutions provide direct information about the magnitude of shocks that could also occur after treatment. This paper proposes a simple inference procedure that uses those pre-trends as the reference distribution for the post-treatment DID. The procedure is closely related to existing conformal inference procedures…
Sabine Wildemann, Daniel Ambach
23 Jul 2026 · Statistics — Applications
Civic AI systems increasingly support democratic participation, yet interactions with them may reveal sensitive political views, creating tension between improving AI models and residents' expectations of privacy and consent. This study examines the conditions of transparency and user control under which Swiss residents are willing to donate their anonymized chatbot conversations to train an open-source AI model. A 2x2 between-subjects factorial design e…
22 Jul 2026 · Econometrics
Spillovers and interference pose fundamental challenges for causal inference, as treatment assigned to one unit may affect the outcome of others, violating the no-interference assumption underlying most empirical strategies. Existing approaches, based on partial interference, exposure mapping, spatial, network, or structural frameworks, typically rely on strong assumptions about interaction structures or require the existence of uncontaminated control un…
22 Jul 2026 · Econometrics
This paper develops efficient difference-in-differences (DID) estimation under partial interference with a cluster incremental propensity score (CIPS) policy. We define direct and spillover average treatment effects on the treated, establish their identification, and derive their efficient influence functions, from which we construct a cross-fitted estimator. Simulations confirm its finite-sample validity, and an application to China's New Rural Pension…
Jeffrey D. Michler, Anna Josephson, Elinor Benami, Patrick Behrer, Michael J. Cecil, Sydney Gourlay and 4 more
22 Jul 2026 · Econometrics
A central task in conducting impact evaluations is determining who or what was exposed to a treatment, when, and to what degree. These questions can be especially complex in geospatial settings, where many reasonable definitions of exposure may exist. This chapter introduces treatment geometry as a core concept in geospatial impact evaluation (GIE): the spatial and temporal footprint of a treatment as represented in data. How this footprint is defined sh…
22 Jul 2026 · Econometrics
Difference-in-differences with staggered adoption identifies group-time average treatment effects ATT(g,t) by comparing each cohort to units not yet treated, which avoids the "forbidden comparisons" that bias two-way fixed-effects estimators when effects are heterogeneous. This paper studies the covariate-conditional version of that object, tau_{g,t}(x), and estimates it with a fixed-effects causal forest. Within each (g,t) comparison block, the outcome…
Yingxing Li, Aureo De Paula, Weining Wang
21 Jul 2026 · Econometrics
This paper analyzes spillover effects in spatial (network) models when the neighborhood (adjacency) matrix is contaminated by measurement error from reporting, aggregation, or disclosure imperfections, leading to inconsistent estimation of network effects. We introduce a regularization framework for the latent network that allows for sparse and/or low-rank structure and accommodates potential correlation between measurement errors and outcomes. We propos…
Irene Aldridge
21 Jul 2026 · Econometrics
Building on the identity that expected regret equals the covariance between costs and decisions, this paper develops the complete derivative theory of the covariance regret functional. We derive the Gâteaux derivative, showing that the universal steepest-descent direction is the contrarian policy $-(c-\bar{c})$, while ascent yields momentum. For linear policies $\hatπ(c) = Ac+b$, the gradient is the cost covariance matrix $Σ_c$, with a zero Hessian imply…
21 Jul 2026 · Econometrics
We study the optimal design and analysis of experiments for estimating spillover effects. Assuming a known (e.g., linear) exposure mapping, we characterize the treatment-assignment distribution and regression-based estimator that minimize worst-case asymptotic variance against a broad class of distributions of unobservables. The design problem yields an intuitive solution in which the planner trades off spillover signal strength against diffusion of spil…
Masahiro Kato, Taka Kato
20 Jul 2026 · Econometrics
We propose one-step and two-step methods for policy learning with retrieval-augmented generation (RAG). We formulate RAG-based action selection under the potential outcome framework. In the two-step method, vector search retrieves action-specific neighboring evidence in an embedding space, the generator estimates conditional expected outcomes or their contrasts, and a plug-in rule selects an action. This formulation connects action-specific vector search…
20 Jul 2026 · Econometrics
A large literature uses exogenous variation to estimate how assignment to classrooms or other groups shapes social networks. Yet most of these analyses remain dyadic, treating each link in isolation, even though ties often form through triadic closure, as a friend of a friend also becomes a friend. Using fine-grained data on phone calls, text messages, physical co-location, and social-media ties, we estimate the network formation effects of randomly assi…
Fangzhou Yu
20 Jul 2026 · Econometrics
Instrumental-variables estimation increasingly pools many or high-dimensional instruments into a single machine-learned first stage, with rich controls partialled out. The resulting estimand, the partialled-out IV coefficient built from any signal of the instruments, is a signal-weighted average of the heterogeneous effects, which gives an opaque first stage a precise structural meaning. The average is convex whenever a covariance-monotonicity condition…
Fangzhou Yu
20 Jul 2026 · Econometrics
This paper proposes a robust nonparametric hypothesis test for the existence of heterogeneous treatment effects. We focus on the variance of the Conditional Average Treatment Effect (CATE) as a natural omnibus parameter, where a non-zero variance implies the presence of relevant heterogeneity. Standard inference for this parameter faces a fundamental theoretical challenge. On one hand, evaluating variance components on the same sample leads to null degen…
19 Jul 2026 · Econometrics
Unstructured data, such as images and text, are increasingly used in empirical economics. Since training machine-learning models on unstructured data is costly, economists often use off-the-shelf pre-trained deep learning models developed by computer scientists to extract embeddings, which are then used as covariates in target economic analyses. Despite the popularity of this practice, its theoretical foundations remain limited. There are two main diffic…
19 Jul 2026 · Econometrics
Standard pre-tests of normality on reduced-form innovations are insufficient to detect two or more Gaussian shocks and hence, the failure of identification in non-Gaussian SVARs. We instead propose a bootstrap-based approach to evaluate the asymptotic validity of this condition by measuring the divergence between the conditional bootstrap distribution of a maximum likelihood estimator and its limiting distribution under valid identification. We show that…
Onil Boussim
18 Jul 2026 · Econometrics
This paper develops a synthetic control estimator for compositional outcomes, vectors of shares generated by an underlying categorical process. Derived from a random utility model with interactive fixed effects on relative systematic utilities, the estimator maps compositions to log-odds, where the standard convex hull condition identifies the counterfactual as a convex combination of donor log-odds. Equivalently, it recovers the Fréchet barycenter under…
18 Jul 2026 · Econometrics
Experimenters often run pilots, but how much a small pilot should shape the main-wave design has no settled answer. This paper shows how noisy pilot evidence should guide treatment assignment probabilities in two-wave experiments. Two canonical rules mark the extremes. Balanced assignment guards against worst cases but ignores evidence that one arm is noisier. Feasible Neyman allocation adapts, but with a finite pilot it can overreact to noise, producing…
Yuhao Li, Haokun Lu, Xiaojun Song
18 Jul 2026 · Econometrics
We propose a unified Kernel Minimum Distance (KMD) framework for estimating and testing models defined by conditional moment restrictions. By embedding conditional moments into a Reproducing Kernel Hilbert Space (RKHS), we construct a closed-form $V$-statistic objective function that quantifies the distance from the restrictions. We establish the $\sqrt{n}$-consistency and asymptotic normality of the associated minimum distance estimator. Within this fra…
16 Jul 2026 · Econometrics
We introduce the mnorm package, which allows one to calculate conditional multivariate normal densities and probabilities and to differentiate them with respect to various parameters including covariances and integration limits. The package also supports parallel (multi-core) computing, handles non-normal marginals via the Gaussian copula, and provides fast routines for the calculation of bivariate and trivariate normal probabilities. The package is of s…
16 Jul 2026 · Econometrics
This paper studies when high-resolution signals aggregated to administrative units can recover unobserved local economic activity. We develop a reverse-regression framework for signals generated by activity but used to predict it at coarser spatial supports. The main theorem decomposes predictive elasticity into elementary elasticity, reverse-regression attenuation, and a spatial aggregation term driven by unit size and within-unit dispersion, showing ag…
15 Jul 2026 · Econometrics
Experiments may, by design, prevent one from observing on a single subject both the response to a treatment and to its absence. Because of this, marginal distributions for both cases may be observable but not their joint distribution, thus obscuring the distribution of the treatment effect. We examine the case where we impose that the treatment effect is nonnegative, also called monotone treatment response, a common assumption relevant to many practical…
15 Jul 2026 · Econometrics
Forecasting is a central goal of time-series analysis. This review centers on three major developments in recent AI-based time-series forecasting: transformers, large pretrained models for zero-shot forecasting, and diffusion-based generative forecasters. We connect these methods to the econometric tradition built around the vector autoregression (VAR) through a common object: the conditional distribution of the future given the past. The review is organ…
15 Jul 2026 · Econometrics
The paper investigates Bayesian Model Averaging and Selection (BMA/S) under non-standard stochastic assumptions, focusing on stochastic frontier analysis (SFA). We propose fast, reliable procedures for inference in the normal-exponential stochastic frontier model and examine whether accounting for asymmetric disturbances affects model averaging and/or selection outcomes relative to the conventional Gaussian-error BMA/S. Particular attention is given to m…

See all papers →

Corpus current to 27 Jul 2026. Updated daily from arXiv.