Every paper in arXiv econ.EM, its bibliography parsed from the LaTeX source, and a weighted citation graph over the whole corpus — linked to author profiles and publication records.
When natural disasters strike, individuals, communities, and even entire countries can suffer. Researchers have studied the impacts of disasters on various factors of interest, from mental health, to poverty, to economic activity. However, the impact of disasters on the nonprofit sector is understudied despite the nonprofit sector's perhaps surprising role in local or national economies as well as its role in disaster response and recovery. Thus, we stud…
We give an exact randomization-based confidence set for the average treatment effect (ATE) in matched-pair studies with a binary outcome, requiring neither monotonicity nor any distributional assumption beyond the within-pair coin flip. At its core is an analytic solution to the worst-case allocation of attributable effects: two binomial-symmetry lemmas identify the pattern hardest to reject as a single boundary corner, so testing null hypotheses needs n…
In the presence of interference, where the treatment assigned to one unit can affect the outcomes of others, many causal estimands depend on the treatment-assignment policy under which the experiment is conducted. This policy dependence creates a fundamental challenge for off-policy estimation, where the goal is to estimate causal quantities under a hypothetical intervention policy different from the one used to collect data. We study this problem of off…
We establish the consistency and asymptotic normality of a two-step estimator of conditional expectiles in the context of conditional scale models. We first estimate the conditional variance parameters by quasi-maximum likelihood and then compute the unconditional expectile of the innovations using the empirical distribution of the standardized residuals. We show how replacing true innovations with standardized residuals affects the asymptotic variances…
To study a scalar parameter, a researcher may consider multiple research designs. Based on the evidence across designs, the researcher may wish to formulate a headline estimate of the parameter. I examine how to choose this headline when it is unclear which design is most appropriate for studying the parameter. I model this setting by assuming that (i) exactly one of the designs is valid for the parameter and (ii) the researcher has ambiguity about which…
Inference in linear regression commonly treats OLS residuals as proxies for unobserved errors. This approximation can fail when the regression projection is nonlocal relative to the error-dependence structure. Residualization then shifts covariance information across observations and clusters, while conventional heteroskedasticity-consistent (HC) and cluster-robust variance estimators (CRVE) retain only diagonal or within-cluster residual moments and may…
Policies with a common objective and implementation date may differ in details or context. We distinguish the aggregate average treatment effect on the treated (ATT) from sub-aggregate ATTs defined by implementation cohort, jurisdiction, period, or policy type. UN-DID and DID-INT, two DiD estimators that construct jurisdiction-by-time effects, estimate these ATTs under parallel-trends conditions matched to the aggregation. In CPS placebo-law simulations,…
Market efficiency relies fundamentally on stable liquidity. Consequently, forecasting liquidity dynamics is a priority for both investors and regulators. We introduce a new tail-risk metric, Illiquidity-at-Risk (IlliQaR), designed to quantify the magnitude of extreme liquidity dry-ups. Relying upon the realized Amihud (a precise illiquidity measurement derived from high-frequency data as the ratio of realized volatility to trading volume) we assess the p…
Valentina Norambuena-Guzman, Cong Chen, Lang Tong, Timothy D. Mount
1 Sep 2026 · eess.SY
In a network with ramp-limited generators and inaccurate net-demand forecasts, practical rolling-window dispatch can drive locational marginal prices (LMPs) below generators' bid-in offers. In such cases, out-of-market (OOM) settlements are used to compensate generators and maintain dispatch-following incentives, but OOM can have negative consequences, including nontransparent real-time price signals, discriminatory compensation, and incentives for untru…
We propose the first manipulation test designed for boundary discontinuity designs (BDDs) with general boundary shapes. A BDD is a multidimensional extension of the regression discontinuity design (RDD) in which treatment assignment is determined by whether the multidimensional running variable crosses a lower-dimensional boundary set. The test avoids multivariate density estimation and builds on the observation that, in the absence of manipulation, obse…
We study identification and estimation of moments of random coefficients in short linear panels, allowing the number of heterogeneous coefficients to exceed the number of equations observed for each unit. Under moment homogeneity, different regressor histories impose restrictions on the same moment vector. We give necessary and sufficient conditions for these restrictions to identify moments of a given order, stated in terms of the row spaces generated b…
When do text embeddings work as inputs to empirical analysis? Their use rests on an assumption: that we can trade text for its low-dimensional embedding, and lose little in doing so. I make that assumption precise under a generative model in which documents are mixtures of latent topics. I study two uses---clustering units in embedding space and controlling for high-dimensional text. A cluster of embeddings is a set of documents with similar topic mixtur…
Foundation models promise accurate forecasts with little or no task-specific training, but whether they can replace models designed specifically for electricity price forecasting remains unclear. We compare nine variants from five foundation model families, evaluated in zero-shot mode, with two state-of-the-art electricity price forecasting benchmarks in Germany, Poland, and Spain over 2021-2025. Their performance is assessed in terms of point and probab…
Stablecoins, typically pegged to fiat currencies, cannot achieve true stability because they inherit fluctuations in the underlying unit of account. To overcome this limitation, we introduce a stablecoin pegged to the Maximum Likelihood Value (MLV), a newly defined unit of account derived as the most probable configuration of latent real-value movements that explains observed nominal-value (price) changes. Grounded in inferential statistics and modern po…
Industrial recommenders give new content initial views through budgeted exploration, then use early performance to decide further delivery. On many short-video platforms, exploration is the primary way new videos reach viewers. Viewer-side tests measure consumption; the published budget objectives we review omit creator response. We analyze four experiments on a major short-video platform. An eight-month creator ablation finds production exploration rais…
In this paper, we study causal non-causal state space models to model time series characterised by a local explosive increase followed by a sharp decrease such as stock prices. To motivate the use of causal non-causal state space models, we show that the causal non-causal convolution autoregressive model introduced by Gourieroux and Zakoian (2017) can be consistent with the rational expectations stock price model. As in a causal state space model, a cent…
We present a suite of R packages for macroeconomic forecasting that leverages advanced Bayesian, structural, multivariate, dynamic, hierarchical, non-linear, and non-Gaussian models. The suite enables both structural and predictive analyses, and is adapted to time series data across various types, dimensions, and sampling frequencies. Each additional feature increases computational complexity. To address this challenge, our software design incorporates a…
We study identification of two-way unobserved heterogeneity in the nonparametric panel regression $G_{it}=g(α_i,γ_t)+\varepsilon_{it}$, where identification of the latent types reduces to constructing identified, injective proxies for them. To this end we consider the singular value decomposition (SVD) of the bivariate regression function $g(α,γ)$ on a product domain $Ω_α\timesΩ_γ$, whose left singular functions ${u_r}$ serve as proxies for the unobserve…
Recent work encourages political scientists to move from post-only toward within-subject designs for improved precision from repeated measurements. We formalize a potential-outcomes framework for two-period within-subject designs that allows for unequal allocation and heterogeneous treatment and carryover effects. We characterize the pooled estimator and evaluate the carryover test used to justify pooling. We find: first, pooling identifies the average t…
Which visual choices make a post perform better? A growing literature answers this question with pooled coefficients estimated across many creators, which platforms translate into creative recommendations. We show that these coefficients blend two distinct patterns that can point in opposite directions for the same attribute. The first, audience preference, arises because creators who favor a style attract differently composed audiences, so their posts p…
We develop a novel method of inference for network-dependent high-dimensional random vectors. Dependence is characterized via a functional dependence measure based on graph distance, allowing the approximation theory to capture the interaction between the decay of dependence and the growth of network neighborhoods. We establish Gaussian approximation results for the maximum norm under finite-moment and sub-Weibull conditions, providing explicit condition…
In this study, we construct the first orthogonal basis for additively consistent subspace in pairwise comparisons theory. This construction is based on our representation of additively consistent best approximations of skew-symmetric matrices with respect to a tensor basis having minimal support. The orthogonal basis establishes the logarithmic consistent projection for the orthogonal windowing of pairwise comparisons matrices. It is compared with the wi…
Many empirical investigations of long-run relations are based on cross-section regressions in averaged or long differenced data that effectively have the time dimension of a $T\times N$ panel compressed. We analyze a class of time-compressed I(1) data and show that they have magnified variability stemming from the fact that the cross-section variance of a non-stationary panel `fans out' with time. Cross-section regressions in time compressed data can pot…
This paper proposes a nonparametric Bayesian inference framework for partially identified discrete response models. The key observation is that these models map a reduced-form conditional choice probability to an identified set. Consequently, nonparametric Bayesian inference for the conditional probability mass function leads to Bayesian inference for the identified set. The inference framework nests conditional moment inequalities and linear systems wit…
This paper develops a new econometric framework to identify and estimate policy-relevant causal effects in contexts with endogenous selection into treatment and spillovers within single large networks or spatial settings. Conventional causal inference methods relying on either unconfoundedness or no-interference assumptions are generally inadequate in these scenarios. We introduce a Spillover Roy model that jointly models endogenous treatment selection a…
Multinomial choice models allow flexible substitution patterns but become computationally demanding with many alternatives or observations. With a fixed per-observation simulation budget, simulated maximum likelihood introduces simulation bias, while each optimization step requires a full-sample likelihood evaluation. We propose Stochastic Approximation with Unbiased Simulated Scores (SAUSS), an averaged stochastic approximation based on conditionally un…
Comparing economic models requires balancing fit against flexibility. Yet standard selection criteria often proxy flexibility using parameter counts, overlooking differences due to functional form and experimental design. This paper introduces a data compression approach for evaluating economic models. Following the Minimum Description Length principle, models are interpreted as codes and evaluated by how effectively they compress data. This perspective…
Richard Bernknopf, Leila Gonzales, Christopher Keane
25 Aug 2026 · Econometrics
Natural hazards are a nonmarket disamenity that affects an individual's search for employment resulting in a negative environmental impact that produces an economic inefficiency. We develop a seek-and-screen job search approach that uses a discrete choice simulation to examine how salary, crime, and natural hazard risk influence job choice. We model the job decision process as a series of elimination events using a Cox hazard model grounded in a Random U…
This article surveys two deeply connected literatures that approach the same fundamental problem from different disciplinary traditions: dynamic discrete choice (DDC) in structural econometrics and inverse reinforcement learning (IRL) in machine learning. Both seek to infer the preferences of decision makers from observed sequential behavior, assuming that individuals act to maximize an expected reward function within a dynamic, uncertain environment for…
We study dynamic linear panel data models in which the lagged outcomes and strictly exogenous covariates carry individual-specific coefficients and the time-varying errors have a flexible covariance structure. With a fixed number of time periods, we point-identify the joint distribution of the random coefficients and the structural errors under a distributional form of strict exogeneity, and propose a closed-form, multi-step estimator based on the invers…
Econometric models offer parsimonious but inexact approximations to data-generating processes. This paper studies the generalized method of moments (GMM) when exchangeable specification errors of order $n^{-1/2}$ contaminate the moment conditions. I develop estimators for the mean and variance of these specification errors, establishing their consistency in an asymptotic framework where the number of overidentifying restrictions grows with the sample siz…
Sebastián Souyris, Jason A. Duan, Anantaram Balakrishnan, Varun Rai
24 Aug 2026 · Econometrics
Problem definition: Solar electricity generation is a strategic component of energy portfolios designed to meet growing demand and reduce carbon emissions. Governments and municipalities encourage household photovoltaic (PV) adoption through upfront rebates and tax credits. Limited budgets require principled, data-driven policies that account for the drivers of adoption and the effects of incentives on adoption rates. Methodology/results: We develop a dy…
This paper develops a procedure for uncovering the common cyclical factors that drive a mix of stationary and nonstationary variables. The method does not require knowing which variables are nonstationary or the nature of the nonstationarity. An application to the FRED-MD macroeconomic dataset demonstrates that the approach offers similar benefits to those of traditional principal component analysis with some added advantages.
In panels with sample selection (that may occur due to attrition, nonresponse, etc.), the assumption of selection on observables (missing at random, MAR) is commonly imposed despite often being implausible. However, this assumption becomes testable when a refreshment sample is available. We develop a statistical test of MAR based on a distance between two estimated distributions: one obtained using the standard inverse probability weighting (IPW) that is…
Hamid Bekamiri, Jan Auernhammer, Milad Abbasiharofteh, Jesper Lindgaard Christensen
24 Aug 2026 · Econometrics
Green-patent indicators based on Cooperative Patent Classification Y02 tags increasingly inform research, industrial policy, and climate-oriented investment, yet their construct validity has not been evaluated at corpus scale. We ask whether Y02 classification errors are random measurement noise or systematic, direction-specific bias. We introduce an Error-as-Signal framework in which disagreement between an administrative label and an independent model…
Social scientists rely on hypothesis testing to support their research conclusions, but the standard tests are designed for testing one hypothesis rather than adjudicating between rival possibilities. We develop a new framework, "classification testing", as an alternative. Instead of selecting one hypothesis to test, a researcher conducting a classification test decides what qualitative distinctions ("classes") are most substantively relevant; the test e…
In the empirical sciences, significance thresholds often determine whether findings are treated as evidence of an effect. This paper studies how likely findings that just meet conventional significance thresholds are to remain significant in replications of the same sample size. To answer this question, we estimate the expected replication probability conditional on a given p-value among published studies for experimental economics, psychology, and socia…
Traditional statistical tests are poorly suited for the sequential evaluation of probabilistic forecast calibration. We address this limitation in macroeconomic forecasting by applying a new sequential testing method based on e-values. The e-value-based methodology enables anytime-valid inference. It allows practitioners to test against calibration continuously without invalidating statistical guarantees. To illustrate the framework's practical value, we…
Understanding the propagation of extreme events is important in many economic and environmental applications, yet most econometric methods for causal inference focus on average effects rather than tail behavior. This paper studies the identification of causal relations in extremes and derives resulting estimators and their asymptotic inference. As measure of causal dependence between extreme realizations of variables, we analyze the asymptotic behavior o…
Analysis of experimental data becomes challenging when the underlying population is connected by a network. Exposure mapping is a common tool in the literature for defining and estimating spillover effects. These mappings reduce the dimensionality of the estimand, thereby facilitating identifiability. It is assumed that this mapping is correctly specified, leaving the choice of the exposure mapping to the analyst. This makes estimators of the spillover e…