Every paper in arXiv econ.EM, its bibliography parsed from the LaTeX source, and a weighted citation graph over the whole corpus — linked to author profiles and publication records.
In this study, we construct the first orthogonal basis for additively consistent subspace in pairwise comparisons theory. This construction is based on our representation of additively consistent best approximations of skew-symmetric matrices with respect to a tensor basis having minimal support. The orthogonal basis establishes the logarithmic consistent projection for the orthogonal windowing of pairwise comparisons matrices. It is compared with the wi…
Many empirical investigations of long-run relations are based on cross-section regressions in averaged or long differenced data that effectively have the time dimension of a $T\times N$ panel compressed. We analyze a class of time-compressed I(1) data and show that they have magnified variability stemming from the fact that the cross-section variance of a non-stationary panel `fans out' with time. Cross-section regressions in time compressed data can pot…
This paper proposes a nonparametric Bayesian inference framework for partially identified discrete response models. The key observation is that these models map a reduced-form conditional choice probability to an identified set. Consequently, nonparametric Bayesian inference for the conditional probability mass function leads to Bayesian inference for the identified set. The inference framework nests conditional moment inequalities and linear systems wit…
This paper develops a new econometric framework to identify and estimate policy-relevant causal effects in contexts with endogenous selection into treatment and spillovers within single large networks or spatial settings. Conventional causal inference methods relying on either unconfoundedness or no-interference assumptions are generally inadequate in these scenarios. We introduce a Spillover Roy model that jointly models endogenous treatment selection a…
Multinomial choice models allow flexible substitution patterns but become computationally demanding with many alternatives or observations. With a fixed per-observation simulation budget, simulated maximum likelihood introduces simulation bias, while each optimization step requires a full-sample likelihood evaluation. We propose Stochastic Approximation with Unbiased Simulated Scores (SAUSS), an averaged stochastic approximation based on conditionally un…
Comparing economic models requires balancing fit against flexibility. Yet standard selection criteria often proxy flexibility using parameter counts, overlooking differences due to functional form and experimental design. This paper introduces a data compression approach for evaluating economic models. Following the Minimum Description Length principle, models are interpreted as codes and evaluated by how effectively they compress data. This perspective…
Richard Bernknopf, Leila Gonzales, Christpher Keane
25 Aug 2026 · Econometrics
Natural hazards are a nonmarket disamenity that affects an individual's search for employment resulting in a negative environmental impact that produces an economic inefficiency. We develop a seek-and-screen job search approach that uses a discrete choice simulation to examine how salary, crime, and natural hazard risk influence job choice. We model the job decision process as a series of elimination events using a Cox hazard model grounded in a Random U…
This article surveys two deeply connected literatures that approach the same fundamental problem from different disciplinary traditions: dynamic discrete choice (DDC) in structural econometrics and inverse reinforcement learning (IRL) in machine learning. Both seek to infer the preferences of decision makers from observed sequential behavior, assuming that individuals act to maximize an expected reward function within a dynamic, uncertain environment for…
We study dynamic linear panel data models in which the lagged outcomes and strictly exogenous covariates carry individual-specific coefficients and the time-varying errors have a flexible covariance structure. With a fixed number of time periods, we point-identify the joint distribution of the random coefficients and the structural errors under a distributional form of strict exogeneity, and propose a closed-form, multi-step estimator based on the invers…
Econometric models offer parsimonious but inexact approximations to data-generating processes. This paper studies the generalized method of moments (GMM) when exchangeable specification errors of order $n^{-1/2}$ contaminate the moment conditions. I develop estimators for the mean and variance of these specification errors, establishing their consistency in an asymptotic framework where the number of overidentifying restrictions grows with the sample siz…
Sebastián Souyris, Jason A. Duan, Anantaram Balakrishnan, Varun Rai
24 Aug 2026 · Econometrics
Problem definition: Solar electricity generation is a strategic component of energy portfolios designed to meet growing demand and reduce carbon emissions. Governments and municipalities encourage household photovoltaic (PV) adoption through upfront rebates and tax credits. Limited budgets require principled, data-driven policies that account for the drivers of adoption and the effects of incentives on adoption rates. Methodology/results: We develop a dy…
This paper develops a procedure for uncovering the common cyclical factors that drive a mix of stationary and nonstationary variables. The method does not require knowing which variables are nonstationary or the nature of the nonstationarity. An application to the FRED-MD macroeconomic dataset demonstrates that the approach offers similar benefits to those of traditional principal component analysis with some added advantages.
In panels with sample selection (that may occur due to attrition, nonresponse, etc.), the assumption of selection on observables (missing at random, MAR) is commonly imposed despite often being implausible. However, this assumption becomes testable when a refreshment sample is available. We develop a statistical test of MAR based on a distance between two estimated distributions: one obtained using the standard inverse probability weighting (IPW) that is…
Hamid Bekamiri, Jan Auernhammer, Milad Abbasiharofteh, Jesper Lindgaard Christensen
24 Aug 2026 · Econometrics
Green-patent indicators based on Cooperative Patent Classification Y02 tags increasingly inform research, industrial policy, and climate-oriented investment, yet their construct validity has not been evaluated at corpus scale. We ask whether Y02 classification errors are random measurement noise or systematic, direction-specific bias. We introduce an Error-as-Signal framework in which disagreement between an administrative label and an independent model…
Social scientists rely on hypothesis testing to support their research conclusions, but the standard tests are designed for testing one hypothesis rather than adjudicating between rival possibilities. We develop a new framework, "classification testing", as an alternative. Instead of selecting one hypothesis to test, a researcher conducting a classification test decides what qualitative distinctions ("classes") are most substantively relevant; the test e…
In the empirical sciences, significance thresholds often determine whether findings are treated as evidence of an effect. This paper studies how likely findings that just meet conventional significance thresholds are to remain significant in replications of the same sample size. To answer this question, we estimate the expected replication probability conditional on a given p-value among published studies for experimental economics, psychology, and socia…
Traditional statistical tests are poorly suited for the sequential evaluation of probabilistic forecast calibration. We address this limitation in macroeconomic forecasting by applying a new sequential testing method based on e-values. The e-value-based methodology enables anytime-valid inference. It allows practitioners to test against calibration continuously without invalidating statistical guarantees. To illustrate the framework's practical value, we…
Understanding the propagation of extreme events is important in many economic and environmental applications, yet most econometric methods for causal inference focus on average effects rather than tail behavior. This paper studies the identification of causal relations in extremes and derives resulting estimators and their asymptotic inference. As measure of causal dependence between extreme realizations of variables, we analyze the asymptotic behavior o…
Analysis of experimental data becomes challenging when the underlying population is connected by a network. Exposure mapping is a common tool in the literature for defining and estimating spillover effects. These mappings reduce the dimensionality of the estimand, thereby facilitating identifiability. It is assumed that this mapping is correctly specified, leaving the choice of the exposure mapping to the analyst. This makes estimators of the spillover e…
Spatial autoregressive inference is typically conditional on the spatial weights matrix, W, even though the underlying interaction structure is often unknown and empirical conclusions can be sensitive to its specification. This paper develops double/debiased machine learning inference for low-dimensional SAR parameters when the spatial interaction operator is learned flexibly from potentially endogenous characteristics. Within a maintained admissible sup…
During clinical trials evaluating a drug's effect on a survival endpoint, intermediate events often occur in addition to the primary event. The treatment can exert its effect on the primary endpoint along multiple pathways through intermediate events. Assumptions for identifying mediation effects, such as sequential ignorability in natural effects or the dismissible components condition in separable effects, fail because intermediate events act as treatm…
This paper studies the nonparametric identification and estimation of additively separable triangular models with continuous endogenous and instrumental variables, allowing for a nonseparable first-stage equation. Under the independence of instrumental variables and unobservables, we show that the outcome function possesses a closed-form expression as a functional of conditional cumulative distribution functions. The resulting plug-in estimators require…
We study a dynamic spatial panel model with observed regressors, interactive effects, and contemporaneous and lagged dependence in a large-$N$, fixed-$T$ framework. The spatial model constitutes an $N$-dimensional simultaneous-equations system. In this $N$-equation view, the interactive effects introduce $N$ unit-specific loading vectors. Estimating them individually when $T$ is fixed creates the type of incidental-parameters problem underlying Nickell b…
This paper studies quantile treatment and spillover effects in network experiments. Average spillover effects reveal how treating a unit's neighbors affects its outcome on average, but mask the heterogeneity of these effects across the outcome distribution. We define structural quantile effects that compare outcome quantiles between exposure states, characterizing how own treatment and exposure to treated neighbors affect different parts of the outcome d…
Asymptotic normality approximations often fail to hold for extremum estimators when the true value of the parameter is at or close to the boundary of a parameter space. I analyze and develop tests using a quasi-unconstrained estimator, which is asymptotically normal even when the true parameter vector is near or at the boundary. These results generalize previous work with this estimator by allowing for more types of constraints and showing how the method…
Rejection sampling requires a proposal that dominates the target by a known constant, generally unavailable for non-Gaussian state space models. We construct such a proposal for the latent state path, yielding independent exact smoothing draws and an unbiased likelihood estimator whose relative variance is at most $1/p-1$ per draw at acceptance probability $p$. The method covers scalar states with affine Gaussian dynamics and log-concave observation dens…
Short observational pricing panels can contain many observations while offering only a small number of distinct price movements. This paper studies the inferential consequences of that distinction in a synthetic data-generating process calibrated to a sparse pricing regime. We separate uncertainty conditional on a realised price trajectory from variation in estimation error across alternative trajectories generated by the same pricing process. In the bas…
We study causal moderation when treatment assignment is randomized but the moderator is not. We combine the parallel estimation framework with front-door adjustment to identify an average mediated treatment moderation effect. We apply this approach to 165 municipal assembly elections in Tokyo (1987-2023), where pamphlet positions are assigned by lottery and total pamphlet pages are mechanically determined by candidate set size and fixed municipal rules.…
Firms perform online experiments with multi-armed bandits to personalize what consumers are shown while balancing exploration and exploitation. However, third-parties can infer consumers' underlying segments from observing which banners, ads, or recommendations consumers receive. To control this inference, we propose a privacy risk budget that firms can set ex ante to bound such third party belief updating using differential privacy. To spend this privac…
Moment restrictions provide a flexible basis for quasi-Bayesian inference when a full likelihood is unavailable, but the weighting matrix in a quadratic moment criterion determines both the relative importance of the moments and the information scale of posterior updating. We propose curvature-calibrated quasi-Bayesian updating, which uses the inverse of the covariance (or long-run covariance) of the moment conditions evaluated at a self-consistent quasi…
Pairwise randomization can yield substantial efficiency gains in experiments. Yet methodological guidance cautions against pairwise randomization, especially in settings with attrition, partly because common practices for estimation (i.e., pair fixed effects) imply discarding data from incomplete pairs thus exacerbating data loss from attrition. This practice of dropping incomplete pairs reduces statistical power of tests as well as precision of estimate…
This paper develops asymptotic theory and feasible inference for unbounded-kernel order-k U-statistics under clustered sampling and weakly dependent time-series sampling. The analysis first builds the complete order-2 pipeline, moving from clustered data to exact m-dependence and then to near-epoch dependence. The same logic is subsequently extended to general order k greater than or equal to 2. Under clustered sampling, the theory allows arbitrary withi…
We study difference-in-differences (DiD) designs in which a binary treatment changes an endogenous time-varying (continuous, discrete, or mixed) mediator that in turn affects an outcome. Under our model assumptions, we show that the usual DiD estimand mixes the average direct effect on the treated, the average indirect effect, and a trend bias term. A two-way fixed effects (TWFE) regression that controls for the mediator does not recover the average dire…
We study how fast experimental designs can approach the semiparametric efficiency bound in finite samples, as measured by the excess variance of unadjusted treatment effect estimation. We prove an impossibility theorem: under weak conditions, no design can approach the variance bound uniformly over smooth outcome models unless covariate dimension $d \ll \log n$. Even in experiments with thousands of units, this permits only a handful of covariates. Motiv…
When parallel trends fails for some treated cohorts but not others, the average treatment effect on the treated (ATT), an average over all of them, is exactly the target that becomes hard to recover. We propose changing the estimand rather than defending it. The credible-subpopulation local ATT (LATT) is the effect for the subpopulation of cohorts whose parallel trends is credible, and it is point-identified under parallel trends for the selected cohorts…
This paper analyzes when choice probabilities reveal rankings of deterministic utility indices in semiparametric discrete choice models. It begins with binary choice, where quantile thresholds guarantee ranking recovery, and shows that such thresholds can arise either from behavioral departures from utility maximization (e.g., limited attention) under exchangeable unobservables, or from non-exchangeable unobservables under standard utility maximization.…
Empirical studies of peer effects often exploit conditional random assignment to peer groups within urns. We develop a GMM framework for estimation and inference in this setting. The framework separately identifies endogenous and contextual peer effects and nests tests of random peer-group assignment as a special case. It permits unknown heteroskedasticity and corrects finite-urn bias in variance estimation. Its asymptotic theory allows the number of pee…
We ask whether COVID-19 lockdown stringency altered national Olympic performance between Rio 2016 and Tokyo 2020, using the Oxford Stringency Index and the 99 countries that won a medal in either edition. As in \citet{liu2024}, mean performance is unaffected: stringency is insignificant in every OLS and ANOVA specification. The distribution is not. Among the 84 non-traditionally dominant nations, medal changes are three to six times more dispersed in hig…
We study the existence of a Regional Differential in rugby sevens: whether, in tournaments where no competing team enjoys formal home status, some national sides systematically over- or under-perform depending on where the event is staged. Using the universe of 2672 men's and women's matches from international rugby sevens tournaments played between 2016 and 2025, principally the World Rugby Sevens Series, but also the World Cup Sevens and the Olympic Ga…
We study how many observations are needed to determine the causal direction between two linearly related variables. Classical LiNGAM theory shows that independent non-Gaussian disturbances identify the direction, but does not quantify the difficulty when the causal effect is weak or the disturbances are nearly Gaussian. Let $β$ bound the absolute structural coefficient from below, let $ν$ measure each standardized disturbance's distance from Gaussianity,…