Every paper in arXiv econ.EM, its bibliography parsed from the LaTeX source, and a weighted citation graph over the whole corpus — linked to author profiles and publication records.
Richard Bernknopf, Leila Gonzales, Christpher Keane
25 Aug 2026 · Econometrics
Natural hazards are a nonmarket disamenity that affects an individual's search for employment resulting in a negative environmental impact that produces an economic inefficiency. We develop a seek-and-screen job search approach that uses a discrete choice simulation to examine how salary, crime, and natural hazard risk influence job choice. We model the job decision process as a series of elimination events using a Cox hazard model grounded in a Random U…
This article surveys two deeply connected literatures that approach the same fundamental problem from different disciplinary traditions: dynamic discrete choice (DDC) in structural econometrics and inverse reinforcement learning (IRL) in machine learning. Both seek to infer the preferences of decision makers from observed sequential behavior, assuming that individuals act to maximize an expected reward function within a dynamic, uncertain environment for…
We study dynamic linear panel data models in which the lagged outcomes and strictly exogenous covariates carry individual-specific coefficients and the time-varying errors have a flexible covariance structure. With a fixed number of time periods, we point-identify the joint distribution of the random coefficients and the structural errors under a distributional form of strict exogeneity, and propose a closed-form, multi-step estimator based on the invers…
Econometric models offer parsimonious but inexact approximations to data-generating processes. This paper studies the generalized method of moments (GMM) when exchangeable specification errors of order $n^{-1/2}$ contaminate the moment conditions. I develop estimators for the mean and variance of these specification errors, establishing their consistency in an asymptotic framework where the number of overidentifying restrictions grows with the sample siz…
Sebastián Souyris, Jason A. Duan, Anantaram Balakrishnan, Varun Rai
24 Aug 2026 · Econometrics
Problem definition: Solar electricity generation is a strategic component of energy portfolios designed to meet growing demand and reduce carbon emissions. Governments and municipalities encourage household photovoltaic (PV) adoption through upfront rebates and tax credits. Limited budgets require principled, data-driven policies that account for the drivers of adoption and the effects of incentives on adoption rates. Methodology/results: We develop a dy…
This paper develops a procedure for uncovering the common cyclical factors that drive a mix of stationary and nonstationary variables. The method does not require knowing which variables are nonstationary or the nature of the nonstationarity. An application to the FRED-MD macroeconomic dataset demonstrates that the approach offers similar benefits to those of traditional principal component analysis with some added advantages.
In panels with sample selection (that may occur due to attrition, nonresponse, etc.), the assumption of selection on observables (missing at random, MAR) is commonly imposed despite often being implausible. However, this assumption becomes testable when a refreshment sample is available. We develop a statistical test of MAR based on a distance between two estimated distributions: one obtained using the standard inverse probability weighting (IPW) that is…
Hamid Bekamiri, Jan Auernhammer, Milad Abbasiharofteh, Jesper Lindgaard Christensen
24 Aug 2026 · Econometrics
Green-patent indicators based on Cooperative Patent Classification Y02 tags increasingly inform research, industrial policy, and climate-oriented investment, yet their construct validity has not been evaluated at corpus scale. We ask whether Y02 classification errors are random measurement noise or systematic, direction-specific bias. We introduce an Error-as-Signal framework in which disagreement between an administrative label and an independent model…
Social scientists rely on hypothesis testing to support their research conclusions, but the standard tests are designed for testing one hypothesis rather than adjudicating between rival possibilities. We develop a new framework, "classification testing", as an alternative. Instead of selecting one hypothesis to test, a researcher conducting a classification test decides what qualitative distinctions ("classes") are most substantively relevant; the test e…
In the empirical sciences, significance thresholds often determine whether findings are treated as evidence of an effect. This paper studies how likely findings that just meet conventional significance thresholds are to remain significant in replications of the same sample size. To answer this question, we estimate the expected replication probability conditional on a given p-value among published studies for experimental economics, psychology, and socia…
Traditional statistical tests are poorly suited for the sequential evaluation of probabilistic forecast calibration. We address this limitation in macroeconomic forecasting by applying a new sequential testing method based on e-values. The e-value-based methodology enables anytime-valid inference. It allows practitioners to test against calibration continuously without invalidating statistical guarantees. To illustrate the framework's practical value, we…
Understanding the propagation of extreme events is important in many economic and environmental applications, yet most econometric methods for causal inference focus on average effects rather than tail behavior. This paper studies the identification of causal relations in extremes and derives resulting estimators and their asymptotic inference. As measure of causal dependence between extreme realizations of variables, we analyze the asymptotic behavior o…
Analysis of experimental data becomes challenging when the underlying population is connected by a network. Exposure mapping is a common tool in the literature for defining and estimating spillover effects. These mappings reduce the dimensionality of the estimand, thereby facilitating identifiability. It is assumed that this mapping is correctly specified, leaving the choice of the exposure mapping to the analyst. This makes estimators of the spillover e…
Spatial autoregressive inference is typically conditional on the spatial weights matrix, W, even though the underlying interaction structure is often unknown and empirical conclusions can be sensitive to its specification. This paper develops double/debiased machine learning inference for low-dimensional SAR parameters when the spatial interaction operator is learned flexibly from potentially endogenous characteristics. Within a maintained admissible sup…
During clinical trials evaluating a drug's effect on a survival endpoint, intermediate events often occur in addition to the primary event. The treatment can exert its effect on the primary endpoint along multiple pathways through intermediate events. Assumptions for identifying mediation effects, such as sequential ignorability in natural effects or the dismissible components condition in separable effects, fail because intermediate events act as treatm…
This paper studies the nonparametric identification and estimation of additively separable triangular models with continuous endogenous and instrumental variables, allowing for a nonseparable first-stage equation. Under the independence of instrumental variables and unobservables, we show that the outcome function possesses a closed-form expression as a functional of conditional cumulative distribution functions. The resulting plug-in estimators require…
We study a dynamic spatial panel model with observed regressors, interactive effects, and contemporaneous and lagged dependence in a large-$N$, fixed-$T$ framework. The spatial model constitutes an $N$-dimensional simultaneous-equations system. In this $N$-equation view, the interactive effects introduce $N$ unit-specific loading vectors. Estimating them individually when $T$ is fixed creates the type of incidental-parameters problem underlying Nickell b…
This paper studies quantile treatment and spillover effects in network experiments. Average spillover effects reveal how treating a unit's neighbors affects its outcome on average, but mask the heterogeneity of these effects across the outcome distribution. We define structural quantile effects that compare outcome quantiles between exposure states, characterizing how own treatment and exposure to treated neighbors affect different parts of the outcome d…
Asymptotic normality approximations often fail to hold for extremum estimators when the true value of the parameter is at or close to the boundary of a parameter space. I analyze and develop tests using a quasi-unconstrained estimator, which is asymptotically normal even when the true parameter vector is near or at the boundary. These results generalize previous work with this estimator by allowing for more types of constraints and showing how the method…
Rejection sampling requires a proposal that dominates the target by a known constant, generally unavailable for non-Gaussian state space models. We construct such a proposal for the latent state path, yielding independent exact smoothing draws and an unbiased likelihood estimator whose relative variance is at most $1/p-1$ per draw at acceptance probability $p$. The method covers scalar states with affine Gaussian dynamics and log-concave observation dens…
Short observational pricing panels can contain many observations while offering only a small number of distinct price movements. This paper studies the inferential consequences of that distinction in a synthetic data-generating process calibrated to a sparse pricing regime. We separate uncertainty conditional on a realised price trajectory from variation in estimation error across alternative trajectories generated by the same pricing process. In the bas…
We study causal moderation when treatment assignment is randomized but the moderator is not. We combine the parallel estimation framework with front-door adjustment to identify an average mediated treatment moderation effect. We apply this approach to 165 municipal assembly elections in Tokyo (1987-2023), where pamphlet positions are assigned by lottery and total pamphlet pages are mechanically determined by candidate set size and fixed municipal rules.…
Firms perform online experiments with multi-armed bandits to personalize what consumers are shown while balancing exploration and exploitation. However, third-parties can infer consumers' underlying segments from observing which banners, ads, or recommendations consumers receive. To control this inference, we propose a privacy risk budget that firms can set ex ante to bound such third party belief updating using differential privacy. To spend this privac…
Moment restrictions provide a flexible basis for quasi-Bayesian inference when a full likelihood is unavailable, but the weighting matrix in a quadratic moment criterion determines both the relative importance of the moments and the information scale of posterior updating. We propose curvature-calibrated quasi-Bayesian updating, which uses the inverse of the covariance (or long-run covariance) of the moment conditions evaluated at a self-consistent quasi…
Pairwise randomization can yield substantial efficiency gains in experiments. Yet methodological guidance cautions against pairwise randomization, especially in settings with attrition, partly because common practices for estimation (i.e., pair fixed effects) imply discarding data from incomplete pairs thus exacerbating data loss from attrition. This practice of dropping incomplete pairs reduces statistical power of tests as well as precision of estimate…
This paper develops asymptotic theory and feasible inference for unbounded-kernel order-k U-statistics under clustered sampling and weakly dependent time-series sampling. The analysis first builds the complete order-2 pipeline, moving from clustered data to exact m-dependence and then to near-epoch dependence. The same logic is subsequently extended to general order k greater than or equal to 2. Under clustered sampling, the theory allows arbitrary withi…
We study difference-in-differences (DiD) designs in which a binary treatment changes an endogenous time-varying (continuous, discrete, or mixed) mediator that in turn affects an outcome. Under our model assumptions, we show that the usual DiD estimand mixes the average direct effect on the treated, the average indirect effect, and a trend bias term. A two-way fixed effects (TWFE) regression that controls for the mediator does not recover the average dire…
We study how fast experimental designs can approach the semiparametric efficiency bound in finite samples, as measured by the excess variance of unadjusted treatment effect estimation. We prove an impossibility theorem: under weak conditions, no design can approach the variance bound uniformly over smooth outcome models unless covariate dimension $d \ll \log n$. Even in experiments with thousands of units, this permits only a handful of covariates. Motiv…
When parallel trends fails for some treated cohorts but not others, the average treatment effect on the treated (ATT), an average over all of them, is exactly the target that becomes hard to recover. We propose changing the estimand rather than defending it. The credible-subpopulation local ATT (LATT) is the effect for the subpopulation of cohorts whose parallel trends is credible, and it is point-identified under parallel trends for the selected cohorts…
This paper analyzes when choice probabilities reveal rankings of deterministic utility indices in semiparametric discrete choice models. It begins with binary choice, where quantile thresholds guarantee ranking recovery, and shows that such thresholds can arise either from behavioral departures from utility maximization (e.g., limited attention) under exchangeable unobservables, or from non-exchangeable unobservables under standard utility maximization.…
Empirical studies of peer effects often exploit conditional random assignment to peer groups within urns. We develop a GMM framework for estimation and inference in this setting. The framework separately identifies endogenous and contextual peer effects and nests tests of random peer-group assignment as a special case. It permits unknown heteroskedasticity and corrects finite-urn bias in variance estimation. Its asymptotic theory allows the number of pee…
We ask whether COVID-19 lockdown stringency altered national Olympic performance between Rio 2016 and Tokyo 2020, using the Oxford Stringency Index and the 99 countries that won a medal in either edition. As in \citet{liu2024}, mean performance is unaffected: stringency is insignificant in every OLS and ANOVA specification. The distribution is not. Among the 84 non-traditionally dominant nations, medal changes are three to six times more dispersed in hig…
We study the existence of a Regional Differential in rugby sevens: whether, in tournaments where no competing team enjoys formal home status, some national sides systematically over- or under-perform depending on where the event is staged. Using the universe of 2672 men's and women's matches from international rugby sevens tournaments played between 2016 and 2025, principally the World Rugby Sevens Series, but also the World Cup Sevens and the Olympic Ga…
We study how many observations are needed to determine the causal direction between two linearly related variables. Classical LiNGAM theory shows that independent non-Gaussian disturbances identify the direction, but does not quantify the difficulty when the causal effect is weak or the disturbances are nearly Gaussian. Let $β$ bound the absolute structural coefficient from below, let $ν$ measure each standardized disturbance's distance from Gaussianity,…
Route and activity choice are connected levels of a common sequential mobility decision problem: activity choice determines what people do, where, and when, while route choice governs how they move between activities. This review develops a unified framework connecting transportation choice modeling with inverse reinforcement learning (IRL) and imitation learning (IL). Under explicit assumptions, recursive logit, logit dynamic discrete choice, and maximu…
We propose a family of control variate estimators for variance reduction in design-based survey sampling and causal inference, with and without interference. In these settings, inverse probability weighting (IPW) estimators are widely used, but may have large variance when sampling, treatment, or exposure probabilities are small. Building on the observation that several common estimators, including the Hajek, normalized, and augmented inverse probability…
This study investigates the methodological and theoretical properties of session handover in applications that use large language models. A task may continue in a new session when the context reaches the model's input limit, when the application restarts, or when another agent is asked to finish the task. The application must then decide which information from the earlier session to pass on. We formulate handover as the transfer of a task-relative in-con…
This paper employs a Threshold Bayesian Vector Autoregression (TBVAR) to estimate the regime-dependent macroeconomic effects of capital regulation in Hungary. Using the Factor-based Index of Systemic Stress (FISS) as the threshold variable, the model identifies normal and stress regimes consistent with the occasionally binding constraints literature. The TBVAR offers a practical multivariate alternative to Growth-at-Risk for data-constrained economies. G…
We develop a method for estimating and testing a single block of a macroeconomic model with heterogeneous agents, without placing assumptions on the structure of the rest of the economy. In a large class of models, individual agents' decisions depend on the macroeconomy only through their expectations of the evolution of a finite-dimensional vector of "sufficient statistics" (e.g., asset returns or aggregate earnings). Our estimator selects the structura…
Limited dependent variable models are central to empirical economics, but likelihood-based inference is infeasible when likelihoods involve high-dimensional integration over latent variables. This paper proposes Stochastically Estimated Gradient Ascent (SEGA), a scalable estimation approach for limited dependent variable models. Using Fisher's identity, SEGA replaces the intractable likelihood score with an unbiased augmented-data score evaluated at a si…