Every paper in arXiv econ.EM, its bibliography parsed from the LaTeX source, and a weighted citation graph over the whole corpus — linked to author profiles and publication records.
Policy learning aims to determine who should be treated based on individual characteristics. In high-stakes settings such as medicine and public policy where safety is a central concern, improving the average outcomes alone may not be sufficient: decision makers may also seek to protect individuals from harm, in line with the Hippocratic principle of “do no harm.” In this paper, we propose conformal policy learning (CPL), a policy learning procedure with…
I study dynamic treatment effects in panel data under staggered adoption when treatment timing depends jointly on unobserved time-invariant heterogeneity and time-varying pretreatment covariates, including lagged outcomes. Untreated potential outcomes follow a nonparametric dynamic panel model that allows flexible interactions between time-varying covariates and latent heterogeneity. I use pretreatment outcome histories to find individuals with similar t…
We investigate Empirical Bayes (EB) methods in the context of compound adaptive experiments, where the arm distribution in each experiment follows a normal distribution with an unknown mean that we seek to estimate. There are two main EB strategies: $g$-modeling, which estimates the prior by maximizing the marginal likelihood, and $f$-modeling, which derives posterior means directly from the empirical distribution of the observations. We show that $g$-mo…
This paper develops Wald inference for least-squares estimation of linear regression models on dyadic data, accommodating configurations where multiple observations share the same pair of units (e.g., directed flows, multilayer networks, and dyadic panels). We establish that the dyadic-robust Wald statistic is asymptotically $χ^2_q$ for an arbitrary nonrandom sequence of full-rank restrictions, under a single condition on the accumulation of dependence.…
Model robustness analysis estimates an effect across a multiverse of specifications that pools control sets identifying the declared estimand with sets that condition on mediators or colliders. We propose stating rival assumptions about contested controls as a small set of candidate causal graphs, enumerating the adjustment sets each graph licenses, and reporting robustness metrics conditional on each graph. A finite-mixture identity splits the licensed…
Anthropogenic forcing components follow different long-run paths, while persistent temperature change can involve distributional changes beyond the mean. Scalar regressions aggregate these components and retain only mean temperature, obscuring how distinct forcing paths relate to persistent distributional change. We develop new testing, estimation, and inference methods for long-run relations between an integrated predictor vector and a density-valued re…
Likelihood-based network models are often fitted under links' independence and low-order constraints, while empirical networks frequently exhibit systematic higher-order structures such as triangles and wedges, characterizing the observed clustering patterns. Real-world core-periphery networks such as the interbank market or the air transportation system represent key examples, with cores displaying complex and nonlinear features. We formalize Penalized…
The random-coefficient demand model of Berry, Levinsohn, and Pakes (1995) is commonly estimated by the generalized method of moments (GMM), using an unconditional moment restriction with a fixed set of instruments. Identification of the model, however, rests on a conditional moment restriction. The two are not equivalent: the unconditional restriction may admit additional parameter values. We construct a counterexample in which the model is identified by…
The Money Pump Index (MPI) of Echenique et al. (2011) measures the severity of consumer irrationality, but computing the exact mean and median MPI over all revealed preference cycles is NP-hard (Smeulders et al., 2013). Existing solutions rely on heuristic proxies, such as evaluating only shorter cycles or bounding the MPI. By framing revealed preferences as a directed graph, this paper projects choice violations onto fundamental cycle bases, which are m…
Predict-then-optimize methods such as Smart "Predict, then Optimize" (SPO+) of Elmachtoub and Grigas (2022) learn a mapping from contextual features to unknown edge costs and then solve the induced combinatorial problem on the predicted costs. This approach is powerful but relies on the predictive model being well specified: when the true cost-generating process is nonlinear in the features and the predictor is linear, SPO+'s performance degrades as the…
Two frequent approaches for identifying structural VARs are external instruments, which carry economic content but are often weak, and non-Gaussianity of the shocks which provides statistical identification but carries no economic meaning. We combine the two strategies in a single generalized method of moments framework that stacks proxy exclusion restrictions with higher-order moment conditions of the structural shocks. This hybrid approach point-identi…
Observed choice in a dynamic game mixes current profit with continuation value. A rival adds a second problem: the same comparison averages over the rival's equilibrium policy. Changing the primitive transition rewrites continuation technology; changing the rival's Markov policy, holding that law fixed, rewrites the mixture over rival-contingent payoffs. The two are not substitutes. For a rival-feature payoff of rank $K$, rank identification up to locati…
In social and economic networks linked agents often share connections in common. There are two competing explanations for this phenomenon. First, agents may have a structural taste for transitive links - the returns to linking may be higher if two agents share a common connection. Second, agents may assortatively match on unobserved attributes, a process called homophily. We study parameter identifiability in a simple model of dynamic network formation w…
A switchback experiment alternates an entire system between treatment and control over time. It is especially useful when interactions between units can undermine standard unit-level experiments. Switchback experiments are commonly implemented on fixed temporal grids, which impose highly structured restrictions on when treatment can switch. We study a broader class of random-duration switchbacks, in which treatment and control alternate across runs whose…
We develop a causal framework for matrix completion under missing not at random (MNAR) data. Drawing on synthetic controls from the econometric panel data literature, our approach relaxes two assumptions common in MNAR matrix completion: positivity and independence of observation indicators. Unlike traditional panel data models, which often require prescribed block-sparse geometries, our framework accommodates flexible, heterogeneous observation patterns…
This paper studies identification and inference in a triangular system with an endogenous regressor when exclusion restrictions are unavailable and the dependence between structural disturbances is modeled through an unrestricted control function. In this setting, standard orthogonality conditions do not deliver point identification, as the unknown control function can rationalize a wide range of structural coefficients. We show that identifying informat…
We derive the exact likelihood of the Berry (1994) share-form nested logit and show it contains a Jacobian term that depends on nest size and the nesting parameter. Omitting the term biases within-nest substitution estimates wherever choice sets vary across markets. Commonly-used count instruments for the within-nest share fail exclusion when product counts enter demand directly. The corrected likelihood identifies substitution without them, making crowd…
Probabilistic forecasting is central to decision-making under uncertainty, yet its methodological landscape has become increasingly fragmented across temporal and spatiotemporal forecasting, statistical modeling, machine learning, and deep generative modeling. This survey develops a unified perspective by organizing probabilistic forecasting methods according to where and how uncertainty is introduced into the forecasting pipeline. Our taxonomy connects…
This paper investigates the construction of moment restrictions in dyadic network formation models with unobserved individual heterogeneity under nontransferable utility. Using observed links and covariates from five-node pentads, we construct moment restrictions that do not depend on individual fixed effects. For a broad class of covariate specifications, the construction is minimal in the sense that it uses the least possible number of nodes and dyads.…
Including covariates alongside strictly monotone encodings can change a causal forest's treatment decisions without adding information. Random feature selection favors covariates represented by multiple columns. Under stated conditions, I show that this imbalance can persist as samples grow. Treatment effect components associated with other covariates are omitted, attenuated, or recovered depending on their inclusion probabilities and tree depth. Simulat…
Identification in vector autoregressions involves two distinct choices: how extensively to orthogonalize shocks and whether orthogonality is imposed sequentially or simultaneously. Generalized identification imposes no orthogonality; Sims (1980} imposes full orthogonality sequentially; Francis et al. (2026) impose full orthogonality simultaneously; and Buchwalter et al. (2026a) provide clustered partial orthogonalization sequentially. We fill the remaini…
We establish a sharp pathwise signal-to-noise criterion for quasi-maximum-likelihood (QML) estimation of a dominant breakpoint in the second-moment structure of a multivariate time series: the QML estimator is consistent whenever the between-regime contrast exceeds the within-regime fluctuation by an explicit factor, and below this threshold global recovery can fail. Two innovations drive the result. First, the framework is pathwise: no stochastic model…
Rafael Macalaba, Aivin V. Solatorio, Patrick Michael Brock, Olivier Dupriez
10 Sep 2026 · cs.CL
Development and humanitarian organizations produce and support surveys, administrative registries, and other data resources to inform research, policy, and operations, yet systematically identifying where these datasets are referenced remains difficult. Such references are dispersed across research papers, project documents, humanitarian reports, and other unstructured text, limiting both the ability to trace data use and to identify potential gaps in da…
Generative artificial intelligence changes how firms reach customers, but standard marketing data do not record how often users see and notice a firm's name in generated answers. We develop Generative Marketing Mix Modeling (GMMM) to estimate the causal effects of Generative Engine Optimization (GEO) and Generative Engine Marketing (GEM). For GEO, GMMM combines repeated generated answers with question counts, shares of use across generative systems, and…
Modeling the joint distribution of extreme values in high-dimensional financial time series is challenging because extremes are sparse and locally extreme observations are not necessarily extreme relative to their full marginal distribution. To address this, we introduce a time-dependent network Hüsler-Reiss model in which market-informed adjacency matrices determine how strongly observations contribute to the estimation. We propose binary and weighted s…
Graphical causal inference supplies a complete theory of efficient covariate adjustment for the average treatment effect: one adjustment set, computable from the graph, is optimal under every compatible distribution. We show that this is a property of the average treatment effect's inverse-prevalence weights, not of causal estimands in general. For the average treatment effect on the treated we index the efficiency bound by the adjustment set and derive…
We show how to optimally design experiments when the resulting data will be used to choose a welfare-maximizing policy subject to constraints. A decision maker seeks to maximize Bayes expected welfare by choosing a policy whose effects depend on an unknown finite-dimensional parameter. The decision maker has access to a first wave of experimental data with a fixed design but may choose the design of a second wave that will be collected before choosing th…
This paper studies identification in linear quantile panel models with unrestricted individual heterogeneity when the number of time periods is fixed and small. We impose strict exogeneity, whereby the conditional quantile restriction holds given the individual's complete regressor history and latent individual effect, but otherwise allow the disturbances to be arbitrarily dependent over time.
We recast statistical significance as a choice between making an immediate policy recommendation and deferring it until further evidence is collected. We show that the welfare-optimal decision corresponds, under minimax regret, to a statistical test whose level depends on the cost and precision of additional evidence. Inverting this rule, we introduce and recommend reporting the abstention-value (A-value) alongside traditional p-values to determine where…
This paper studies difference-in-differences with staggered adoption and a continuous, time-invariant dose. Each cohort-time comparison contains two margins. The level margin is the average treatment effect at realized doses. Under level parallel trends it equals the level contrast between the treated cohort and not-yet-treated controls. The response margin is the within-cohort slope of the outcome change on dose. It uses no controls, and its causal inte…
Many real-world policies and business interventions require assessing short-term effects to inform timely decisions, even though most causal inference methods focus on long-term average treatment effects. In this paper, we introduce average treatment effect localization (ATEL), which captures localized, short-term policy impacts in panel data settings with a single treated unit and provides early indicators of policy impact. To accommodate both time-vary…
Should researchers adjust for covariates in randomized experiments, and if so, how? The literature offers three distinct prescriptions: do not adjust because randomization guarantees unbiasedness; adjust for outcome-prognostic covariates to improve precision; or adjust for covariates imbalanced between treatment arms. These competing prescriptions create confusion and uncertainty. We develop a unified framework for decision and practice. Given available…
This paper presents a causal decision-making framework for estimating price elasticity in retail channels, a process typically confounded by promotions, competitor movements, and market frictions. Rather than forcing a calculation when data is ambiguous, the system introduces decision abstention (\textsc{wait}) as an active diagnostic tool rather than an estimation failure. Combining Double Machine Learning and conformal prediction, the tool evaluates wh…
When social assistance is scarce, should it prioritize households in greatest need or those expected to benefit most? Using panel data from the China Household Finance Survey, we distinguish allocation principles by combining predicted policy gains with entry into China's Minimum Living Standard Guarantee (Dibao). We estimate heterogeneous predicted gains in consumption and education and then recover the conditional priorities revealed by recipient selec…
Spatial treatments are interventions assigned to locations potentially distinct from those of the responding units. We study their optimal design under a general model in which a unit's response diminishes with distance to a treated site. Our estimand of interest is an “uncontaminated” effect equal to the average impact of a single intervention site over all hypothetical sites. We propose a novel design based on a Matérn point process which separates tre…
Yonghong Zhang, Ricardo Correia, Isabel M. Parra, Yong Xie
7 Sep 2026 · Artificial Intelligence
Existing causal-inference benchmarks for LLMs mostly score method descriptions or whether generated code runs, not whether the executed workflow recovers the target causal estimate. CausalVerify studies this verification problem for structured econometric causal-estimation workflows by separating realistic interpretation from verifiable computation. It pairs 259 published economics papers (reconstructed research question, data description, institutional…
Heavy-tailed and skewed outcomes are common in the randomized experiments and observational studies used to estimate heterogeneous treatment effects, yet the mean-squared-error criterion that guides splitting in honest causal trees is sensitive to the extreme values they generate. Building on the causal forest framework (Athey and Imbens, 2016; Wager and Athey, 2018), we introduce the Median Squared Deviation (MSD) criterion, which replaces the leafwise…
A/B tests are standard in firm decision making. In the standard pipeline, experimental data is converted to a deployment decision by applying a t-test of the difference in means (the lift) and deploying the treatment if lift is positive and statistically significant. This common workflow answers the wrong question. We argue that firms need a decision rule for economic payoffs in the future deployment environment, not a test of equality in the experimenta…
We develop a filter for time series, defined at each time $t$ as the minimizer of a discounted convex combination of observed and expected losses. The filter can be estimated by simulation to an arbitrary level of accuracy in $O(1)$ flops at each time point $t$ and can be run for all values $t=1,...,T$ in parallel. These methods are applied to robustly compute a preaveraged price process from the more than 1.5 million trades made on a single financial as…
Estimating the first stage of an instrumental variables (IV) model with the least absolute shrinkage and selection operator (LASSO) requires choosing a dictionary of technical instruments and a penalty level. First-order asymptotic theory offers no guidance on these choices, as any consistent implementation yields a structural parameter estimator with the same limiting distribution. In finite samples, however, these choices can have a substantial impact…