Every paper in arXiv econ.EM, its bibliography parsed from the LaTeX source, and a weighted citation graph over the whole corpus — linked to author profiles and publication records.
Machine learning (ML) techniques are increasingly drawing interest in the choice modelling (CM) field. The focus has primarily been on comparing the performance of these contrasting approaches or on improving behavioural insights for ML techniques, rather than translating ideas from one field into the other. In the present paper, we specifically focus on knowledge transfer from ML into CM in the context of model performance evaluation. In CM, model perfo…
We examine semiparametric solutions to contamination bias for nonbinary treatments. Deepening the discussion by Goldsmith-Pinkham et al. (2024), we detail how spline functions approximate conditional expectation and propensity score functions under weak functional-form assumptions. Reanalyzing 18 regressions across 11 studies, we compare standard linear regressions against parametric and semiparametric versions of three contamination-robust estimators. W…
Structural similarity, the extent to which economic mechanisms respond alike to common shocks, matters for economic forecasting, policy transfer, and other decisions that rely on evidence from comparable settings. For mechanisms approximated by linear models, this idea has a simple geometric representation: The smaller the angle between their coefficient vectors, the more similarly they respond to the same shock. This paper introduces the cosine of this…
This paper develops inference for a Gaussian-nested hypergeometric family of distribution functions. The family \[ G_c(z) = \frac12 + z\,\frac{Γ(c-1/2)}{2\sqrt2\,Γ(c)}\,{}_1F_1\!\left(\frac12;c;-\frac{z^2}{2}\right),\quad c\ge\frac32, \] contains the standard normal distribution at the boundary $c=3/2$. Away from the boundary, the density has algebraic tail behaviour $g_c(z)\sim(c-3/2)|z|^{-3}$, so the parameter $c$ indexes a directed heavy-tailed deform…
Economic networks are often estimated separately over two periods, and changes in their edge sets are interpreted as structural rewiring. Since both networks are estimated, observed turnover also reflects graph-selection error. We study the two-snapshot Hamming-turnover functional under a homogeneous edge-misclassification model. With known sensitivity and specificity and conditional independence of the estimated edge indicators across periods, latent tu…
Local asymptotic minimax (LAM) risk is a foundational efficiency criterion in statistics and econometrics. The literature uses two definitions of LAM risk: one which appears in classical lower bounds and another which appears in arguments establishing attainment of those bounds. Conventional efficiency arguments are consistent with any estimator-dependent weighted average of the two, and consequently do not reveal which of these generalized $α$-LAM risk…
Multivariate longitudinal data may exhibit non-Gaussian margins, nonlinear dynamics, and response vectors with composition that varies across waves. To account for these features, we introduce a vector drawable vine (VD-vine) copula that extends conventional drawable vine copulas from scalar to vector-valued nodes. Here, the response vector at each wave forms a multivariate marginal, and serial dependence is captured through a sequence of linking vector…
In evidence-based policymaking, typically one experimental sample is observed, then a learned policy recommendation is implemented at scale. Policies learned from the experimental data can perform well in expected welfare, yet random sampling in the experiment can produce recommendations with poor welfare outcomes. In this paper, we ask: how should policy learning algorithms balance expected welfare against sampling risk? Our main contribution is to show…
Consider an increasing number of consistent estimators to be averaged when only estimated weights are available. The underlying parameter of interest can be identical across estimators (homogeneity) or not (heterogeneity). The contribution of the paper is threefold. First, it is shown that the interaction of the estimated weights with the estimators can generate specific bias terms. This constrains the number of estimators that can be aggregated when wei…
Juejue Wang, Pedro H. C. Sant'Anna, Victor Chernozhukov, Carlos Cinelli
16 Sep 2026 · Statistics — Methodology
We study the omitted variable bias (OVB) problem in canonical difference-in-differences (DiD) designs when unobserved confounding induces departures from the parallel trends assumption. Our results provide a novel characterization of the OVB formula for the average treatment effect on the treated (ATT), which is of independent interest. We show how the ATT bias is mainly governed by the strength of confounding in the treatment assignment mechanism and pr…
In the canonical Difference-in-Differences design, the control group's post-treatment change serves as an imputation of the treated group's counterfactual change in the same period, an imputation justified by parallel trends. However, differences in group composition can produce between-group differences in how outcomes would evolve over time, rendering this imputation vulnerable to confounding. An alternative imputation -- such as one based on the treat…
Discovering interpretable subgroups whose complier effects deviate from the average is a central goal of instrumental variable analysis under imperfect compliance, yet existing tree-based methods degrade when most covariates are irrelevant to the effect. We propose Shrinkage Bayesian Causal Forest with Instrumental Variable (SBCF-IV) for discovering and estimating subgroups with heterogeneous Complier Average Causal Effects (CACE) in sparse high-dimensio…
We develop a class of linear state space models for matrix-valued time series data where the state is a latent matrix normal process. We derive matrix versions of the Kalman filter, log-likelihood, and smoother enabling estimation of the latent state matrix as well as the model's parameters. To conduct Bayesian inference, we provide algorithms that draw from the joint posterior distribution of the latent state matrices conditional on the observed data an…
This paper develops a framework for individualized treatment allocation when interventions shift equilibrium prices and generate spillovers across treated and untreated units. The planner chooses which units receive a subsidy while allowing equilibrium prices to adjust endogenously. We show that the resulting welfare function is supermodular under broad and interpretable conditions, implying complementarity across treatment assignments and enabling exact…
Applied instrumental variables (IV) practice reports a first-stage F, now often the conditional F of Sanderson and Windmeijer (2016), and reads a large value as license to interpret the second stage. We show that no first-stage diagnostic can provide it. With a scalar instrument, a scalar treatment, and covariates entered linearly, the 2SLS estimand splits into a signal that a saturated specification would target and a contamination, the covariance betwe…
Modern economic and financial data are increasingly organized as multiway arrays, with observations indexed simultaneously by geographic regions, industrial sectors, asset categories, and other economic characteristics. Representing such data as tensor-valued time series preserves their intrinsic multiway structure. Although substantial effort has been devoted to modeling the conditional mean of tensor-valued time series, comparatively less attention has…
Instrumental variable analyses often rely on the assumption that instruments affect the outcome only through the endogenous regressor. In many applications, researchers can defend only a plausible range for direct effects of instruments, while conventional sensitivity analyses may be unreliable when instruments are weak. This paper proposes the profiled Anderson--Rubin (pAR) test, which considers all direct effects within a prespecified range and retains…
Exploiting the equality restrictions that nested Markov models encode requires their tangent-space geometry. For strictly positive finite-state models on arbitrary acyclic directed mixed graphs (ADMGs), we differentiate the recursive-head chart and prove that the range of its score map is the full tangent space. Intrinsic-set coordinate blocks form an algebraic direct sum, blocks of distinct districts are orthogonal, and the resulting Gram projection nee…
Policy learning aims to determine who should be treated based on individual characteristics. In high-stakes settings such as medicine and public policy where safety is a central concern, improving the average outcomes alone may not be sufficient: decision makers may also seek to protect individuals from harm, in line with the Hippocratic principle of “do no harm.” In this paper, we propose conformal policy learning (CPL), a policy learning procedure with…
I study dynamic treatment effects in panel data under staggered adoption when treatment timing depends jointly on unobserved time-invariant heterogeneity and time-varying pretreatment covariates, including lagged outcomes. Untreated potential outcomes follow a nonparametric dynamic panel model that allows flexible interactions between time-varying covariates and latent heterogeneity. I use pretreatment outcome histories to find individuals with similar t…
We investigate Empirical Bayes (EB) methods in the context of compound adaptive experiments, where the arm distribution in each experiment follows a normal distribution with an unknown mean that we seek to estimate. There are two main EB strategies: $g$-modeling, which estimates the prior by maximizing the marginal likelihood, and $f$-modeling, which derives posterior means directly from the empirical distribution of the observations. We show that $g$-mo…
This paper develops Wald inference for least-squares estimation of linear regression models on dyadic data, accommodating configurations where multiple observations share the same pair of units (e.g., directed flows, multilayer networks, and dyadic panels). We establish that the dyadic-robust Wald statistic is asymptotically $χ^2_q$ for an arbitrary nonrandom sequence of full-rank restrictions, under a single condition on the accumulation of dependence.…
We study exact-cardinality, equally weighted minimum-variance portfolio selection under a one-factor covariance model supplied in factor form. In the nonnegative homoskedastic regime, selecting the K smallest loadings is optimal. Allowing strictly positive asset-specific idiosyncratic variances makes the decision problem NP-complete even with positive integer loadings and a strictly positive-definite covariance matrix; with identity residual covariance,…
Model robustness analysis estimates an effect across a multiverse of specifications that pools control sets identifying the declared estimand with sets that condition on mediators or colliders. We propose stating rival assumptions about contested controls as a small set of candidate causal graphs, enumerating the adjustment sets each graph licenses, and reporting robustness metrics conditional on each graph. A finite-mixture identity splits the licensed…
Anthropogenic forcing components follow different long-run paths, while persistent temperature change can involve distributional changes beyond the mean. Scalar regressions aggregate these components and retain only mean temperature, obscuring how distinct forcing paths relate to persistent distributional change. We develop new testing, estimation, and inference methods for long-run relations between an integrated predictor vector and a density-valued re…
Likelihood-based network models are often fitted under links' independence and low-order constraints, while empirical networks frequently exhibit systematic higher-order structures such as triangles and wedges, characterizing the observed clustering patterns. Real-world core-periphery networks such as the interbank market or the air transportation system represent key examples, with cores displaying complex and nonlinear features. We formalize Penalized…
The random-coefficient demand model of Berry, Levinsohn, and Pakes (1995) is commonly estimated by the generalized method of moments (GMM), using an unconditional moment restriction with a fixed set of instruments. Identification of the model, however, rests on a conditional moment restriction. The two are not equivalent: the unconditional restriction may admit additional parameter values. We construct a counterexample in which the model is identified by…
The Money Pump Index (MPI) of Echenique et al. (2011) measures the severity of consumer irrationality, but computing the exact mean and median MPI over all revealed preference cycles is NP-hard (Smeulders et al., 2013). Existing solutions rely on heuristic proxies, such as evaluating only shorter cycles or bounding the MPI. By framing revealed preferences as a directed graph, this paper projects choice violations onto fundamental cycle bases, which are m…
Predict-then-optimize methods such as Smart "Predict, then Optimize" (SPO+) of Elmachtoub and Grigas (2022) learn a mapping from contextual features to unknown edge costs and then solve the induced combinatorial problem on the predicted costs. This approach is powerful but relies on the predictive model being well specified: when the true cost-generating process is nonlinear in the features and the predictor is linear, SPO+'s performance degrades as the…
Two frequent approaches for identifying structural VARs are external instruments, which carry economic content but are often weak, and non-Gaussianity of the shocks which provides statistical identification but carries no economic meaning. We combine the two strategies in a single generalized method of moments framework that stacks proxy exclusion restrictions with higher-order moment conditions of the structural shocks. This hybrid approach point-identi…
Observed choice in a dynamic game mixes current profit with continuation value. A rival adds a second problem: the same comparison averages over the rival's equilibrium policy. Changing the primitive transition rewrites continuation technology; changing the rival's Markov policy, holding that law fixed, rewrites the mixture over rival-contingent payoffs. The two are not substitutes. For a rival-feature payoff of rank $K$, rank identification up to locati…
In social and economic networks linked agents often share connections in common. There are two competing explanations for this phenomenon. First, agents may have a structural taste for transitive links - the returns to linking may be higher if two agents share a common connection. Second, agents may assortatively match on unobserved attributes, a process called homophily. We study parameter identifiability in a simple model of dynamic network formation w…
A switchback experiment alternates an entire system between treatment and control over time. It is especially useful when interactions between units can undermine standard unit-level experiments. Switchback experiments are commonly implemented on fixed temporal grids, which impose highly structured restrictions on when treatment can switch. We study a broader class of random-duration switchbacks, in which treatment and control alternate across runs whose…
We develop a causal framework for matrix completion under missing not at random (MNAR) data. Drawing on synthetic controls from the econometric panel data literature, our approach relaxes two assumptions common in MNAR matrix completion: positivity and independence of observation indicators. Unlike traditional panel data models, which often require prescribed block-sparse geometries, our framework accommodates flexible, heterogeneous observation patterns…
This paper studies identification and inference in a triangular system with an endogenous regressor when exclusion restrictions are unavailable and the dependence between structural disturbances is modeled through an unrestricted control function. In this setting, standard orthogonality conditions do not deliver point identification, as the unknown control function can rationalize a wide range of structural coefficients. We show that identifying informat…
We derive the exact likelihood of the Berry (1994) share-form nested logit and show it contains a Jacobian term that depends on nest size and the nesting parameter. Omitting the term biases within-nest substitution estimates wherever choice sets vary across markets. Commonly-used count instruments for the within-nest share fail exclusion when product counts enter demand directly. The corrected likelihood identifies substitution without them, making crowd…
Probabilistic forecasting is central to decision-making under uncertainty, yet its methodological landscape has become increasingly fragmented across temporal and spatiotemporal forecasting, statistical modeling, machine learning, and deep generative modeling. This survey develops a unified perspective by organizing probabilistic forecasting methods according to where and how uncertainty is introduced into the forecasting pipeline. Our taxonomy connects…
This paper investigates the construction of moment restrictions in dyadic network formation models with unobserved individual heterogeneity under nontransferable utility. Using observed links and covariates from five-node pentads, we construct moment restrictions that do not depend on individual fixed effects. For a broad class of covariate specifications, the construction is minimal in the sense that it uses the least possible number of nodes and dyads.…
Including covariates alongside strictly monotone encodings can change a causal forest's treatment decisions without adding information. Random feature selection favors covariates represented by multiple columns. Under stated conditions, I show that this imbalance can persist as samples grow. Treatment effect components associated with other covariates are omitted, attenuated, or recovered depending on their inclusion probabilities and tree depth. Simulat…
Identification in vector autoregressions involves two distinct choices: how extensively to orthogonalize shocks and whether orthogonality is imposed sequentially or simultaneously. Generalized identification imposes no orthogonality; Sims (1980} imposes full orthogonality sequentially; Francis et al. (2026) impose full orthogonality simultaneously; and Buchwalter et al. (2026a) provide clustered partial orthogonalization sequentially. We fill the remaini…