Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
80,226 characters · 15 sections · 36 citation commands
Causal Identification in Multi-Task Demand Learning with Confounding
Large retail firms routinely make pricing decisions across a vast collection of selling contexts. These contexts can take many forms: they could correspond to a single product sold by a brand across dispersed stores and channels within an omnichannel network, or to distinct products within a single catalog of an e-commerce retailer. In all such cases, each decision environment is characterized by heterogeneous demand determinants, including customer demographics, competitive conditions, local preferences, seasonality, and product-specific attributes. As a result, price sensitivities---and hence optimal prices---vary substantially across decision contexts. Effective pricing policies therefore require reliable estimates of context-specific price-response functions.
Estimating such heterogeneous demand responses poses a fundamental statistical challenge. At the level of an individual store or product, historical price variation is typically limited: operational frictions, coordinated pricing policies, and managerial practices often result in only a small number of distinct prices receiving non-negligible exposure over long periods of time. In contrast, the cross-sectional dimension is large, with many stores or products observed and rich covariates available. This data structure naturally suggests a multi-task learning or partial pooling approach caruana1997multitask, baxter2000model, gelman2013bayesian, in which a shared model maps observable characteristics to task-specific demand parameters, borrowing strength across tasks to overcome limited within-task variation.
However, this approach encounters a central obstacle that fundamentally limits its validity in observational pricing data: confounding in the price-setting process. Retail prices are not randomly assigned; they are chosen endogenously in response to managerial beliefs, anticipated demand fluctuations, promotional calendars, and other local or product-specific shocks. Many of these factors are unobserved or only partially observed by the econometrician. As a result, prices may be systematically correlated with unobserved determinants of demand even after conditioning on observed covariates. This endogeneity is pervasive in practice and undermines the causal interpretation of price coefficients obtained from standard regression or learning-based methods angrist2009mostly, wooldridge2010econometric. We establish that in a multi-task demand learning setting with arbitrarily confounded prices, commonly used approaches---including pooled regression and modern meta-learning techniques---fail to identify causal task-specific price effects. Even as the number of tasks and the number of observations per task grow, these methods converge to estimands that reflect the endogenous pricing policy rather than the underlying causal price-response function. In this sense, policy-induced confounding creates a fundamental identification failure that cannot be remedied by additional data alone.
We then ask whether improved causal identification through cross-task transfer learning is possible in this environment, and if so, under what conditions. We study a setting in which each store/product has a (locally) linear price-response curve with a context-dependent intercept and slope. The causal objects of interest are the context-specific price coefficients and intercepts. The econometric challenge is that prices are chosen endogenously using both observed covariates and latent, task-level information, so standard pooling or meta-learning targets policy-confounded objects.
Our starting insight is that, because managers and pricing algorithms set prices using local information, the realized price history is itself informative about latent demand determinants. Causal transfer learning in observational pricing data therefore must treat the price path as part of the conditioning information, rather than as a regressor whose variation can be taken as quasi-experimental. But conditioning on the full decision history creates a second, less obvious obstacle: if the learner can deterministically infer which decision is used for supervision, the training signal depends on the underlying demand parameters only through a one-dimensional projection, leaving many observationally equivalent solutions and destroying identification.
We resolve this tension with a simple information-design principle: condition on the endogenous price path to absorb confounding from latent fundamentals, while obfuscating just enough outcome information so that the supervised decision cannot be uniquely pinned down from the inputs. This yields {\it Decision-Conditioned Masked-Outcome Meta-Learning} (DCMOML), which masks outcomes at two candidate query points and randomizes which one is used for training. Under a mild exogeneity condition on each task that restricts the adaptivity of the final price point to the demand at the penultimate price point, this design consistently recovers the conditional expectation of the causal parameters given the information set. Overall, the paper makes three main contributions.
More broadly, the paper bridges empirical demand estimation under endogenous prices and cross-task learning in many small-data environments. It shows that transfer learning does not, by itself, resolve endogeneity, and it offers a simple identification-based design that makes causal estimation feasible at a large scale when randomized tests are unavailable or only partial. Identifying these context-specific demand primitives can enable principled downstream pricing and promotions decisions in retail operations.
Our work sits at the intersection of (i) empirical demand estimation under endogenous pricing, (ii) partial pooling and multi-task/meta-learning for many small-data environments, and (iii) machine learning methods for causal inference.
\paragraph{Demand estimation and price endogeneity in empirical IO and marketing.} A large empirical IO and quantitative marketing literature emphasizes that prices are typically endogenous, so recovering causal demand elasticities requires explicit identification. Canonical approaches use instrumental variables or control-function corrections to break the correlation between price and unobserved demand shocks. In differentiated products, the random-coefficients logit framework of berry1994estimating,berry1995automobile operationalizes this idea using instruments based on cost shifters and rivals’ characteristics; nevo2001measuring provides a prominent application highlighting how endogeneity correction changes inferred elasticities and market power. Related IO and marketing work similarly stresses that omitted quality and demand shocks can bias price coefficients absent valid instruments (e.g., hausman1996valuation).
A complementary strand replaces linear IV with control functions (or proxy residualization) embedded in nonlinear demand systems. For discrete choice models, petrin2010control develop a control-function approach compatible with flexible first-stage price models, and train2009discrete surveys the broader discrete-choice toolkit. In marketing, hierarchical Bayes and random-effects models are widely used to capture heterogeneity across products or customers, but they do not, by themselves, resolve price endogeneity: identification still hinges on instruments, exclusion restrictions, or structural assumptions about the pricing process (e.g., rossi2005bayesian).
Our setting is distinct from this classical emphasis along two axes. First, we study large-scale multi-task demand learning—many stores/products with rich covariates but very few distinct price exposures per task—where task-by-task IV, panel fixed effects, or flexible first-stage control-function estimation is often statistically or operationally infeasible. Second, rather than positing specific cost shifters, exclusion restrictions, or a parametric pricing rule, we allow the logged pricing policy to embed arbitrary latent task-level information and instead pursue identification through the information revealed to the learner: conditioning on the endogenous decision history while designing supervision so the target causal effect remains identified without instruments.
\paragraph{Learning and optimization in data-scarce pricing environments.} Operational pricing and revenue management often operate under limited data. A complementary line of work to ours studies optimization with extremely limited historical data and develops robust decision rules allouah2022pricing, allouah2023optimal. These results connect closely to robust optimization ben2009robust and to sample-complexity analyses that quantify how many observations are needed to guarantee a fixed fraction of the full-information objective cole2014sample, huang2015making.
Our setting shares the same per-task data-scarcity motivation, but addresses it via cross-task transfer: we exploit many related tasks, each with only a few effective price points, to learn shared structure and adapt task-specific demand parameters. This is fundamentally different from single-environment distributional robustness, and it is also where endogeneity becomes a critical bottleneck.
\paragraph{Partial pooling, Empirical Bayes, and hierarchical modeling.} When each store/product has too few observations to estimate its own demand curve reliably, a standard statistical remedy is to partially pool information across units. In statistics and econometrics this is formalized via hierarchical and random-effects models gelman2013bayesian,laird1982random,pinheiro2000mixed and their Empirical Bayes (EB) counterparts, which estimate a population prior from the data and shrink task-level coefficients toward it robbins1956empirical,efron1973stein,carlin2000bayes,efron2012large. In demand estimation, these ideas are pervasive: hierarchical Bayesian models stabilize store- or customer-level parameters while allowing systematic variation with covariates rossi2005bayesian. This literature directly targets the small-data regime by trading variance for bias through shrinkage.
The success of these approaches hinges on obtaining unbiased causal estimates at the task level. Recent work studies how to conduct valid inference with adaptive data collection and pricing, including inference for adaptive linear models deshpande2018accurate, hadad2021confidence,wang2021uncertainty. Our work differs in two key respects. First, DCMOML targets unobserved confounding from decision dependence on latent task fundamentals, not just adaptivity to observables. Second, even when unbiased task-level estimation is feasible, classical EB/random-effects methods typically use realized prices only as regressors in the outcome likelihood and do not extract information encoded in the price path. In contrast, DCMOML conditions on the full realized price sequence and learns across tasks, extracting signal from pricing decisions without specifying a model of how prices are set, improving predictive performance.
\paragraph{Multi-task learning and meta-learning under endogenous decisions.} In machine learning, cross-task transfer appears as multi-task learning (MTL), which learns shared representations or inductive biases to reduce sample complexity and improve generalization caruana1997multitask,baxter2000model,evgeniou2004regularized. Meta-learning extends this logic to few-shot adaptation: it learns an adaptation rule that maps a small within-task dataset to task-specific parameters bengio1990learning,finn2017maml, hospedales2021metalearning. The renewed prominence of meta-learning is also connected to the success of large foundation models and in-context learning, which can be viewed through a meta-learning lens radford2019language.
Despite strong predictive performance, standard MTL and meta-learning formulations typically do not confront endogeneity. Our approach is motivated by the fact that, under endogenous pricing, the decision history itself carries information about latent task fundamentals. DCMOML conditions on the entire realized price path and learns across tasks, using how managers move prices as an auxiliary signal about task heterogeneity, while its outcome-obfuscation/randomization design preserves causal identification when identification is feasible (e.g., with two distinct prices per task). Thus, we can view DCMOML as a decision-conditioned, causally-identified transfer-learning analogue to standard MTL/meta-learning objectives for demand prediction.
\paragraph{Causal machine learning and heterogeneous treatment effects.} Our model resembles heterogeneous treatment effect (HTE) estimation in that the price coefficient varies across tasks and depends on covariates. Under unconfoundedness (ignorability), given observed features, a large literature develops flexible estimators of conditional average treatment effects, including causal trees and forests athey2016recursive, wager2018estimation and orthogonal / double-robust methods such as Double Machine Learning (DML) chernozhukov2018double and the R-learner nie2021quasi. These approaches rely on the key condition that, after conditioning on observables, treatment assignment is independent of potential outcomes (or at least that the relevant orthogonality moments hold). Textbook treatments and modern syntheses appear in imbens2015causal and chernozhukov2024applied.
In contrast, our setting permits unobserved confounding: prices may depend on latent store/product-specific information that also affects demand. This violates the orthogonality conditions that underpin DML/R-learner-style residualization. While classical econometrics offers remedies via instrumental variables, panel fixed effects, and difference-in-differences angrist2009mostly, wooldridge2010econometric, these tools are often difficult to deploy at the granularity of many stores/products with only two or three effective price points. Recent causal inference work studies identification under hidden confounding using proxy variables and proximal identification tchetgen2020proximal, or uses instruments within flexible ML pipelines (e.g., DeepIV; hartford2017deepiv). Our estimator takes a different route: rather than requiring instruments or proxies, we exploit the multi-task structure with repeated decisions and design the learner's information set so that the realized decision history can be conditioned on without destroying identifying variation.
\paragraph{Tasks, Data, and Notation}We consider a multi-task demand learning environment with $N$ tasks, indexed by $i = 1,\dots,N$. For concreteness, one can imagine that each task corresponds to a distinct store selling a fixed product. Tasks are heterogeneous and are characterized by observable contextual covariates $Z_i \in \mathbb{R}^d$, which may include demographics, competition information, geographic attributes, and other store-level features. We assume that $\{Z_i\}_{i=1}^N$ are drawn independently from an unknown distribution $\mathcal{P}_Z$. For each task $i$, we observe $K$ historical price--demand pairs \[ \{(p_{ik}, D_{ik})\}_{k=1}^K, \] where $p_{ik}\in\mathbb{R}$ denotes the price charged and $D_{ik}\in\mathbb{R}$ denotes realized demand. Here, $k$ represents a sequential pricing path---e.g., $p_{ik}$ is the price posted on day $k$ over a window of length $K$ and $D_{ik}$ is the corresponding demand. Prices may repeat across the $K$ observations. We focus on regimes with limited within-task price variation, where each task exhibits only a small number of distinct price levels (e.g., two or three). Throughout, we assume that at least two distinct prices are observed for every task.
\paragraph{Structural Demand Model} We assume that demand is linear in price within each task. Specifically, for each task $i$ and observation $k$,
where $\Theta_i \triangleq (\theta_i^0, \theta_i^1)^\top \in \mathbb{R}^2$ denotes the task-specific demand parameters, and $\epsilon_{ik}$ is an idiosyncratic demand shock. Define the regressor vector $P_{ik} \triangleq (1,\, p_{ik})^\top,$ so that (ref) can be written compactly as
We assume that the noise terms $\epsilon_{ik}$ are mean zero, have finite variance, and are independent across $i$ and $k$.
\paragraph{Heterogeneity and Shared Structure Across Tasks} Task-specific demand parameters vary systematically with observed context and idiosyncratically through unobserved factors. We formalize this by decomposing $\Theta_i$ as
where:
We assume that $\mathbb{E}[\Omega_i \mid Z_i] = 0$ and that $\{\Omega_i\}_{i=1}^N$ are i.i.d.\ across tasks with finite second moments. The function $g(\cdot)$ represents the systematic component of heterogeneity that can, in principle, be learned from data by pooling information across tasks, while $\Omega_i$ captures persistent store-specific shocks that are not explained by observed covariates. Substituting (ref) into the demand equation yields
The unobserved components of the model are therefore the task-level random effect $\Omega_i$ and the observation-level noise $\epsilon_{ik}$. We assume that the collections $\{Z_i, \Omega_i\}_i$ and $\{\epsilon_{ik}\}_{i,k}$ are mutually independent.
\paragraph{Price Assignment and Confounding} A central feature of our setting is that prices are endogenously assigned. At the outset, we impose no structure on the price-setting process beyond basic measurability and causality. In particular, we require that the period-$k$ price $p_{ik}$ is chosen as a measurable function of the information available up to period $k-1$, namely \[ (Z_i,\; p_{i1},D_{i1},\ldots,p_{i,k-1},D_{i,k-1},\; \Omega_i). \] Crucially, we do not assume that prices are conditionally exogenous: even after conditioning on the observable history $(Z_i,p_{i1},D_{i1},\ldots,p_{i,k-1},D_{i,k-1})$, the price $p_{ik}$ may remain statistically dependent on the unobserved demand component $\Omega_i$. In other words, prices may reflect managerial judgment or local information unavailable to the econometrician, so endogeneity persists despite conditioning on the context and the history of observables. We will, however, later impose a mild restriction on the adaptivity of the final price $p_{iK}$, formalized in Assumption (ref).
\paragraph{Causal Estimand and Objective} Our primary object of interest is the causal task-specific demand parameter vector $\Theta_i$, which characterizes the demand response to counterfactual price interventions at store $i$. Importantly, $\Theta_i$ is defined purely by market fundamentals, independent of the historical pricing policy. This distinguishes it from standard regression estimands, which capture endogenous correlations between prices and unobservables.
We note that because $\Theta_i$ includes a random task-specific deviation $\Omega_i$, it cannot be perfectly recovered from finite task-level data, even as the number of tasks grows. Accordingly, our objective is to consistently estimate the {\it conditional mean} of $\Theta_i$ given an admissible subset of the observable history in task $i$. Through transfer learning, we aim to recover this predictive causal functional, effectively denoising the task parameters while maintaining causal validity.
A natural starting point for estimating the task-specific parameter $\Theta_i$ is to fit an ordinary least squares (OLS) regression separately for each task using its $K$ datapoints $\{(P_{ik},D_{ik})\}_{k=1}^K$. This corresponds to the empirical risk minimizer (ERM)
There are two challenges here. First, when prices are adaptively chosen, the regression estimates are generally biased lai1982least.\footnote{Although debiasing techniques exist (see deshpande2018accurate, hadad2021confidence,wang2021uncertainty), these techniques trade off bias for variance and cannot reliably recover $\Theta_i$ with small per-task data.} Second, even under non-adaptive pricing, OLS estimates can be noisy and unstable with limited per-task data. This motivates approaches that leverage shared structure across tasks.
At the opposite extreme from task-specific estimation, one may attempt to ignore task heterogeneity altogether and learn only the shared mapping $g(\cdot)$. This corresponds to a pure representation-learning approach in which store-level variation is explained entirely by observable context, and unobserved task-level deviations are treated as noise.
A natural baseline in this spirit is to posit a parameterized function class $\{g_\Lambda(\cdot)\}_{\Lambda \in \mathcal{H}}$ (e.g., linear models, kernel methods, or deep neural networks), and to estimate $\Lambda$ by pooling all observations across tasks:
This estimator is appealing in high-dimensional settings because it avoids within-task estimation altogether and relies purely on cross-task variation. However, under the model introduced in Section (ref), this approach is generically inconsistent. To see this, note that substituting the data-generating process yields the population regression residual \[ D_{ik} - P_{ik}^\top g(Z_i) = P_{ik}^\top \Omega_i + \epsilon_{ik}. \] Because prices are endogenously assigned, the regressor $p_{ik}$ may be correlated with the unobserved task-level deviation $\Omega_i$. Consequently, $\mathbb{E}\!\left[P_{ik}^\top \Omega_i \mid Z_i\right] \neq 0$ in general, violating the orthogonality condition required for consistency of (ref). This is a classical endogeneity problem: prices embed latent managerial or local information through $\Omega_i$, and learning $g(\cdot)$ via pooled regression therefore recovers a policy-dependent object rather than the causal mapping from context to demand parameters.
Meta-learning provides a natural middle ground between the two extremes discussed above hospedales2021metalearning, finn2017maml: estimating demand separately for each store using only within-store data, and ignoring store-specific deviations by learning only the shared mapping $g(\cdot)$. The central idea is to use cross-store regularities to guide how limited within-store data should be interpreted. When little or no store-level data are available, the method effectively defaults to a pooled estimate; as more store-specific observations accumulate, the estimate adapts toward a store-specific demand curve. In this sense, meta-learning smoothly interpolates between shared-model learning and per-task estimation.
Operationally, meta-learning treats the estimation of each store’s demand parameters as a conditional prediction problem: given the store’s context $Z_i$ and a small set of observed price–demand pairs, the goal is to predict the underlying parameter vector $\Theta_i$. This mapping from $(Z_i,\text{data})$ to $\Theta_i$ is learned using data from many stores, allowing the model to discover how demand parameters typically vary with the context and early price--demand signals.
Formally, we consider a parameterized adaptation map $g_\Lambda : (Z_i,\mathcal{S}_i) \mapsto \widehat{\Theta}_i,$ where the support set $\mathcal{S}_i$ consists of the first $K-1$ observations for store $i$, $\mathcal{S}_i = \{(P_{i1}, D_{i1}), \dots, (P_{i,K-1}, D_{i,K-1})\}.$ The remaining observation $(P_{iK}, D_{iK})$ is treated as a query point and is used to train the model via the empirical risk minimization problem\footnote{It is inevitable that the support set cannot be the entire observation history for a task, as some data must be given up to provide a supervisory signal for learning the adaptation map.}
With sufficient expressive power and data, this procedure could be naively anticipated to converge to the Bayes predictor $g_{\Lambda}(Z_i, \mathcal{S}_i) \;\approx\; \mathbb{E}[\Theta_i \mid Z_i, \mathcal{S}_i], $ thereby optimally combining cross-store information and limited within-store observations. However, despite this flexibility, this convergence doesn't generally hold true under our model, and meta-learning remains inconsistent for causal demand estimation. The fundamental issue is that conditioning on $\mathcal{S}_i$ is not sufficient to remove confounding. Indeed, the population residual associated with (ref) under the target $\mathbb{E}[\Theta_i \mid Z_i, \mathcal{S}_i]$ satisfies
which is non-zero since, even after conditioning on all observed data up to time $K-1$ within the task, the query price $p_{iK}$ remains correlated with the unobserved slope component $\omega_i^1$. Thus, although meta-learning adapts estimates toward store-specific parameters, the adaptation is driven by endogenous signals. The resulting estimator converges to a policy-dependent object rather than the conditional mean causal demand curve.
In short, per-task OLS fails due to insufficient data, while both shared-model and meta-learning approaches fail due to endogeneity.
This section develops the key conceptual and technical ideas underlying our approach.
The preceding sections established that conditioning on limited task-level information— such as $(Z_i,\mathcal{S}_i)$ in standard meta-learning—cannot in general eliminate confounding induced by endogenous pricing. We therefore turn to a natural next question: whether expanding the conditioning set to include the full price history of a task is sufficient to restore causal identification.
In particular, the realized price vector $(p_{i1},\dots,p_{iK})$ contains information about the latent task parameter $\Theta_i$ through the (unknown) pricing policy, beyond what is revealed by the input--output pairs $\{(p_{ik},D_{ik})\}_{k=1}^{K-1}$ viewed as a regression problem. This motivates targeting the richer conditional object
which explicitly conditions on {\it all} endogenous price realizations; in particular, including the query price point $p_{iK}$. One might hope that passing the entire price history into a sufficiently expressive learner would control for confounding and resolve the endogeneity problem. To formalize this idea, consider the following idealized meta-learning objective. Using one observation as a query point and the remaining observations as covariates, define
where $P_{iK}=(1,p_{iK})^\top$. The essential difference relative to the standard meta-learning objective is that the entire price vector $(p_{i1},\dots,p_{iK})$ is included among the covariates (as opposed to only $(p_{i1},\dots,p_{iK-1})$).
While this modification addresses confounding at the level of conditional means, it results in a new identification challenge. The issue arises because the meta-learner has explicit access to the query price point $p_{iK}$. The supervision signal for task $i$ enters only through the scalar inner product $P_{iK}^\top \Theta_i$. Consequently, the empirical risk in (ref) is invariant to shifts of the predicted parameter vector along directions orthogonal to $P_{iK}$. Formally, for any candidate prediction $\widehat{\Theta}_i$ and any measurable scalar function $\phi$ of the inputs, define \[ \widetilde{\Theta}_i \;\triangleq\; \widehat{\Theta}_i + \phi(inputs_i)
. \] Since $P_{iK}^\top [\,p_{iK},-1\,]^\top = 0$, both $\widehat{\Theta}_i$ and $\widetilde{\Theta}_i$ achieve exactly the same objective value in (ref). When the function class $g_\Lambda$ is sufficiently expressive and because it has explicit access to $p_{iK}$, the learning problem admits many such observationally equivalent solutions, and the causal parameter $\Theta_i$ is not identified. This highlights an important point: conditioning on the full decision history is necessary to address confounding, but empirical risk minimization alone is insufficient to pin down the underlying structural parameters.
We now describe a simple information design that restores identification. The objective is to define a conditioning set that is rich enough to address policy-induced confounding, while avoiding the degeneracy that arises when the learner can pin down the query regressor.
\paragraph{Step 1: Randomizing the query index to avoid a uniquely identifiable query regressor.} A first issue with standard meta-learning is that the query index is fixed by design (e.g., $k=K$), so the learner knows a priori which price point will be used in the loss. In our setting, this is problematic because it makes the query regressor perfectly identifiable from the inputs. To prevent this, we randomize which observation is designated as the query.
To reduce the variance introduced by this randomization, we randomize the query index only within a fixed two-point subset of indices. Formally, for each task $i$, let $K_i^*$ be the largest index such that $p_{iK_i^*} \neq p_{iK}$, i.e., $p_{iK_i^*}$ is the last chosen price that was distinct from $p_{iK}$ and $K_i^*$ is the corresponding index. Under the assumption that there are at least two distinct prices, $K_i^*$ is well defined. Then draw \[ k_i^1 \sim \mathrm{Unif}\{K_i^*,\,K\}, \] independently across tasks and independently of all other variables. We designate $(p_{ik_i^1},D_{ik_i^1})$ as the query pair. At this stage, a natural analogue of standard meta-learning is to give the learner the full price vector and all demand outcomes except the query demand:
which suggests the conditional target $\mathbb{E}[\Theta_i\mid X_i^{(-k_i^1)}]$.
\paragraph{Step 2: Dropping a second demand outcome to obfuscate the query regressor.} A little thought reveals that query randomization is not sufficient. Even though the learner does not know $k_i^1$ in advance, it can infer $k_i^1$ from the observed data in (ref) because the demand vector is indexed and therefore aligned with prices. Concretely, the learner observes $(p_{ij},D_{ij})$ for every $j\neq k_i^1$, and it observes no demand associated with exactly one price index. Hence, the missing index is identified, and so is the query price; in other words, $p_{ik_i^1}$ is measurable with respect to the sigma-field generated by $X_i^{(-k_i^1)}$. The same non-identification from Section (ref) therefore persists.
This argument shows that to prevent the query regressor from being uniquely identified from the covariates, it is not enough to withhold a single demand observation: doing so leaves exactly one unmatched price index. We therefore withhold two demand outcomes so that there are two unmatched prices and the learner cannot determine which one corresponds to the query.
In our two-point design, the second withheld index is simply the other element of $\{K_i^*,K\}$. Specifically, after drawing $k_i^1$, set \[ k_i^2 \;\triangleq\; \{K_i^*,K\}\setminus\{k_i^1\}. \] The learner observes all prices and all demands except those at indices $k_i^1$ and $k_i^2$, i.e., $K_i^*$ and $K$:
where $\mathbf{D}_i^{-(K_i^*,K)}$ collects $\{D_{ij}: j\notin\{K_i^*,K\}\}$ together with their indices.
Under (ref), both the prices $p_{iK_i^*}$ and $p_{iK}$ are unassociated with demand outcomes; thus, $p_{ik_i^1}$ is no longer measurable with respect to the sigma-field generated by $X_i^{-(K_i^*,K)}$, and the query regressor cannot be deterministically recovered from the covariates. This restores identifying variation that is destroyed when the query regressor is revealed, and motivates the conditional target
However, we need to ensure that we are not introducing a new form of confounding by withholding the additional demand outcome $D_{iK_i^*}$. Formally, we impose the following assumption within each task, requiring that after conditioning on the observable covariates (after demand obfuscation), the query-period demand shock remains mean zero.
Under Assumption (ref), we show in our main result that the target in (ref) is identifiable with a sufficiently expressive meta-learner function class.
\paragraph{Discussion.} Assumption (ref) is an exogeneity restriction on the final two distinct price-points. It is naturally satisfied whenever the final pricing decision---i.e., the terminal price and how long it remains in place---does not respond to the demand realization at the query index $K_i^*$, i.e., to the demand outcome at the penultimate distinct price point.
To see this, note that by definition of $K_i^*$, all prices after $K_i^*$ are identical: for every $k > K_i^*$, $p_{ik} = p_{iK}$. Under a causal pricing policy, mean-zero shocks imply that $\epsilon_{ik}$ is centered given the price history up to time $k$ and the demand history up to time $k-1$. Therefore, the conditional mean of $\epsilon_{iK}$ given the entire history of observables is always $0$. The problem is that the conditional mean of $\epsilon_{iK_i^*}$ given the {\it entire} price path can deviate from zero if the decision to set $p_{iK}$ and its duration depends on $D_{iK_i^*}$. In that case, $p_{iK}$ and its duration reveal information about the shock $\epsilon_{ik_i^1}$ at the query time in the event that $k_i^1 = K_i^*$. Consequently, Assumption (ref) holds as long as the terminal price $p_{iK}$ and its duration are chosen independently of $D_{iK_i^*}$. Importantly, these decisions can {\it still} depend on all of the rest of the historical observations and the latent parameter $\Omega_i$. We next discuss two common price-generating schemes under which this condition is plausible:
{(i) {\bf Non-adaptive confounded pricing.}} If the decision-maker chooses the price path $p_{i1:K}$ independent of past demand observations, although still potentially depending on the unobserved $\Omega_i$, then $(p_{i1:K},K_i^*)$ can be determined prior to the realization of the shocks $\{\epsilon_{ik}\}_{k=1}^K$. In particular, under the mean-zero noise condition, $$\mathbb{E}\!\left[\epsilon_{ik_i^1}\,\middle|\, Z_i,\ p_{i1:K},\ \mathbf{D}_i^{-(K_i^*,K)},\ p_{ik_i^1}\right] = \mathbb{E}\!\left[\epsilon_{ik_i^1}\,\middle|\, Z_i,\ p_{i1:K},\ p_{ik_i^1}\right] =0.$$ This is the case, for instance, in Examples (ref)--(ref) discussed earlier.
{(ii) {\bf Adaptive confounded pricing with successive two-price experimentation.}} Assumption (ref) also holds for adaptive pricing routines that successively experiment at two price levels as part of a gradient descent routine on the profit curve (e.g., two-point stochastic approximation; kiefer1952stochastic, salem2025algorithmic, spall1992multivariate,flaxman2005online). Such schemes follow a canonical online learning template in which a previous gradient-step update yields a new incumbent price $p_t$, and the policy then deploys two nearby probes (e.g., $p_t\pm\delta$) to estimate the gradient for the next update.
In this view, if the indices $K_i^*$ and $K$ correspond to such a two-point experimental block---so that the probe prices (and their duration) are chosen based on earlier history and latent information $\Omega_i$, but not on the realized demand within the block---then Assumption (ref) holds. More generally, if $(K_i^*,K)$ do not correspond to a probe block, one can instead restrict each task to its most recent two-point experiment: take the last two distinct price points that were part of a two-price probe, and ignore any subsequent data. Tasks then have heterogeneous effective lengths, which can be handled with standard devices (e.g., masking/padding in the meta-learner inputs or architectures that operate on variable-length sequences/sets).
We now present our main estimator.
We now present our main results on the identifiability of our estimand and the consistency of our estimator.
The proof, provided in Appendix (ref), leverages the two-outcome withholding design to show that the query regressor cannot be recovered from the covariates and therefore retains identifying variation conditional on the learner's information. Specifically, we show that this induces a positive definite conditional second-moment matrix $\mathbb{E}[P_{ik_i^1}P_{ik_i^1}^\top\mid X_i^{-(K_i^*,K)}]$, which rules out the orthogonal-shift non-identification present when the query regressor is revealed. Therefore, any candidate predictor that attains the population risk minimum must agree with $\mathbb{E}[\Theta_i\mid X_i^{-(K_i^*,K)}]$ almost surely.
This consistency result follows from standard uniform convergence arguments: under realizability and a Glivenko--Cantelli (uniform law of large numbers) condition for the induced squared-loss class, any empirical risk minimizer has population risk converging to the population minimum and therefore converges (in risk, and in $L^2$ under the eigenvalue condition) to the identified target. The proof is presented in Appendix (ref). The eigenvalue condition is a well-conditioning requirement that the squared price gap $(p_{K^*} - p_K)^2$ remains strictly positive and sufficiently large relative to the scale of the prices; see Appendix (ref).
We evaluate DCMOML and design alternatives on a more realistic variant of Example (ref). Each task $i$ is a store-specific linear demand curve with both slope and intercept varying across tasks. Prices are chosen endogenously by a manager who (i) forms a noisy estimate of the task’s revenue-optimal price and (ii) locally experiments around that estimate. We vary the noise in the manager’s optimal-price estimate to control the strength of confounding: when the estimate is accurate, prices are more tightly coupled to latent parameters; as this noise increases, prices contain more quasi-exogenous variation and confounding weakens. We consider three regimes: high, medium, and low confounding (HC, MC, LC, resp.). Throughout, within-task experimentation remains small. The precise definition of the setup is presented in the Appendix (ref). Note that in this environment, all $K$ prices per task and hence the last two price points, are almost surely distinct. As a result, $K_i^\ast=K-1$ (almost surely) for all tasks.
\paragraph{Methods.} The alternative methods are designed to test the design choices we made in DCMOML.
We focus on $K=2$ since it is the most stringent regime for outcome obfuscation, as no demand values are revealed to the learner under DCMOML. All transfer-learning methods use a simple feedforward neural network architecture (MLP with a hidden dimension of 128 and depth 4). We report MSE with standard error for recovering the task-level slope $\theta_i^1$ and intercept $\theta_i^0$ in Figure (ref).
\paragraph{Findings.} Three patterns emerge. First, under high confounding, DCMOML substantially outperforms all baselines on both slope and intercept estimation. In contrast, DCML performs quite poorly (omitted from the plots to improve readability; see Table (ref) in the Appendix) as conditioning directly on the query decision renders the causal estimand unidentifiable. TASK-OLS is also unstable in this regime due to limited within-task price variation (Table (ref) in the Appendix). As expected, EB-GLS significantly improves performance relative to TASK-OLS and SHARED; however, it is outperformed by META and DCMOML, suggesting the power of meta-learning techniques in extracting signal from decision choices that EB ignores.
Second, while DCMOML remains competitive throughout, its relative advantage is naturally largest in the regime it is designed for: highly endogenous pricing with minimal within-task experimentation. When confounding weakens, obfuscating outcomes discards useful signal, so the performance of DCMOML, META-LEARNING, EB-GLS, and SHARED moves closer.
Finally, DCUOML exhibits poor performance despite query randomization and information obfuscation, presumably since it leaks enough information for the learner to effectively infer the query decision, reintroducing the identifiability failure that plagues DCML.
We evaluate {DCMOML} on UK-online-retail, an online retail dataset containing transaction data spanning 01/12/2010--09/12/2011 for a UK-based gift retailer uci_online_retail. It contains $\sim 4{,}070$ products, each with an average of 3.78 distinct posted prices (median 4). Price exposure is highly concentrated: the modal price accounts for 65.48% of observed days on average, and the top two prices together account for 89.74%.
\paragraph{Tasks and holdout protocol.} We construct two product-level task definitions. In both, the context is $Z_i\in\mathbb{R}^{1024}$, the sentence-transformer embedding reimers2019sentencebert of the product title. Since ground-truth demand parameters $\Theta_i$ are unobserved, performance of demand estimation methods is evaluated by predictive accuracy on a held-out price point.
\paragraph{Practical considerations.} Both tasks induce heteroskedasticity because $D_{ik}$ averages over variable exposure lengths; we accordingly minimize exposure-weighted squared losses during training. In the temporal task, exposure lengths $e_{ik}$ are themselves decision variables and may encode latent demand conditions. DCMOML applies the same information-design principle: condition on the endogenous decision history (prices and exposures), but obfuscate outcomes so the learner cannot deterministically infer which decision is used for supervision.
We focus on $K=2$ for two reasons: (i) it is the most challenging regime for DCMOML, which reveals no demand outcomes to the meta-learner; and (ii) in the temporal task it prioritizes the most recent history, mitigating concerns about non-stationarity.
\paragraph{Methods.} We compare:
\paragraph{Training and evaluation.} For transfer-learning methods, we split products into 80% train and 20% validation sets and use the validation RMSE for early stopping (same criterion across methods). All methods use the same feedforward network (depth 2, hidden size 256). Losses are exposure-weighted and normalized within each product in every batch (see Appendix (ref) for details). We repeat the full pipeline over 100 random seeds; PER-TASK solves weighted least squares per product and requires no repeats. We report exposure-weighted RMSE on the held-out point: $\text{RMSE}=\sqrt{(\sum_i e_i \,(y_i-\hat y_i)^2)/(\sum_i e_i)},$ and 95% confidence intervals across seeds.
\paragraph{Findings.} The results are presented in Figure (ref). Across both task definitions, {DCMOML} achieves the lowest held-out RMSE, outperforming outcome-conditioned {META} and the pooled {SHARED} model. This pattern is consistent with pricing endogeneity: posted prices appear correlated with latent, product-specific demand factors not captured by $Z_i$, plausibly reflecting decentralized pricing with private demand information across the product catalog.
The RMSE in the Exposure-Sequence task is generally lower than that in the Static-Top3 task, aligned with the fact that (a) the Static-Top3 task tests on a holdout that is not seen in training, while for the Exposure-Sequence model, the holdout point may have appeared in the training price-points, and (b) conditioning on recent history yields a more temporally relevant signal. The performance of SHARED is competitive compared to META, outperforming it on the Exposure-Sequence task. This could possibly reflect a combination of endogeneity and the fact that META has to learn a more complex model with the same supervisory budget as SHARED. META-NA, moreover, only has effectively about half this budget, resulting in a substantially worse performance. Finally, {PER-TASK} performs poorly on Static-Top3 (RMSE 200.50; omitted for readability), highlighting the gains from cross-task transfer even when products are heterogeneous.
This paper studies multi-task demand learning when historical prices are endogenously chosen and may be arbitrarily correlated with unobserved demand drivers. In this regime, standard pooling and outcome-conditioned meta-learning can converge to policy-confounded estimands and fail to recover causal price effects, even with many tasks.
We propose a simple information-design principle to restore identification: the learner must condition on the endogenous decision history, yet it must not be able to deterministically infer which decision is used for supervision. This yields Decision-Conditioned Masked-Outcome Meta-Learning (DCMOML), which conditions on the full within-task price path while masking outcomes at two candidate query indices and randomizing (equivalently, averaging) the query choice. Under a non-selection condition on the query-period shock and with only two distinct historical prices per task, we show the population risk has a unique minimizer equal to the desired conditional target, and that ERM is consistent under standard regularity conditions. We validate the performance of our approach on both synthetic data with varying degrees of confounding and a retail e-commerce transaction dataset.
Future work includes extending DCMOML to richer demand models (nonlinearities, cross-price effects), characterizing efficiency and optimal masking when more than two prices are available, integrating identification with downstream pricing optimization under distribution shift, and developing diagnostics to decide when outcome masking is warranted in practice.