The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
129,672 characters
ACT, WAIT, or EXPERIMENT: A Causal Governance Framework for Retail Price Optimization Under Abstentions
\maketitle
\begin{abstract}
\noindent In intermediated retail channels, estimating consumer price elasticity from wholesale list prices is confounded by retail pass-through, promotions, competitor movements, and local market frictions. Rather than focusing solely on point estimation, this paper introduces a unified causal governance framework under which decision abstention (\textsc{wait}) is treated as a primary, diagnostic output rather than an estimation failure. Grounded in the companion theory of design-based uncertainty (\emph{Delgado, 2026}), we show that conventional within-panel bootstrap intervals suffer severe empirical subcoverage on this panel and that a hierarchical variance-component construction using Paule-Mandel estimation closes most of the gap; that construction is reported as validated rather than adopted, and every figure in this paper is produced under the bootstrap the implementation currently runs.
We present an \textsc{act/wait} decision system that combines Double Machine Learning, cost-shock instrumental contrasts, and conformal prediction to decouple weekly panel estimation from monthly decisions. The system evaluates recommendations through an admissibility gate and eight guards, jointly partitioning ten abstention causes, with the structural-break test (Guard~5) operating as a passive temporal drift sensor rather than an active veto, and the asymmetric Dual Beta estimator retired due to physical noise floor limits. The decision layer partitions all exit reasons exhaustively into three mutually exclusive verdict types under a fixed precedence: \textsc{wait-contamination} ($\mathcal{G}_{\text{CONT}}$), \textsc{wait-transient} ($\mathcal{G}_{\text{TRANS}}$), and \textsc{wait-identification} ($\mathcal{G}_{\text{ID}}$). The third is the operationally consequential one, as it marks a product presentation as a candidate for an actively designed pricing experiment rather than a terminal rejection.
This paper makes three main contributions: (1) an operational framework converting selective prediction into an active pricing diagnostic, leaving no presentation in an unclassified terminal state; (2) transforming seven tacit econometric assumptions into active system diagnostics; and (3) showing that aggregating unidentifiable presentation-level items to brand or category levels restores usable elasticity estimates (reducing RMSE from $0.571$ to $0.159$ against a true magnitude of $1.1$).
The system is tested against synthetic data-generating processes with known ground truth under thirty pre-registered engineering rules (11 pass, 18 fail, 1 not applicable); it has not been run against a commercial panel, and no result reported here is an observation of a real product category. Informative failures show that component discipline matters more than algorithmic complexity. On identifiable panels, the family-wise false-veto rate of the whole harness is estimated at $0.053$, one veto among nineteen independently seeded replications. Given the thin data typical of retail channels, reporting when evidence supports—or fails to support—a decision offers a practical and secure alternative to uncritical estimation.
\vspace{0.6em}
\noindent\textbf{Keywords:} Price optimization; causal inference; double machine learning; decision under abstention; demand systems; intermediated retail channels.
\end{abstract}
\clearpage
\section{Introduction}\label{sec:intro}
A manufacturer selling through decentralized retail sets the list price and does not set the
price the consumer pays. The shopkeeper does, subject to local margin targets, menu costs and
rigid psychological price points \citep{levy2011price}. Any elasticity estimated from
list-price variation in this channel is therefore a composite: the consumer's response to the
shelf price convolved with the shopkeeper's pass-through of the list price to the shelf
\citep{besanko2005own}. The operational problem is nonetheless posed at the list price, since
that is the only lever the manufacturer holds: set monthly list prices, bounded by
$\pm 20\%$, so as to maximize contribution without sacrificing volume or share.
Naive regressions of volume on historical prices fail here because they confound the causal
slope with promotions, competitor moves, seasonality and nominal inflationary drift
\citep{villas1999endogeneity}. Since an optimizer is a causal intervention machine
\citep{pearl2009causality}, the binding problem in this channel is not optimization but the
isolation of the slope the optimizer consumes. The revenue-management literature and the
enterprise deployments it documents \citep{talluri2006theory, hormby2010marriott,
deng2023alibaba, llenas2026pepsico} largely take the observational price--demand relation as
sufficient for prescription. Decentralized Retail violates that premise through three structural
conditions: scarce list-price variation confounded with promotion, indirect and heterogeneous
pass-through, and shelf prices captured by audit rather than by transaction.
Section~\ref{sec:related-work} situates that gap against six bodies of literature this work
draws on.
\paragraph{From estimation to decision.} Most of that literature, and most practice built on
it, treats the object of interest as a number: an elasticity to be estimated as precisely as
possible and handed to an optimizer. This paper treats it instead as an input to a decision
that is not always safe to make. Three questions organize what follows, and they replace ``what
is the elasticity'' as the paper's central question.
\begin{enumerate}
\item \textbf{When can a price change be executed on the estimate, and when should the system
withhold a recommendation instead?} This is the \textbf{ACT/WAIT} question. Treating
\emph{WAIT} as a first-class, reportable output, rather than as evidence of a system that
failed to produce a number, is the paper's first contribution
(Section~\ref{subsec:decision-layer}); Section~\ref{subsec:artifact} shows that a WAIT verdict
is also diagnostic of the remedy it calls for, not only a report that one is needed.
\item \textbf{What conditions, on the data and on the estimator, have to be made explicit and
checked before an ACT verdict is trusted?} Answering this converts assumptions such systems
ordinarily leave tacit into diagnostics the system executes rather than presumes. That
conversion, together with a declared account of which diagnostics adapt an existing
econometric tool and which are original to this setting, is the paper's second contribution
(Section~\ref{subsec:tacit}).
\item \textbf{At what level of commercial aggregation, presentation, brand or category, is a
decision legitimate when the evidence does not support it at the level the business would
prefer to act on?} Showing when and how far the unit of decision should move is the paper's
third contribution (Section~\ref{subsec:unit-of-identification}).
\end{enumerate}
This paper reports a system built to answer those three questions in one specific setting and,
more consequentially, the validation run that measured whether it does. The system decouples
weekly panel estimation from a monthly decision cycle and chains Double Machine Learning
\citep{chernozhukov2018double}, a cost-shock instrumental contrast, hierarchical pooling,
conformal predictive intervals, and a decision layer that abstains when identification is
unviable. It is validated against two synthetic generators with known ground truth under
thirty engineering-verification rules registered before the run. These thirty rules are a
quality-assurance device internal to the validation run, each pairing a claim the manuscript
makes with a threshold that would settle it (Section~\ref{subsec:validation-design}); they are
not the paper's scientific hypotheses, which are the three contributions listed above. Eleven
of the thirty rules pass, eighteen fail, and one does not apply, and several of the failures are
informative findings in their own right rather than simple defects.
Two further results support the third contribution and the abstention rule together.
Aggregating the estimate from the presentation to the brand or the category level substantially
improves its reliability, because design-specific bias among independently priced
presentations partly averages out: a category-level estimate reaches a root mean squared error
of $0.159$ against a true elasticity of magnitude $1.1$, where a presentation-level one reaches
$0.571$ (Section~\ref{subsec:unit-of-identification}). And on a panel this thin, a resampled
interval computed within one realized design cannot always be certified as honest, because it
cannot see the dispersion that would arise from having observed a different set of price
movements altogether; that limitation, not a deficiency of estimation, is why some
presentations, at some levels of aggregation, are correctly answered with WAIT rather than with
a number (Section~\ref{subsec:impossibility}). This paper motivates that limitation only to the
extent needed to justify the abstention rule; its full econometric treatment, the decomposition
of the shortfall, how it scales with aggregation, and the case for closing it experimentally, is
developed in companion methodological work \citep{delgado2026acrossdesignuncertaintyshortpricing} rather than here.
That evidentiary shape is not peculiar to this system or this sponsor. A hundred and twenty
weekly periods, a handful of presentations, list prices that move a few times a year and a
shelf price captured by audit describe the evidence available to most manufacturers in this
channel. Reporting when such evidence does, and does not, support a decision is more useful
than reporting a system that always appears to clear it.
The methodological claim the paper defends is broader than pricing. The failure modes that
survive a competent build are rarely errors of estimation; they are unstated assumptions
about the environment, and each can be converted into a quantity that costs less to measure
than its consequence costs to absorb. Four such assumptions were found by auditing the
design and three more by executing the validation. Table~\ref{tab:tacit} states all seven
with the contrast each became and, for each, whether the test adapts an existing tool or is
original to this work.
The remainder is organized as follows. Section~\ref{sec:related-work} situates the three
contributions against the literature and states the gap. Section~\ref{subsec:data} describes
the institutional setting and the panel. Section~\ref{sec:methodology} gives a global view of
the system that answers the first two questions above, with formal specifications deferred to
Appendix~\ref{app:methodological} and the extended exposition to
Appendix~\ref{app:extended-methodology}. Section~\ref{subsec:validation-design} states the
validation design. Section~\ref{sec:results} reports what the run measured, with extended
contrasts in Appendix~\ref{app:extended-results}. Section~\ref{subsec:ablations} reports
robustness checks and ablations. Section~\ref{sec:discussion} draws the governance and
deployment consequences, and Section~\ref{sec:conclusion} closes.
\section{Related Work and the Gap}\label{sec:related-work}
This work sits at the intersection of six literatures, and states plainly, at the outset, what
none of them by itself supplies.
\textbf{Structural demand estimation} in industrial organization
\citep{berry1994estimating, berry1995automobile, nevo2001measuring} identifies substitution
patterns and market power from aggregate market shares under strong assumptions about
functional form and instrument availability; it targets the estimate itself and is agnostic
about when a firm should act on it. \textbf{Vertical pass-through}
\citep{villasboas2007vertical, nakamura2010accounting} documents that retailers rarely
transmit a wholesale price change to the shelf one-for-one, which motivates this paper's
treatment of pass-through as a process rather than a constant
(Section~\ref{subsec:decentralized-trade}), but does not address how a manufacturer without
transaction-level shelf data should decide whether to move a price at all.
\textbf{Non-exchangeable conformal inference} \citep{gibbs2021adaptive, barber2023conformal}
supplies the machinery this paper uses to bound the volume response under a chronological,
non-stationary split (Section~\ref{subsec:conformal-layer}); it is a tool the framework adopts,
not a decision rule in itself. \textbf{Selective prediction and learning-to-defer}
\citep{elyaniv2010foundations, mozannar2020consistent} formalize the general principle that a
system can decline to answer rather than answer badly, and the ACT/WAIT gate is an instance of
that principle applied to a pricing decision; the contribution here is the instantiation, the
declared diagnostics that trigger abstention in this setting, not the general principle.
\textbf{Multiway panel clustering} \citep{chiang2022multiway} and the broader inference
literature it belongs to inform the resampling and clustering choices audited in
Section~\ref{subsec:causal-identification}. And \textbf{performative prediction}
\citep{perdomo2020performative}, alongside the general critique that a relation estimated under
one policy regime need not survive a new one \citep{lucas1976econometric, mackenzie2006engine},
names the risk that once this system moves prices, the observational identification it relies
on will erode; Section~\ref{subsec:cross-system} treats that risk as a governance problem this
paper flags rather than solves.
A seventh, more practically oriented literature sits alongside these six and is worth
separating from them: enterprise deployments of pricing and revenue-management systems
\citep{talluri2006theory, marn1992managing, hormby2010marriott, deng2023alibaba,
llenas2026pepsico}. These report what a deployed system achieved, typically a revenue or margin
gain, and are the closest thing to a track record this literature has for whether firms use
such systems at all and what they get from them. They are also, for the purpose of this paper,
the sharpest illustration of the gap: none of them reports which of the system's identifying
assumptions were checked, whether the reported gain would survive an audit of confounding, or
at what level of aggregation the underlying estimates were trustworthy. The comparison is
scoped to the three cited here and is not offered as a survey finding, but it is a repeated
pattern and not a coincidence: a literature organized around outcomes achieved has little
occasion to report on evidence that was insufficient, because a system that only ever reports
what it recommended has no channel through which to report having declined to recommend
anything. The adjacent literature on algorithm aversion
\citep{dietvorst2015algorithm, dietvorst2018overcoming} suggests why that omission is costly
rather than harmless: trust in an automated recommendation is fragile once the system has been
seen to err, and is rebuilt by giving the user visible, legible control over an imperfect
system, of which a declared and costed abstention is the clearest form.
\textbf{The gap.} None of the seven literatures above treats the decision to abstain as a
first-class object with its own diagnostics, its own measured cost, and its own conditions for
reversal; the closest, selective prediction, formalizes the principle in a general
classification setting and does not address what the relevant diagnostics are in an
observational pricing panel with retail intermediation. None specifies the conditions under
which a decision illegitimate at one level of commercial aggregation becomes legitimate at a
coarser one, or measures how far aggregation actually closes that gap. And none converts the
assumptions a causal pricing system ordinarily leaves tacit, that controls are pre-treatment,
that an instrument acts through one channel, that a comparison group is inert, that no adjacent
optimizer touches the same unit, into diagnostics the system itself executes and reports.
Sections~\ref{subsec:decision-layer}, \ref{subsec:tacit} and
\ref{subsec:unit-of-identification} address the three gaps in that order, each corresponding to
one of the questions posed in Section~\ref{sec:intro}.
\section{Institutional Setting and Data}\label{subsec:data}
The institutional detail behind the identification problem stated in
Section~\ref{sec:intro} is what makes this channel harder than the modern-trade setting most of
the causal-pricing and revenue-management literature addresses: the manufacturer's price is not
the consumer's price, the intermediary between them is a liquidity-constrained small business
rather than a chain with scanner data, and the panel available to measure any of it is short by
the standards of that literature. This section describes the panel this paper's system runs
on, before Section~\ref{sec:methodology} describes what the system does with it.
The valuation dataset is a weekly panel at the presentation~$\times$~region~$\times$~week
level, integrating sell-in, market-audit sell-out, distributor point-of-sale, competitor
pricing, promotional schedules and variable costs, with weather and holiday covariates. Gap
imputation and publication rules belong to an upstream data contract: gap imputation by forward fill, capped at two weeks and carrying an
indicator flag, together with non-null constraints, plausible price--volume bounds,
price--cost coherence and weekly partitioning. The system neither generates nor repairs
missing imputations, and validates only a diagnostic alert for week-over-week volume drops
above $80\%$ within a cell, as a proxy for unmarked stockouts; that alert is specified here
and executable exclusively against a production panel, so it is not exercised by any result
in this paper. The boundary is methodologically load-bearing rather than administrative: if
imputation flags concentrate at month-end among smaller retailers, forward fill is masking
order refusals rather than random missingness, which is evidence for the extensive-margin
mechanism of Appendix~\ref{subsec:constraints} rather than a capture defect.
The atomic unit is the \textbf{presentation} (brand $\times$ format family $\times$ pack
size), not the SKU: list prices are assigned per presentation, so intra-presentation SKU
variation carries no identifying price variance. The brand grounds the pooling hierarchy, the
format family supplies the instrumental contrasts, and the pack size governs consumption
occasions and price stickiness. The framework organizes around the structural hierarchy
$\text{Category} > \text{Brand} > \text{Presentation}$ rather than around any category domain,
and assumes a portfolio architecture with multiple brands, distinct pack sizes, uninstrumented
competitor prices and retail intermediation. The generators of
Section~\ref{subsec:validation-design} instantiate one portfolio exhibiting that topology and
the system is shown to run end to end on it; transfer to other portfolios sharing the
structure remains a design expectation rather than a verified property. This hierarchy is also
what Section~\ref{subsec:unit-of-identification} later aggregates over when presentation-level
evidence is not enough to decide.
The \textbf{week} is the temporal unit of inference throughout: cross-fitting folds, block
bootstrap, conformal splits, break tests and recency windows all operate on weekly
boundaries, while the month governs only the decision cadence. That division is an
identification decision and not a convenience. Under monthly aggregation the quota-pressure
guard loses its contrast identically, the calendar component of the cash-pressure index
disappears, and the break tests fall from a hundred and twenty observations to twenty-eight.
The one genuine argument for the monthly grain, that it would cancel forward-buying bias
within the observation, is already answered by taking sell-out rather than sell-in as the
outcome. Appendix~\ref{subsec:grain} develops the argument in full.
\paragraph{Feasibility and admissibility.} A presentation~$\times$~region cell is
\textbf{modelable} when five joint conditions hold on \emph{deflated} prices: a coefficient of
variation above $0.03$; at least four discrete price changes ($|\Delta\log p| > 0.005$); an
absolute price--promotion correlation below $0.6$; an absolute price--competitor correlation
below $0.9$; and non-frozen shelf status, defined conjunctively as
$\mathrm{CV}(P^{so}) \ge 0.01$ together with at most $5\%$ zero-volume weeks. A minimum of
forty usable observations applies. A presentation is then \textbf{admissible} when the
diagnostic passes in at least half its regions, the treatment partial $R^2$ after
residualization exceeds $0.10$, and the bootstrap interval width for the raw elasticity falls
below $0.6$. Admissibility is evaluated before the guards, so a failing presentation reaches
WAIT without a guard being evaluated, and it conditions conformal calibration and QA-Gate
scoring. Non-modelable units inherit a brand-level pooled elasticity and are marked
non-actionable. Appendix~\ref{subsec:data-diagnostic} states the diagnostic in full, including
the reduced-form competitor reaction function $\hat r$ that makes explicit which estimand the
competitive bracket contrasts. These two thresholds, jointly with the eight guards of
Section~\ref{subsec:decision-layer} and the ten abstention causes they define, are the diagnostics referred to in the second question of
Section~\ref{sec:intro}: Table~\ref{tab:tacit} and Section~\ref{subsec:tacit} state which of
them adapt an existing tool and which are original to this setting.
\section{Methodology}\label{sec:methodology}
The system executes a sequential, artifact-driven process: diagnosis, causal identification,
hierarchical pooling, conformal guarantees, and decision optimization.
Table~\ref{tab:layers} states each layer in one line together with the verdict the validation
run returned on it, and the subsections below develop the argument each layer rests on.
Formal specifications and operating parameters are in Appendix~\ref{app:methodological}, and the
extended exposition from which this section is condensed is in
Appendix~\ref{app:extended-methodology}.
\begin{figure}[H]
\centering
\sbox\pandoc@box{\includegraphics[keepaspectratio,alt={The Pricing Intelligence system}]{Figure1_Pipeline_Compact.png}}
\Gscale@div\@tempa{\textheight}{\dimexpr\ht\pandoc@box+\dp\pandoc@box\relax}
\Gscale@div\@tempb{\linewidth}{\wd\pandoc@box}
\ifdim\@tempb\p@<\@tempa\p@\let\@tempa\@tempb\fi
\ifdim\@tempa\p@<\p@\scalebox{\@tempa}{\usebox\pandoc@box}
\else\usebox{\pandoc@box}
\fi
\caption{Schematic overview of the system. Solid arrows represent primary algorithmic flow,
dashed arrows denote control and feedback mechanisms, and the two terminal verdicts are
distinguished by outline: solid for ACT, dashed for WAIT.}\label{fig:pipeline}
\end{figure}
\begin{table}[H]
\centering\small
\caption{The system in one line per layer, with the verdict the validation run returned.
The right-hand column is the paper's organizing discipline applied to its own architecture:
a layer is reported as validated only where an experiment measured its claim against a known
truth.}\label{tab:layers}
\begin{tabular}{@{}p{0.20\linewidth}p{0.33\linewidth}p{0.39\linewidth}@{}}
\toprule
Layer & What it does & Verdict of the run \\
\midrule
Variation diagnostic and admissibility gate
& Five cutoffs on observable variation plus three on the estimate; routes non-identifiable cells to WAIT
& Separates cleanly on the configuration it was calibrated on; precision falls to $0.000$ once evaluated outside it (\S\ref{subsec:gate-generalization}) \\
Causal identification (DML)
& Partialling-out with contiguous block cross-fitting; estimand $\theta=\varepsilon\rho$
& Recovers $\theta$ in the identifiable regime at a bias of $0.143$; does not degrade gracefully elsewhere (\S\ref{subsec:recovery}) \\
Instrumental contrast and exclusion audit
& Cost shocks as instrument; Hausman-type contrast with a residual-correlation audit
& Exclusion likely violated on this panel; remedy is data acquisition, not modeling \\
Hierarchical pooling
& Normal--normal shrinkage over Category $>$ Brand $>$ Presentation
& Improves the point estimate; its interval properties relative to the bootstrap are treated in companion work (\S\ref{subsec:hierarchical-coverage}) \\
Conformal layer
& CQR, Mondrian by region, chronological split
& The one uncertainty component whose measured behaviour matches its claim, verified out-of-time at rolling origins (\S\ref{subsec:coverage-experiment}) \\
Demand system (LA-AIDS on residuals)
& Cross-elasticities identified through the same residualization kernel
& Homogeneity and symmetry each rejected in $0.625$ of replications (\S\ref{subsec:ablations}) \\
Pass-through layer
& $\rho$ as a process with states; deconvolution $\varepsilon = \theta/\rho$
& Estimable and well recovered ($\hat\rho$ bias $-0.003$); immaterial to the decision, at $\le 0.6\%$ of band variance \\
Optimizer
& Exact $\arg\max$ of contribution over a discrete corridor
& Not the binding component; deterministic and global by construction \\
Decision layer (gate + eight guards, ten causes)
& Four-way verdict per presentation (ACT / \textsc{wait-cont} / \textsc{wait-trans} / \textsc{wait-id}), voting on the raw estimate
& Family-wise false-veto rate $0.053$ on an identifiable panel ($N=19$, \S\ref{subsec:funnel}); Guard~5 withdrawn as not operative \\
Experimental layer
& Holdout, synthetic control, permutation inference, contamination test
& Calibrated between twelve and thirty independently assigned units (\S\ref{subsec:ablations}) \\
\bottomrule
\end{tabular}
\end{table}
\subsection{Causal identification}\label{subsec:causal-identification}
Estimation is the partialling-out estimator with cross-fitting from Double/Debiased Machine
Learning \citep{chernozhukov2018double} applied to a partially linear model
\citep{robinson1988root} and grounded in Frisch--Waugh--Lovell
\citep{frisch1933partial, lovell1963seasonal}. Gradient-boosted learners
\citep{friedman2001greedy} model the joint dependence of treatment
$T = \log P^{si}$ and outcome $Y = \log Q^{so}$ on a confounder set $W$ carrying promotional
flags and intensity, log competitor prices, temperature, holidays, working days,
week-of-year and log inflation, extensible to distribution, out-of-stock, media and trade
investment where the panel supplies them.
The estimand requires one statement that is often left implicit. The coefficient of the
residualized outcome on the residualized treatment is not the structural consumer elasticity:
it is the convolution $\theta = \varepsilon\cdot\rho$ of consumer response with shopkeeper
pass-through. Separating the two is the job of the pass-through layer, not of a better
nuisance model, and comparing $\hat\theta$ against an injected consumer elasticity would
report transmission as estimation bias.
Four properties of the construction carry consequences the run measures.
\emph{(i)}~Cross-fitting uses $K=5$ contiguous temporal blocks
\citep{bergmeir2012use}; this is not forward-chaining and relies on stationarity and weak
dependence \citep{chiang2022multiway} rather than on Neyman orthogonality, which concerns
moment conditions and not information sets. An embargoed variant of that partition is
implemented and its coverage measured, and it is \emph{not} the one this run uses: adopting it
would change the residualizer that two of the ablations compare, so those contrasts would stop
testing what the text says they test. The consequence is declared rather than absorbed, because
it is a leak and not a preference: without the embargo the partition can place a single week on
both sides of the split through a different region. \emph{(ii)}~Nuisance learners are shallow boosted trees, which are piecewise constant
and represent a monotone trend as a staircase, so low-frequency components can survive into
the residuals. \emph{(iii)}~Inference is a moving-block bootstrap clustered by week
\citep{kunsch1989jackknife} with block length
$L = \mathrm{clip}(\mathrm{round}(1.5\,W^{1/3}), 2, \lfloor W/3\rfloor)$, giving $L=7$ at
$W=120$. \emph{(iv)}~The point estimate is the median across folds rather than the pooled
DML2 aggregate. Each of the four is contrasted against its alternative in
Section~\ref{subsec:ablations}, and two of them do not survive as general recommendations: the
choice of nuisance learner reverses between the two generators, and the choice of point
estimator reverses between learners.
\paragraph{Temporal status and the mediator bracket.} Every column of $W$ carries a declared
status, \texttt{pre} or \texttt{post}, relative to the pricing decision, and estimation aborts
on an undeclared control. Conditioning on post-treatment realizations blocks causal pathways
and changes the estimand \citep{rosenbaum1984consequences, vanderweele2015explanation}, so
elasticities are reported with and without post-treatment commercial controls and the
differential is published as a diagnostic rather than used to prune automatically. Under the
sign structure assumed for this channel, trade activity rising with list price, the indirect
path carries the opposite sign to the direct effect, so over-conditioning inflates
$|\hat\theta|$; since the optimizer's interior optimum carries a markup $\theta/(1+\theta)$
falling toward one as $|\theta|$ grows, the failure mode is a systematically conservative
recommendation that stays plausible. The direction depends on the assumed sign and not on the
estimator, and is declared as conditional for that reason.
\paragraph{Two structural contrasts.} Residualization addresses observed confounders only, so
two contrasts probe what it cannot. The \textbf{instrumental contrast} uses log variable
costs as an instrument for supply-side shocks, and the divergence between observational and
instrumental estimates is an operational Hausman test \citep{hausman1978specification}.
Category-wide cost shocks move competitor prices too, violating the exclusion restriction and
displacing both estimators together, which \emph{understates} the contrast
\citep{angrist1996identification}; the framework therefore audits the restriction directly
through the residual correlation between the residualized instrument and the residualized
competitor price, with an explicit time index added because a boosted tree partials a
monotone drift poorly. The \textbf{competitive bracket} decomposes price response into three
declared estimands, conditional, historical-equilibrium and cross, and admits the internal
identity $\theta_{\text{total}} - \theta_{\text{cond}} \approx \theta_{\text{cross}}\cdot
\hat r$. The identity binds only where the competitor moves and reacts, and
Section~\ref{subsec:recovery} establishes on which panel it is contrastable before reporting
it. Both are specified formally in Appendices~\ref{app:pliv-contrast}
and~\ref{app:competitive-bracket}.
\subsection{Regularization, uncertainty and the demand system}\label{subsec:regularization-uncertainty}
The causal layer returns one slope per presentation with an interval, and three problems
remain that no single-unit estimator solves: thin panels borrow strength or return noise, a
point estimate is not a bound on the volume a committee will observe, and an own-price
elasticity says nothing about what the rest of the portfolio does.
\paragraph{Partial pooling.} The hierarchy $\text{Category} > \text{Brand} >
\text{Presentation}$ shrinks each estimate toward its branch through a conjugate
normal--normal update, with three origins for the prior: a leave-one-out brand mean with
between-presentation heterogeneity $\hat\tau^2$ estimated by Paule--Mandel
\citep{paule1982consensus}; the previous cycle, through a local-level dynamic linear model at
discount $\delta = 0.90$ \citep{west1997bayesian}, whose information weight halves in roughly
seven cycles; and a cold-start branch prior. Two pitfalls are avoided by construction, using
the sample variance in place of $\hat\tau^2$ and summing raw precisions rather than
incorporating $\hat\tau^2$ into branch precisions \citep{efron1975data}. The output is a
decision-grade elasticity $\theta_{\text{dec}}$ carrying its posterior variance, own-data
weight and declared origin. Its role is bounded deliberately: the ACT/WAIT gate votes solely
with the raw estimate $\hat\theta$, so pooling can never unlock a movement, while
$\theta_{\text{dec}}$ sets the level of the published price and elasticity curves.
Appendix~\ref{app:partial-pooling} gives the update and two declared approximations.
\paragraph{The heterogeneity component and design variance.} The between-presentation
heterogeneity $\hat\tau^2$ that this layer estimates in order to set the shrinkage weight also
measures something a within-panel bootstrap cannot: dispersion that comes from the realized
design itself, this set of price movements falling in these weeks, rather than from sampling
noise around it (Section~\ref{subsec:hierarchical-coverage} sketches why, and treats the
question at the level this paper needs). An interval of the form
$\hat\theta \pm 1.96\sqrt{s^2 + \hat\tau^2}$ is one way to fold that source back into the
reported uncertainty; whether, and by how much, it improves on the production bootstrap
interval is a question of its own, developed in companion methodological work \citep{delgado2026acrossdesignuncertaintyshortpricing} rather than
here. The change is \emph{identified and not yet adopted} in the production configuration, and
every figure in this paper is produced under the interval the implementation currently runs.
One consequence for this layer is unfavourable and is stated with the rest. Where the raw
estimate is mis-centred, shrinkage can move the interval toward a displaced branch mean while
narrowing it, so the pooling posterior's coverage is not guaranteed to improve on the
bootstrap it shrinks. Partial pooling is not assumed to help the interval merely because it
helps the point estimate, and a system that publishes both should say so.
\paragraph{The conformal layer.}\label{subsec:conformal-layer} Predictive intervals on log volume are built by
Conformalized Quantile Regression \citep{romano2019conformalized}, Mondrian-stratified by
region \citep{vovk2003mondrian}. Where a stratum is too sparse for local quantile estimation
the estimator falls back to global calibration, which relaxes the guarantee from
group-conditional to marginal coverage, and each fallback is recorded by a unit-level indicator. Because random splits in a weekly panel mix past and future
and produce optimistic bands, the partition is strictly chronological ($50/30/20$) and
evaluated across rolling origins. Non-exchangeability forfeits the finite-sample guarantee
\citep{vovk2005algorithmic, lei2018distribution}, so the construction rests on the adaptive
conformal argument for non-stationary environments \citep{gibbs2021adaptive, barber2023conformal}
and on measured out-of-time coverage that clears the nominal $0.90$ level at the rolling origins
tested (Section~\ref{subsec:coverage-experiment}). The band
over-covers, which is the conservative direction and costs width rather than validity, and it
bounds the volume response only: price bounds are enforced independently by the competitive
corridor, and intersecting the two would mix an interval on a response with a bound on a
decision variable.
\paragraph{Cross-elasticities.} Portfolio decisions require knowing how much presentations
cannibalize one another. The system estimates an AIDS demand system \citep{deaton1980almost}
in its linear approximation over revenue shares deflated by the Stone index
\citep{stone1954linear}, and the identification correction is the layer's contribution:
standard practice regresses shares on raw prices, which confounds the price effect with
everything that moved at the same time and invents spurious substitution. Here the system
runs exclusively on the residuals of the same causal model, reusing the single cross-fitting
kernel, so by Frisch--Waugh--Lovell the coefficients coincide with those of the full system
without inheriting the endogeneity of the observational price. The theoretical restrictions
are tested rather than imposed. Section~\ref{subsec:ablations} reports the test and it is
unfavourable. Three declared simplifications, the observed-share Stone deflator
\citep{moschini1995units}, date-level collapse of $W$, and the amplification of noise by
division through small revenue shares, are stated in Appendix~\ref{app:aids}.
\subsection{Intermediated-Retail extensions}\label{subsec:decentralized-trade}
Two extensions describe a single agent, the shopkeeper who sets the shelf price and who buys
under a liquidity constraint.
\paragraph{Pass-through as a process, not a constant.} Treating $\rho$ as an implicit constant
is the standard error in this channel: measured pass-through is neither unitary nor
homogeneous \citep{besanko2005own}, and two presentations with identical observed elasticities
under different transmission regimes call for different decisions. The system estimates $\rho$
per presentation from a residualized regression of the log shelf price on the log list price
and classifies the regime against two cutoffs, \emph{sticky} below $0.30$ where the shelf is
anchored at a round point and the list lever loses traction, and \emph{margin-defending} at or
above $0.80$. To that level it adds a Markov chain over the shopkeeper's margin state,
conditioned on the sign of the list shock, and the deconvolution
$\hat\varepsilon = \hat\theta/\hat\rho$, which is reported as an interpretive diagnostic only:
decoupling introduces a negative structural covariance, so the optimizer and every band
operate directly on $\hat\theta$, and $\hat\varepsilon$ is undefined below $\hat\rho = 0.30$.
What the region~$\times$~week aggregate does not identify is the discrete rigidity of
repricing, the $(S,s)$ behaviour \citep{sheshinski1977inflation, nakamura2008five} in which the
shelf jump depends on untransferred cost accumulated since the last change; aggregation
smooths it into a falsely gradual pass-through \citep{caballero2007price}. A three-layer
structural architecture for store-level panels is specified in
Appendix~\ref{app:shopkeeper} and not estimated here. One asymmetry in the estimation is
declared because a reader will otherwise find it: $\rho$ is residualized against a reduced
confounder set. Excluding the competitor's list price is defensible, since it is not a
confounder of the shopkeeper's own margin decision; excluding the inflation index is not,
since it is a common factor of both prices, and the run measures what the conflation costs
(Appendix~\ref{subsec:passthrough}).
\paragraph{The budget constraint of the intermediate buyer.} A store may decline or truncate
an order for reasons unrelated to how well the product sells. The intuitive measurement, the
store's observed outlay with the manufacturer, is built from the treatment and the outcome
simultaneously: admitted as a control it converts the elasticity into within-basket
substitution at constant expenditure, and recommends a larger increase precisely where
substitution room exists. The construct is therefore decomposed into three objects with three
entry points: a \textbf{cash-pressure index} built only from sources exogenous to the firm's
own sales, admitted to $W$ with \texttt{pre} status and published with its variance
decomposition; \textbf{total store expenditure}, which belongs in the demand system as the
instrumented allocation variable \citep{blundell1999estimation}; and the \textbf{binding state}
of the constraint, which is an effect modifier and not a control. Order refusal is an
extensive-margin event requiring a two-part treatment \citep{cragg1971some} under selection on
the constraint state \citep{heckman1979sample}.
The status of this frontier is stated exactly, because it is an instance of the paper's
governing principle rather than a gap. One component is implemented, a cheap screening test
that regresses portfolio breadth on the cash-pressure index and requires $|t| \ge 2$ before
any further construct is built. Three are specified and \emph{not} estimated, because a
region~$\times$~week panel cannot support them and estimating them badly at this grain would
smooth a lumpy discrete process into a mild and false elasticity. What the screen returns on
this generator, and why, is reported in Section~\ref{subsec:ablations}.
\subsection{Optimization and the ACT/WAIT decision layer}\label{subsec:decision-layer}
The \textbf{optimizer} maximizes contribution $(P-C)\hat Q(P)$ subject to minimum margin,
volume, share and ladder constraints, the operating limit of $\pm 20\%$ on the base list
price, and the competitive corridor $[P^{\text{comp}}(1-\kappa), P^{\text{comp}}(1+\kappa)]$
at $\kappa = 0.10$. The feasible set is the intersection of these, the tighter bound governs,
and the artifact records which one bound. For a single presentation that set is
one-dimensional and discrete, since prices are set in whole units of the smallest circulating
denomination, so exact enumeration delivers the global optimum deterministically and the grid
\emph{is} the feasible set rather than a discretization of it; a metaheuristic would add
stochastic dependence and buy nothing. Genuine non-convexity appears only in the joint
portfolio vector, where cross-cannibalization couples decisions, and differential evolution
\citep{storn1997differential} with L-BFGS-B polishing \citep{byrd1995limited} is reserved for
that scale-up. A \textbf{Score} combining normalized margin and volume ranks scenarios for the
committee and never defines the optimum, since maximizing a normalized score is not
maximizing contribution.
Every recommendation travels with three quantities that a committee can adjudicate: the
\emph{tolerable volume loss} $L = \Delta p/(m + \Delta p)$, which is algebraically exact and
free of functional-form assumptions; the \emph{critical elasticity} at which moving and
holding are indifferent, exact conditional on the log-log specification the optimizer itself
uses; and a flag stating whether the whole uncertainty interval falls on one side of that
threshold. The motivation is empirical: the canonical meta-analysis reports a mean elasticity
of $-1.76$ with a dispersion as wide as the mean across 367 estimates \citep{tellis1988price},
and a later synthesis over 1{,}851 estimates does not agree with it in the first digit
\citep{bijmolt2005new}. Publishing an elasticity to two decimals promises a precision the
method cannot deliver; reporting which side of the break-even the whole range falls on
delivers what the decision requires.
\paragraph{The decision layer.} The decision layer adds no analysis. It converts what has been
estimated into one verdict per presentation and makes the optimizer act only where
identification holds. It is the least common component in enterprise pricing systems and the
one on which organizational trust turns: reluctance to use a model intensifies once it has
been seen to err \citep{dietvorst2015algorithm}, and is recovered by giving the user visible
control over an imperfect system rather than by improving the model
\citep{dietvorst2018overcoming}.
ACT requires passing every \emph{operative} guard simultaneously. The qualifier is
load-bearing: a guard that cannot be evaluated is reported as such rather than counted as
passed, so a presentation reaching ACT on seven evaluable guards is distinguishable from one
reaching ACT on eight. A verdict of WAIT, however, is not a single state. The decision layer
partitions the abstention causes into three mutually exclusive types, and the partition is
what makes the abstention diagnostic rather than terminal: it states what would have to change
for the presentation to become actionable, and in one of the three cases that change is an
experiment this system can specify.
\begin{definition}[Three-way classification of guard-level abstention causes]\label{def:extended-classification}
The admissibility gate and the eight guards define ten guard-level abstention causes, since
Guard~1 evaluates two conditions separately: a regime shift in $\theta$ on excluding broken
weeks (\emph{1a}) and an excessive share of weeks so flagged (\emph{1b}). Let
$\mathcal{S}$ be the vector of diagnostic outcomes for a presentation in a cycle. The ten
causes partition into three disjoint subsets:
\begin{itemize}
\setlength{\itemsep}{0pt}\setlength{\parskip}{0pt}
\item $\mathcal{G}_{\text{ID}}$ (\textsc{identification}, five causes): failing admissibility;
Guard~1a; Guard~1b; Guard~2; Guard~3. Each is a precision shortfall, or natural variation
contaminated by reactive pricing, that independently randomized perturbation directly resolves.
\item $\mathcal{G}_{\text{CONT}}$ (\textsc{contamination}, four causes): Guard~4; Guard~6;
Guard~7; Guard~8. Guard~4 belongs here rather than with precision because a strong-instrument
divergence is direct evidence of residual endogeneity in the observational variation itself,
by the same logic already applied to Guard~7. Each signals a specification or validity problem
that an experiment does not, by itself, repair.
\item $\mathcal{G}_{\text{TRANS}}$ (\textsc{transient}, one cause): Guard~5. Neither more
identifying variation nor a specification fix addresses a regime that has just ended; the
estimate becomes usable again once the new regime has accumulated its own history.
\end{itemize}
The decision layer is the mapping
$\mathcal{D}(\mathcal{S}) \rightarrow \{\text{ACT},\ \textsc{wait-contamination},\
\textsc{wait-transient},\ \textsc{wait-identification}\}$ defined by
\[
\mathcal{D}(\mathcal{S}) =
\begin{cases}
\text{ACT}, & \mathcal{G}_{\text{CONT}} = \emptyset \;\wedge\; \mathcal{G}_{\text{TRANS}} = \emptyset \;\wedge\; \mathcal{G}_{\text{ID}} = \emptyset,\\[4pt]
\textsc{wait-contamination}, & \mathcal{G}_{\text{CONT}} \neq \emptyset,\\[4pt]
\textsc{wait-transient}, & \mathcal{G}_{\text{CONT}} = \emptyset \;\wedge\; \mathcal{G}_{\text{TRANS}} \neq \emptyset,\\[4pt]
\textsc{wait-identification}, & \mathcal{G}_{\text{CONT}} = \emptyset \;\wedge\; \mathcal{G}_{\text{TRANS}} = \emptyset \;\wedge\; \mathcal{G}_{\text{ID}} \neq \emptyset.
\end{cases}
\]
\end{definition}
The precedence is not a tie-breaking convenience. A single contamination-type cause blocks the
experimental route regardless of what else bound, since it signals an assumption violation that
independent randomization alone does not repair; a recent structural break next makes the
current estimate untrustworthy independent of its precision; and a presentation whose only
binding causes are identification-type is the one for which an actively designed price
experiment is the correct remedy. The definition adds no guard, no gate and no threshold: it
only completes an assignment the framework already applied informally to some causes and left
unstated for the rest, and it names a third category for the one cause, a structural break,
that fits neither of the first two.
Figure~\ref{fig:taxonomy} depicts the same partition graphically, together with the precedence
rule that resolves a presentation whose diagnostics trip causes from more than one subset at
once; Table~\ref{tab:decision-taxonomy} below restates it as the four verdicts in order of
precedence.
\begin{figure}[H]
\centering
\sbox\pandoc@box{\includegraphics[keepaspectratio,alt={The ten guard-level abstention causes grouped by subset}]{fig_taxonomy.png}}
\Gscale@div\@tempa{\textheight}{\dimexpr\ht\pandoc@box+\dp\pandoc@box\relax}
\Gscale@div\@tempb{\linewidth}{\wd\pandoc@box}
\ifdim\@tempb\p@<\@tempa\p@\let\@tempa\@tempb\fi
\ifdim\@tempa\p@<\p@\scalebox{\@tempa}{\usebox\pandoc@box}
\else\usebox{\pandoc@box}
\fi
\caption{The ten guard-level abstention causes of Definition~\ref{def:extended-classification},
grouped by the subset each belongs to ($\mathcal{G}_{\text{ID}}$, $\mathcal{G}_{\text{CONT}}$,
$\mathcal{G}_{\text{TRANS}}$) and resolved by the decision layer's fixed precedence:
contamination first, then transient, then identification. Guard~5 is drawn with a dashed
outline because it is computed and classified on every cycle but is not an active veto in the
validated configuration (\S\ref{subsec:guard5-power}). A \textsc{wait-identification} verdict is
the only one that qualifies a presentation as an Experiment Candidate; whether it is actually
assigned a designed experiment in a given cycle depends further on the committee's budget and
the portfolio subset-selection solution.}\label{fig:taxonomy}
\end{figure}
\begin{longtable}[]{@{}
>{\raggedright\arraybackslash}p{(\linewidth - 6\tabcolsep) * \real{0.24}}
>{\raggedright\arraybackslash}p{(\linewidth - 6\tabcolsep) * \real{0.30}}
>{\raggedright\arraybackslash}p{(\linewidth - 6\tabcolsep) * \real{0.46}}@{}}
\caption{The four verdicts of the decision layer under Definition~\ref{def:extended-classification}, in order of precedence. Every presentation in every cycle receives exactly one.}\label{tab:decision-taxonomy}\\
\toprule\noalign{}
Verdict & Nature of the block & What would change it \\
\midrule\noalign{}
\endfirsthead
\multicolumn{3}{@{}l}{\itshape Table D.\arabic{table}{} (continued)}\\
\toprule\noalign{}
Verdict & Nature of the block & What would change it \\
\midrule\noalign{}
\endhead
\bottomrule\noalign{}
\endlastfoot
ACT & None & Identification and stability both hold; the optimizer emits $P^*$ \\
\textsc{wait-contamination} & Structural: a validity assumption fails & Control of the offending channel, not more variation; an experiment alone does not repair it \\
\textsc{wait-transient} & Temporal: the regime just changed & Accumulated history under the new regime; the estimate becomes usable again without any intervention \\
\textsc{wait-identification} \emph{(Experiment Candidate)} & Observational: the available variation cannot resolve the effect & Independently assigned price variation. This is the verdict that qualifies a presentation as a formal candidate for portfolio experimental design \\
\end{longtable}
Table~\ref{tab:guards} states the gate and the eight guards in compact
form and Appendix~\ref{subsec:guards} develops them; several adapt an existing econometric
tool to this decision and others have no close equivalent in that literature, and
Section~\ref{subsec:tacit} states which is which for each of the diagnostics discussed there.
Guards~1--5 address the quality and the
currency of the estimate: contaminated weeks, the Cinelli--Hazlett robustness value
\citep{cinelli2020making}, the stability of the recommendation across the width of the
uncertainty, the instrumental contrast, and structural-break tests
\citep{ploberger1992cusum, chow1960tests, andrews1993tests} on weekly aggregates with a
Bartlett long-run variance \citep{newey1987simple}. Guards~6--8, added after an external
audit, address the ways an estimate can be well measured and still be measuring something
else: over-conditioning on mediators, a failed exclusion restriction, and cross-system
interference. Two mechanisms complement them: the elasticity is treated as a snapshot with an
expiry date, re-estimated on expanding windows with a drift alert; and every ACT-class
recommendation is re-evaluated under three competitive scenarios drawn from the bracket, which
is a reduced-form one-round device declared as such rather than structural game theory.
\paragraph{Cross-system interference.} No decision model in an enterprise operates alone, and
treating adjacent systems as inert is an unstated identifying assumption. The structural
neighbour here is a commercial order-recommendation system that functions as a quota
allocator: it predicts what the firm aims to place in the channel, not what demand will
absorb. Quota pressure is then an unobserved confounder with a temporal signature, since a
salesperson behind target discounts and pushes, and the incentive cycle concentrates that
contamination in the closing weeks where a weekly panel carries most of its price movement
\citep{oyer1998fiscal, larkin2014cost}. Forward buying reverses after a month-end push
\citep{hendel2006measuring}, which is an independent argument for taking sell-out rather than
sell-in as the outcome. And over-inventoried stores are \emph{conjectured} to discount on the
shelf to rotate, which would enter the pass-through layer as consumer behaviour; that third
channel has no source behind it and is declared as an operational conjecture, unlike the first
two. The correction is to admit the incentive cycle into $W$ and test its influence, which is
Guard~8, and to adopt a reporting convention under which every alternative reaching the
committee declares its expected cost to the adjacent system. One risk no econometrics
resolves: \textbf{performativity} \citep{mackenzie2006engine}, the pricing instance of the
general point that a relation estimated under one policy regime need not survive a new one
\citep{lucas1976econometric}. Once the engine moves prices, future variation is the child of
its own estimates. The defense is experimental, and it is instrumented rather than
recommended (Appendix~\ref{subsubsec:experimental-layer}).
\subsection{Architecture, quality gate and artifacts}\label{subsec:system-architecture}
Five principles are enforced in code rather than recommended in documentation: a
\textbf{single residualization kernel} feeding the causal estimator, the demand system and the
pass-through deconvolution; a \textbf{single temporal resampling policy}, because one component
that shuffles rows contaminates everything downstream with optimism; \textbf{labeled
placeholders, never disguised ones}; \textbf{no undeclared control}; and \textbf{one effective
unit of inference}, meaning that the count entering any statistic whose distribution depends
on a sample size is the number of independent \emph{periods}, never the number of panel rows.
The fifth is stated as a principle because its violation was found where the rest of the
system had already got it right. The bootstrap clusters on the week and the break tests
aggregate to it, but the robustness value and the first-stage $F$ statistic each consumed a
raw row count, which inflated their degrees of freedom by the number of regions and rendered
the weak-instrument screens effectively non-binding. That both instances survived a careful
build is the concrete form the paper's general thesis takes inside its own implementation:
the failure that survived was a tacit assumption about the unit of analysis, not an error of
estimation. Section~\ref{subsec:funnel} reports what correcting it does to the veto rate.
A single \textbf{QA Gate} halts the system with an explicit exception, rather than degrading
silently, if the out-of-fold fit of the demand model on the modelable set falls below $0.65$
or mean conformal coverage falls below the nominal level minus $0.03$. One correction to that
design is proposed and stated as such: the out-of-fold fit of a log-volume model is dominated
by seasonality and level, so it protects the band that predicts volume but not the slope that
requires the treatment to be identified. The quantity that protects the second is already
computed, the treatment partial $R^2$ at $0.10$, and the gate is therefore best presented as
two thresholds answering two questions. The deliverable of a cycle is not a table but a
versioned repository: a $\pm 20\%$ lookup table in $1\%$ steps with the percentage applied
once on the base price and a per-region support flag, elasticity curves by tranche, the
cross-elasticity matrix, and a schema-validated dashboard feed, all persisted by cycle and
presentation with quality flags and intervals. The feed's verdict field is closed over exactly
four categorical states, \texttt{ACTIONABLE}, \texttt{WAIT\_CONTAMINATION},
\texttt{WAIT\_TRANSIENT} and \texttt{WAIT\_IDENTIFICATION}, and the schema rejects any other
value, so a downstream consumer cannot silently receive an unclassified verdict; the binding
cause and the operative status of each guard travel alongside it (Appendix~\ref{subsec:architecture}).
\section{Validation design}\label{subsec:validation-design}
Every quantity reported below is measured against a data-generating process whose answers are
known by construction. The design premise is that a methodological claim about identification
is worth what its recovery experiment shows: a system that abstains correctly on a panel
where the truth is unknown is indistinguishable from one that abstains at random.
\paragraph{Two generators.} The \textbf{five-regime generator} simulates $W=120$ weekly periods
across six regions and six presentations spanning three brands and two categories, with volume
generated jointly with prices, costs, promotion, calendar effects, weather, minimum wage and an
inflation index. Five identification regimes vary the amount and kind of price variation
available, from a \emph{clean} regime with ample list movement to regimes built deliberately to
fail on weak variation, confounding with promotion, or collinearity with the competitor; four
further mechanisms inject pass-through, a cross-price response, forward buying and
margin-dependent delisting. Appendix~\ref{subsubsec:dgp} states every parameter.
A generator with known answers is also a generator with chosen answers, so the run uses a
second. The \textbf{extended generator} draws eight presentations across a wider set of brands
and varies pass-through, regime assignment and elasticity magnitude in three ways chosen to
remove what the first generator makes artificially clean; Appendix~\ref{subsubsec:dgp} states
the ranges. Results are reported on both generators wherever they disagree, and the
disagreements are among the more informative outputs of the run: a property that holds on one
and not the other is a property of a configuration and not of the method. Two features both
generators share bound what either can establish, and are declared: both are stationary apart
from an inflationary drift, and in both the confounding is injected at magnitudes the designer
chose.
\paragraph{Four layers plus an experimental one.} Validation runs in four layers: a
\textbf{statistical} layer requiring significance, a negative own-price sign, dominance of own
over cross elasticities and the QA-Gate threshold; a \textbf{commercial sanity} layer
confronting the estimated substitution hierarchy with the consumer decision tree the business
knows; a \textbf{back-testing} layer that deliberately decouples prediction from inference,
because a single predictive model conflates the two and answers both poorly; and a
\textbf{stress} layer at the extremes. The sign requirement carries a selection cost that the
run measures rather than assumes away (Section~\ref{subsec:ablations}).
None of the four can establish that an executed price change caused what followed it. The
\textbf{experimental layer} therefore implements a holdout drawn by a deterministic function of
the cycle identifier, so the assignment is auditable months later; a synthetic control with
convex unit weights reproducing the pre-treatment trajectory \citep{abadie2010synthetic}, with
its pre-fit RMSE published alongside the effect and the intercept adjustment of
\citet{doudchenko2016balancing}, the time weights of \citet{arkhangelsky2021synthetic} being
declared as the natural upgrade rather than claimed; permutation inference over the assignment
mechanism itself, since with six regions no asymptotic standard error would be credible; and a
contamination test of the no-interference half of SUTVA \citep{rubin1980randomization} against
the panel's own placebo distribution plus an economic floor of $0.005$ in log price.
Appendix~\ref{subsubsec:experimental-layer} states the protocol and its resolution limits.
\paragraph{Bookkeeping.} Three conventions govern how every figure below should be read, and
they are what separates a validation run from a demonstration. \emph{Pre-registered verdicts}:
thirty decision rules were written before the run, each pairing a claim the manuscript makes
with a threshold that would settle it; eleven pass, eighteen fail, and one does not apply
because the failure it was written to adjudicate did not reproduce. Table~\ref{tab:verdicts}
lists all thirty against their declared thresholds. These thirty rules are engineering
verification: a device against post-hoc selection, internal to this validation run and scoped
to individual components of the system. They should not be confused with the paper's
scientific contributions, the ACT/WAIT decision framework, the audit of tacit assumptions into
diagnostics, and the conditions for legitimate aggregation, stated in
Section~\ref{sec:intro}, which the run as a whole is designed to support and which no single
row of Table~\ref{tab:verdicts} either makes or breaks. \emph{Separated seeds}: where a design
choice was selected by measurement, the selection ran on one set of seeds and the published
figure on another, so a variant chosen on one panel does not inherit the selection bias of
having been the best of several. \emph{No single $R$}: the replication count is set per
experiment, from $300$ seeds for the panel-width and randomized-shock experiments down to $5$
for the re-run liquidity-screen and pass-through-heterogeneity scenarios, and the counts are
collected in Table~\ref{tab:replications} rather than folded into one
number that would overstate the evidential weight of the thinner blocks.
\section{Results}\label{sec:results}
\subsection{Recovery of the injected elasticities}\label{subsec:recovery}
Table~\ref{tab:recovery} reports, for each regime of the five-regime generator over 63
replications, the injected consumer elasticity and pass-through with the list-price elasticity
they imply, the median estimate, bias, root mean squared error, empirical coverage, and the
shares of replications passing the admissibility gate and reaching ACT.
The two rightmost columns are unambiguous. \emph{Clean} reaches admissibility in $0.968$ of
replications and ACT in $0.944$; every other regime reaches both in exactly $0.00$. In
\emph{clean} the estimate is $-0.962$ against a truth of $-1.105$, a bias of $0.143$ toward
zero, at a coverage of $0.484$ against a nominal $0.95$. Where identification fails the
estimator does not degrade gracefully: in \emph{collinear} the median estimate is $+0.248$,
the wrong sign, and in \emph{confounded} and \emph{weak} the estimates are $-0.103$ and
$-0.183$ against truths of $-1.190$ and $-0.850$, attenuated by an order of magnitude. That is
the failure mode which survives casual inspection, because a small elasticity is a
plausible-looking number.
\paragraph{The round-point regime does not produce a large spurious estimate.} A dedicated
mechanics experiment isolates the two candidate explanations for what a round shelf price could
do to the recovered elasticity, a collapsing denominator that the partial $R^2$ would catch
against a spurious correlation at ordinary treatment variance that it would not, and finds
support for neither: over $19$ replications the estimator recovers $\hat\theta = -0.069$ against
a truth of $0$, at a residual treatment variance of $0.6385$, comparable to the $0.5767$ of the
\emph{clean} reference, with the partial $R^2$ passing in every replication, a coefficient of
variation of the shelf price of $0.068$ confirming that the shelf does not transmit, and a
residual correlation of $-0.031$. The variation diagnostic correctly flags a pinned shelf as
such, but an outsized estimate is not what this experiment measures there: whatever produces one
is a property of a specific configuration and not of the round-point regime in general.
\paragraph{The extended generator.} Table~\ref{tab:recovery-ext} repeats the exercise where
$\rho$ is drawn per region, over 53 replications. Two results qualify the separation above. The
gate is no longer
perfect: the round-point presentation is admitted in $0.981$ of replications under the boosted
learner and the confounded presentation likewise in $0.981$, against $0.00$ for both on the
five-regime generator. Their estimates are not damaging ($0.062$, and a bias of $0.293$), so the
gate does
not admit a disaster, but it admits them, and the near-perfect separation is therefore a
property of the configuration the thresholds were set on. Where it matters most the two agree:
in the clean regime the original returns a bias of $0.143$ and the extended a mean bias of
$0.197$ across its four clean sub-presentations under the boosted learner, a difference of
$0.054$, so the estimation results transfer between generators even where the gating results do not.
\paragraph{The competitive bracket.} Reporting the internal consistency checks requires first
settling whether the panel can support them, and it cannot: the competitor's real log price
varies by a standard deviation of $0.0359$, which propagates to a signal of $0.0108$ in log
volume and, against a demand noise of $0.05$, places the injected $+0.30$ cross term below the
noise floor at a signal-to-noise ratio of $0.215$. The audit identity then holds trivially
because $\hat r \approx 0$ zeroes both sides. Both checks are reported instead on a variant in
which the competitor makes eight discrete moves and follows the own price at $r = 0.5$, raising
the dispersion of its real log price to $0.1161$. Over
31 replications the cross elasticity returns a median of $0.148$ against the injected $+0.30$, the
reaction is recovered at $\hat r = 0.502$, and the identity returns a median left-hand side of $0.153$
against a median right-hand side of $0.059$ at a median absolute error of $0.101$. The three are not
equally successful, and saying so is better than presenting a near-equality the data do not
deliver: the reaction function is recovered almost exactly, the cross elasticity at $49\%$ of
its injected value, and the identity has a residual the size of the quantity it constrains, so
it detects gross violations and nothing finer.
\subsection{Why a resampled interval is not enough on its own}\label{subsec:coverage-experiment}
The block-bootstrap interval that the causal layer reports is a standard resampling
construction \citep{kunsch1989jackknife}, clustered by week to respect the panel's serial and
cross-sectional dependence. Measured against known truth, it undercovers relative to its
nominal level, and testing eight alternative resampling and clustering schemes, at two nuisance
learners, does not close the gap: none reaches the coverage this paper requires before an
interval is allowed to certify a decision. The shortfall turns out to be a matter of where the
interval is centred rather than of how wide it is, which rules out simply widening it as a fix.
The conformal band used elsewhere in the system (Section~\ref{subsec:conformal-layer}) does
not share this problem; measured at rolling origins it remains the one component of the
uncertainty machinery whose behaviour matches its claim.
This finding is a statement about estimator behaviour rather than about any single pricing
decision, and a direct decomposition on this same panel is reported here rather than deferred
in full, because it is already computed and it is what turns the diagnosis above from
qualitative to quantitative. Over four independently drawn designs with forty noise draws taken
within each, holding the data-generating process fixed and redrawing which weeks carried which
price movements: the dispersion of $\hat\theta$ \emph{within} a realized design, relative to the
bootstrap's own standard error, is $0.51$ under the boosted learner and $0.41$ under
the sieve, both below one, so the bootstrap is, if anything, conservative \emph{conditional on
the design it was computed from} -- empirical coverage conditional on the design is $0.881$ for
the boosted learner and $0.475$ for the sieve. The same ratio computed against \emph{total}
dispersion, within designs and between them together, is $0.62$ for the boosted learner and
$2.04$ for the sieve: for the sieve, total dispersion is roughly double the bootstrap's own
width, consistent with a between-design component the resample cannot see; for the boosted
learner it is not, and the two learners are read separately rather than pooled into a single
verdict on this point. Table~\ref{tab:inference-variants}
reports the full eight-construction comparison this implies: the best of the eight, multiway
clustering, reaches $0.68$ under the boosted learner and $0.80$ under the sieve, still short of
the $0.95$ nominal level, and the rest range down to $0.20$, so no resampling or clustering
choice tested closes the gap. Section~\ref{subsec:hierarchical-coverage} takes up the remaining
question directly: whether an interval built on the missing between-design component, rather
than on resampling within one design, can. A full theoretical treatment of that construction
remains companion-paper material
(\citealp{delgado2026acrossdesignuncertaintyshortpricing}); the decomposition above is reported
here, rather than left entirely to that work, because it is what licenses the diagnosis of this
section rather than merely asserting it. What matters for the argument here is the consequence:
an interval computed from a single realized panel cannot always be certified as honest, and a
decision layer that ignored that possibility would sometimes authorize a move on evidence weaker
than it appears.
\subsection{What a single realized panel cannot tell you}\label{subsec:hierarchical-coverage}
Part of why a resampled interval can be too narrow is conceptual rather than computational.
Resampling weeks inside one realized panel can only characterize dispersion \emph{within} that
panel; it cannot characterize the dispersion that would arise from having observed a different
set of price movements altogether, what Section~\ref{subsec:unit-of-identification} treats as
\textbf{design variance}. A hierarchical variance component, estimated across several
presentations each of which is effectively a different realized design, can represent that
second source in a way a within-panel resample structurally cannot.
This system folds such a component into its uncertainty accounting through the
between-presentation heterogeneity $\hat\tau^2$ it already estimates for shrinkage
(Section~\ref{subsec:regularization-uncertainty}), built on a standard between-study
heterogeneity estimator \citep{paule1982consensus}. A full theoretical account of that
construction's coverage properties is companion-paper material; the operational point this
paper needs is narrower, and a direct comparison on this panel is reported below since it is
already computed: some of the
uncertainty this system faces is a property of which design was realized, not of how well it
is estimated, and no amount of resampling the same panel removes it.
\paragraph{A demonstrated comparison on this panel.} What adopting the hierarchical component
would change is not left unquantified, since it is already measured on a real replica of this
panel. Table~\ref{tab:hier-coverage} contrasts the production block-bootstrap against three
constructions built on $\hat\tau^2$, over $70$ replications on eight presentations, under a
scenario with heterogeneous truths and one in which every presentation shares the same truth,
which isolates design variance by construction. The posterior-predictive construction reaches
coverage of $0.914$ in the homogeneous scenario and $0.92$ in the heterogeneous one, against the
bootstrap's $0.484$ in both; the predictive interval before shrinkage is close behind, at
$0.902$ and $0.916$. Folding the between-presentation component back in closes most of the gap
the earlier decomposition predicted it would, which is direct evidence that the missing variance
is design variance and not an artefact of this panel. The posterior point estimate alone tells
the opposite story: shrinking the estimate toward the pooled mean without also propagating the
predictive uncertainty covers \emph{worse} than the bootstrap it would replace, at $0.407$ in the
homogeneous scenario and $0.42$ in the heterogeneous one, with a residual bias of $0.193$ in the
homogeneous scenario and $0.179$ in the
heterogeneous one, because shrinkage narrows the interval without correcting the centre it is
built around. The ratio of between- to within-presentation dispersion, $\hat\tau/\hat s$, runs
from $2.51$ to $2.78$ across scenarios, confirming directly what
Section~\ref{subsec:coverage-experiment} inferred indirectly: on this panel the bootstrap is
answering a question several times narrower than the one a decision needs answered. None of this
is adopted in the production configuration reported throughout this paper; it is reported here
because it is what the earlier diagnosis would look like acted on, and a full account of the
construction's own theoretical properties remains companion-paper material
(\citealp{delgado2026acrossdesignuncertaintyshortpricing}).
\subsection{The unit at which this channel identifies}\label{subsec:unit-of-identification}
If part of what defeats a per-presentation interval is design variance, it should average out
across units whose designs are independent, at a rate the arithmetic predicts, and this is the
paper's third contribution: showing when the unit of decision should move from the presentation
to the brand or the category. That prediction is testable, and testing it turns a decomposition
into a usable prescription for where a decision is legitimate.
Table~\ref{tab:identlevel} aggregates from the presentation to the brand and to the category
over 70 replications. Under the boosted learner the standard deviation of the estimate across
replications falls by a factor of $2.03$ at the brand, where four presentations are pooled, and
by $2.94$ at the category, where eight are; the factors predicted by independent design biases
averaging at the square root of the unit count are $2.00$ and $2.83$. Under the sieve learner
the measured factors are $2.02$ and $3.04$ against the same predictions, and the largest
relative deviation from the square-root law across the four contrasts is $7.4\%$. The mechanism
is therefore confirmed rather than argued.
Two consequences follow that a decomposition alone would not deliver. The first is a usable
estimate: at the category level under the sieve learner the root mean squared error is $0.159$
on an elasticity of magnitude $1.1$, against $0.571$ for a single presentation under the
boosted learner, a factor of $3.59$. Note which learner wins and why: the low-bias learner
looked worst when a single presentation was in view because its problem was variance, and
aggregation is precisely what removes variance. The second is a boundary on where this
purchases anything. Widening the panel in \emph{regions} does not average design bias, because
the list price is common across regions of the same presentation and the design is therefore
shared; widening it in \emph{presentations} does. This is the effective-unit-of-inference
principle applied to identification rather than to degrees of freedom.
Point-estimate accuracy is not the whole of what a decision needs, and aggregation is not a
free repair for the rest of it: whether an interval at the brand or category level also reaches
nominal coverage, and inside a usable width, is a harder and separate question that
Section~\ref{subsec:impossibility} takes up briefly and that companion work treats in full.
What this section establishes on its own is enough to act on: where item-level identification
is unavailable, the brand or the category is where this channel's evidence is strongest, and a
committee choosing where to publish a number should choose accordingly.
\subsection{The limit on per-presentation precision in this panel}\label{subsec:impossibility}
Aggregation improves the point estimate; it does not, by itself, guarantee an interval narrow
enough for a per-presentation decision. Across the combinations of aggregation level and
estimator examined on this panel, none delivers both nominal interval coverage and a width
inside the admissibility threshold defined in Section~\ref{subsec:data} at the presentation
level, and the picture improves only partially at the brand and category levels. Concretely,
the narrowest interval that also covers at the nominal level, the hierarchical predictive
construction of Section~\ref{subsec:hierarchical-coverage}, is never narrower than $1.47$ times
the admissibility width threshold even in the best combination tested, the category level under
a homogeneous scenario and the boosted learner, and runs up to $3.38$ times the threshold at the
presentation level under the sieve learner in the heterogeneous scenario; none of the twelve
level-learner-scenario combinations tested clears the threshold
(Appendix~\ref{app:extended-methodology}). That is a
property of how little independent price variation a hundred-and-twenty-week panel of this
shape supplies, not a defect specific to this implementation.
It is also the operational justification for treating WAIT as a first-class output rather than
a residual category, the paper's first contribution: on some presentations, at some levels of
aggregation, the evidence-honest answer is that the data does not yet support a per-item
decision at the confidence a committee should require. Characterizing exactly where that limit
sits, how it scales with the panel, and what would close it, is beyond the scope of this paper
and is developed in companion methodological work \citep{delgado2026acrossdesignuncertaintyshortpricing}; Section~\ref{subsec:artifact} discusses what
a committee should do while the limit stands.
\subsection{Detectable effect, and how far the gate travels}\label{subsec:mde}
The diagnostic thresholds are cutoffs on observable variation, but the question they exist to
answer is about detectable effect. The generator injects a known magnitude through the same
portfolio mechanism used in Appendix~\ref{subsec:experimental-scale}, swept over six points from
$0.1$ to $1.0$ at $W=120$ and $35$ replications per point, and Figure~\ref{fig:mde} reports, for
six candidate learners, the injected magnitude at which the block-bootstrap interval first
excludes zero in $80\%$ of replications. The learners split sharply on where that point falls:
under the production-adjacent boosted-tree learner it is $0.9243$, close to the top of the range
swept, and under the alternative tree implementation the curve never reaches $80\%$ power inside
the range tested at all; the lasso, ridge and the sieve reach it at $0.5$ or below, and a radial
basis kernel at $0.8804$.
\begin{figure}[htbp]
\centering
\sbox\pandoc@box{\includegraphics[keepaspectratio,alt={Minimum detectable effect curve}]{fig_mde.png}}
\Gscale@div\@tempa{\textheight}{\dimexpr\ht\pandoc@box+\dp\pandoc@box\relax}
\Gscale@div\@tempb{\linewidth}{\wd\pandoc@box}
\ifdim\@tempb\p@<\@tempa\p@\let\@tempa\@tempb\fi
\ifdim\@tempa\p@<\p@\scalebox{\@tempa}{\usebox\pandoc@box}
\else\usebox{\pandoc@box}
\fi
\caption{Detection rate --- the share of replications whose block-bootstrap interval for
$\theta$ excludes zero --- as a function of injected magnitude, for six candidate learners, at
$W=120$ weeks and 35 replications per grid point. The dotted line marks $80\%$ power; each
learner's honestly simulated minimum detectable effect is the magnitude at which its curve first
crosses that line. The text compares this simulated frontier against the textbook formula built
from each learner's own standard error in the same run, which measures precision rather than
simulated detection and is shown there to be substantially optimistic for several
learners.}\label{fig:mde}
\end{figure}
A companion sweep asks the second half of this section's question directly: holding the
injected magnitude fixed and varying only how many price movements the portfolio contains, over
$(2,3,4,5,6,7,8,10,12)$ movements. Under both the boosted-tree and the sieve learner the
admissibility gate stays below the point where half of draws are admitted at six movements or
fewer, crosses that mark at seven, and settles at an admitted share of $0.55$ to $0.675$ for the
boosted-tree learner and $0.525$ to $0.675$ for the sieve by twelve movements; the mean bias
measured only among the admitted draws does not move in one direction as the count of movements
rises, which is itself informative -- the gate is filtering on whether estimation is possible at
all, not on which replications happen to be least biased. Beyond how far the gate travels, the
frontier itself needs correcting: comparing the honestly simulated detectable
effect against the textbook formula built on each learner's own published standard error, the
ratio between them is $2.335$ for the boosted learner, the largest of six learners compared,
against $1.738$ for a radial basis kernel, $1.183$ for the lasso, and close to parity, at
$1.057$ and $1.045$, for ridge and the sieve. The published frontier is optimistic by a factor
that depends strongly on which learner computed the standard error it is built on, from
essentially none for ridge and the sieve to more than double for the boosted tree; it is the
same standard error this paper shows to be too small relative to the dispersion
that matters. It is recorded as a quantity to be recomputed from realized dispersion, and the
correction changes a published figure rather than a running threshold, since the frontier is not
declared anywhere in the operating configuration today.
\subsection{Does the gate generalize beyond the configuration it was calibrated on?}\label{subsec:gate-generalization}
The abstention gate is
the system's contribution, so whether its thresholds encode a property of the problem or of the
generator they were set on is the most consequential question a validation can ask. Freezing the
thresholds at the values calibrated on configuration A and moving the data-generating
configuration answers it directly. Table~\ref{tab:gate-generalization} reports the calibration
configuration A, a low-magnitude variant B, a few-movements variant C and a mid-pass-through
variant D: under the boosted learner precision is $0.5$ in A, undefined in B because the gate
admits nothing at all, $0.0$ in C and $0.75$ in D; under the sieve learner it is $0.5$ in A,
$0.0$ in B, where the gate still admits an eighth of cells but none of them is usable, $0.5$ in
C and $0.6$ in D. The worst precision reached outside the calibration configuration, across both
learners and every configuration where the gate admits anything at all, is exactly $0.0$,
against a rule of $0.80$.
Three consequences are drawn rather than deferred. The near-perfect separation on configuration
A is an artefact of calibration and does not travel as a general claim; what
remains true is that on the configuration the thresholds were set for, the gate separates. The
recalibration this implies is not a matter of moving a cutoff but of publishing the frontier:
an admissibility rule stated as a single number on a single configuration cannot travel, and
the object that travels is the surface of precision over the configuration space. And the
direction of the failure is not uniform: a gate that admits nothing, as under the boosted
learner in B, wastes a panel without authorizing a bad decision, whereas a gate that admits
cells at zero precision, as both learners do somewhere in this sweep, authorizes moves it
should not. Read against the rest of the paper this is not a reason to discard the
gate but to stop describing it as solved and describe it as calibrated.
\subsection{The abstention funnel}\label{subsec:funnel}
The most informative quantity about a system whose contribution is abstention is where its
portfolio is lost. Table~\ref{tab:funnel} reports the funnel over 50 replications and
$300$ presentation evaluations. Its coverage is declared rather than implied: guards 5
through 8 vote on the verdict but the aggregator records no per-guard exit count for them, so
four of the
ten abstention causes of Definition~\ref{def:extended-classification} appear in the table as
\emph{(pending)} and their contribution sits inside the ACT row. Those cells are left empty on
purpose. The extended-taxonomy study of Section~\ref{subsec:extended-taxonomy} does measure
per-cause incidence, but on a different generator, a different estimator and a disjoint draw,
so transplanting its figures into this funnel would manufacture a coherence the two runs do
not have.
The funnel ends at an ACT rate of $0.307$, and the most informative property of that number is
what it is nearly equal to. Two of six presentations are drawn from the \emph{clean} regime, so
$0.333$ of the portfolio is identifiable by construction; the variation diagnostic passes
$0.33$, the admissibility gate $0.307$, and the instrumented guards remove no further share at
this precision. By regime, the clean presentations reach ACT in $0.92$ of evaluations and every
other regime in $0.000$. The dominant recorded exit reasons are an unreliable $\theta$ at
$0.429$, the broken-weeks regime gap at $0.252$, the robustness value at $0.21$ and the
instability of $P^*$ at $0.078$; reasons are recorded per cell and a cell can carry several, so
these are incidences and not a partition.
\paragraph{What the guards cost on a panel that deserves an answer.} A funnel measured on a
portfolio that is two-thirds non-identifiable cannot distinguish discipline from paralysis, so
the run adds the complementary measurement. On a panel identifiable by construction throughout,
every veto is by construction a false positive. On a clean-regime panel with $\rho = 0.85$, ten
movements of $3$ to $7.5\%$ and an exogenous competitor, over $19$ replications, the
family-wise false-veto rate of the whole harness is $0.053$ (one veto out of $19$). The
admissibility gate is
held inert on this panel by construction, so the rate is produced entirely by the eight guards.
The diagnostic classifies the triggering guard: at this sample size a single cause accounts for
the entire rate, the regime-shift guard (Guard~1a, $\theta$ shifting when broken weeks are
excluded), at $0.053$; no other guard contributes. The median estimate on that panel is
$-1.024$. Abstention therefore has a bounded cost: the
harness does not veto identifiable panels, and the $0.307$ of the funnel is the portfolio and
not the guards.
\paragraph{The unit of inference in the robustness value.} Consuming panel rows in place of
independent weeks inflates the degrees of freedom by a factor of $6.0$, exactly the number of
regions, which confirms the inflation is the mechanical one the principle predicts. Correcting
it raises the mean robustness value from $0.168$ to $0.317$ and lowers the veto rate from
$0.575$ to $0.366$; adding a clustered standard error alongside the corrected degrees of
freedom, at a clustering factor of $1.14$, returns $0.402$ and a veto rate of $0.21$. All
three are published, since reporting the middle variant alone would understate the statistic
and overstate the veto rate, and choosing between them is a calibration decision for a real
category rather than a result. The correction widens the separation between regimes rather than
blurring it: in \emph{clean} the robustness value moves from $0.392$ to $0.672$.
The funnel is the quantity a committee should be shown first, because it converts an abstract
property of the design into an operational expectation. A committee that learns from a table
before the first cycle that a substantial share of a real portfolio returns WAIT reads it as
discipline; one that discovers it from an empty price list reads it as malfunction.
\subsection{Does the taxonomy close the classification gap?}\label{subsec:extended-taxonomy}
A taxonomy that partitions ten causes on paper is worth nothing if the population it is applied
to never exercises the cells it adds. This subsection measures the property directly: how much
of a blocked portfolio the classification rule actually reaches, under the rule as it stood
before Definition~\ref{def:extended-classification} and under the definition itself.
The measurement is deliberately separate from the funnel above. It runs $50$ independent
realizations of the decision mechanism under the gradient-boosting nuisance learner on synthetic
panels disjoint from the funnel's draws, for $300$ decision instances in total. Of those, $250$
return a WAIT of some kind, an $83.3\%$ block rate against an ACT rate of $16.7\%$. Two
populations, two estimators, two block rates: the numbers below are a measurement of the
classification rule and not a recomputation of Table~\ref{tab:funnel}.
Applied to those $250$ blocked instances, the informal five-cause assignment the framework
carried before this paper resolves $97.2\%$ of them. Definition~\ref{def:extended-classification}
resolves $100\%$. The difference is $7$ instances, $2.8\%$ of the blocked portfolio, whose only
binding causes were among the five the earlier rule never named: Guard~1's coverage condition,
Guard~3, Guard~4 and Guard~5. Under the earlier rule those seven presentations returned a WAIT
with no attached remedy, which is the operational definition of an unclassified terminal state:
the committee is told to hold the price and told nothing about what would change that.
\begin{figure}[htbp]
\centering
\includegraphics[width=0.7\linewidth]{fig_task2_decision_resolver.png}
\caption{Share of blocked presentation-instances resolved to a verdict type under the
five-cause assignment and under Definition~\ref{def:extended-classification}'s ten-cause
partition, by verdict type, over $50$ replications and $250$ blocked instances.}
\label{fig:task2-decision-resolver}
\end{figure}
Across the ten causes on this population, admissibility binds most often at $68.7\%$ of the
$300$ instances, followed by Guard~6 at $64.3\%$; Guard~1a and Guard~7 come next, at $43.7\%$
and $43.0\%$, with Guard~2 close behind at $35.7\%$. Among the five causes
the definition newly assigns, Guard~1a regime sensitivity is the most frequent at $43.7\%$.
Guard~5 never binds, at $0.0\%$. Contamination-type causes dominate the blocked set overall,
$96.0\%$ against $4.0\%$ identification-type, which is a property of this population's guard
incidence rather than of the classification rule; a portfolio with cleaner channels and thinner
panels would invert it.
Two readings of that last figure deserve to be separated. The taxonomy closes the classification
gap completely on this population, without altering a single guard's threshold or binding rate,
which is what it was written to do. It does not follow that most blocked presentations are
experimentable here: on this draw only $4.0\%$ of the blocked set is identification-type and
therefore eligible for the experimental route. The classification is exhaustive; the eligibility
it produces is an empirical quantity that varies with the portfolio, and this paper reports it
on one population rather than claiming it as general.
\section{Robustness}\label{subsec:ablations}
This section asks a different question than Section~\ref{sec:results}: not whether the
system's estimates and decisions are accurate on the calibration generators, but whether that
accuracy depends on specific implementation choices -- which learner, which cross-fitting
scheme, which point estimator, which filter -- rather than on the identification strategy
itself. Table~\ref{tab:ablations} reports eleven such contrasts, each substituting one
alternative for one production choice and re-running the same generators used in
Section~\ref{sec:results}. These are stress tests of design decisions, not the pre-registered
pass/fail rules of Section~\ref{subsec:validation-design}; a contrast can produce an
informative null, a reversal, or a retirement without any rule failing.
Each contrast in Table~\ref{tab:ablations} removes or replaces one design decision and measures
what that costs against the same generators. The block is net unfavourable to the production
configuration, which is the strongest available evidence that it was run to measure and not to
confirm: two contrasts return nulls, three return findings against the production
configuration, two retire a component, and one returns opposite verdicts on the two generators.
Appendix~\ref{app:extended-results} reports each in full.
\begin{table}[H]
\centering\small
\caption{The ablation block. Each row removes or replaces one design decision. A null is
reported as the outcome rather than as a step toward a positive one, and the two components
that fail their own test are withdrawn from the published output rather than
retained.}\label{tab:ablations}
\begin{tabular}{@{}p{0.24\linewidth}p{0.28\linewidth}p{0.40\linewidth}@{}}
\toprule
Contrast & Alternative tested & Outcome \\
\midrule
Outcome variable (\S\ref{subsec:sellin-sellout})
& Sell-in against sell-out under injected forward buying
& \textbf{Null.} Gap $-0.088$, and of the opposite sign to the mechanism. Sell-out keeps its argumentative support and loses its quantitative one \\
Structural-break guard (\S\ref{subsec:guard5-power})
& Asymptotic against simulated critical value, four residual paths
& \textbf{Guard withdrawn.} Union of both tests rejects at $0.94$ under the asymptotic critical; calibration restores size to between $0$ and $0.107$ on every path but leaves power at $\le 0.161$ at the largest break, $0.286$ inside the recency window against $0.214$ outside it \\
Nuisance learner (\S\ref{subsec:learner-experiment})
& Boosted trees against a regularized trend-harmonic learner, six learners in total
& \textbf{Disagreement across generators.} Dominated on the five-regime panel; on the extended one the best rival is not distinguishable from the production learner ($t = -0.30$). Declared a trade-off; the learner with the lowest error is the one that breaks the break test \\
Cross-fitting scheme (\S\ref{subsec:crossfit-experiment})
& Contiguous blocks against forward chaining, under an injected break
& \textbf{Null.} Difference $+0.063$; both recover a blend of two regimes, so the tolerance is not what binds \\
Point estimator (\S\ref{subsec:coverage-experiment})
& Median across folds against pooled DML2
& \textbf{Learner-dependent, neither significant.} Paired $t = -1.43$ for the boosted learner, $t = 1.99$ for the production (sieve) learner; both short of a conventional threshold, so the choice is declared a convention \\
Sign filter (\S\ref{subsec:sign-filter})
& Significant-and-negative against an a priori partial-$R^2$ criterion
& \textbf{Filter replaced.} It keeps $0.20$ of estimates and displaces the published mean by $-1.293$; the a priori rule keeps $0.96$ and displaces by $0.022$, revealing an attenuation the filter concealed \\
Dual Beta (\S\ref{subsec:dual-beta-result})
& Asymmetric against symmetric response, deflated and nominal
& \textbf{Estimator retired.} Size $0.286$ against power $0.321$ on the nominal specification, $0.321$ against $0.357$ on the deflated one; a noise floor of about $0.25$ on the gap, larger than any asymmetry a committee would act on \\
Pass-through variance (\S\ref{subsec:nulls-retested})
& Regional and state-dependent $\rho$ injected to break the null
& \textbf{Null confirmed.} Pass-through contributes $\le 0.59\%$ of the response band's variance even with heterogeneity and state-dependence injected; the estimator itself recovers at a bias of $-0.007$ to $-0.048$ \\
Cash-pressure screen (\S\ref{subsec:nulls-retested})
& Order refusal injected to give the dependent variable variation
& \textbf{Screen vindicated.} The non-finite $t$ was a constant regressand, not a defect; with refusal injected the screen operates, and declines correctly where there is no signal \\
Demand system (\S\ref{subsec:aids-inference})
& Bootstrap inference on the restrictions
& \textbf{Restrictions rejected.} Adding-up holds by construction; homogeneity and symmetry are each rejected at $5\%$ in $0.625$ of replications. The matrix is published with the rejections attached \\
Experimental layer (\S\ref{subsec:experimental-scale})
& Six, twelve, twenty and thirty regions
& \textbf{Calibrated with scale.} Attainable $p$-floor $0.067 \to 0.002$; pre-fit RMSE $0.189 \to 0.133$; donors carrying non-zero weight $3.9 \to 12.1$; false-positive rate $0.0 \to 0.06$, not monotonic in between \\
\bottomrule
\end{tabular}
\end{table}
Two of these rows deserve to be read together, because the pair sharpens the paper's organizing
claim rather than merely qualifying it. The nuisance-learner contrast reverses between two
generators that differ only in how pass-through is drawn, and no learner wins on all three of
bias, coverage and root mean squared error: a radial basis kernel has both the lowest bias and
the best coverage, while the boosted tree has the lowest error. A system whose value rested on
the estimator would therefore have no defensible way to choose one. What is stable is the discipline of measuring, and it is
the measurement, not the learner, that transfers.
The last row is a validity check on the system's own experimental-comparison machinery rather
than on the pricing estimator. The experimental layer (Appendix~\ref{subsubsec:experimental-layer})
supplies an independent way to confirm the effect of an already-executed price change through a
holdout and synthetic-control contrast, and Table~\ref{tab:sdid} shows that its resolution is
sensitive to how many independently assigned units back it: the same construction that is
uninformative at six regions is usable at thirty. A committee planning a holdout of this kind
should size it against that table rather than assume six regions behave like thirty; the general
relationship between panel width and attainable resolution is developed in companion
methodological work.
\section{Discussion}\label{sec:discussion}
\subsection{Seven tacit assumptions, seven executable tests}\label{subsec:tacit}
The strongest claim this work supports is not the integration of causal econometrics into a
pricing workflow, which is by now a well-populated genre \citep{chernozhukov2018double}, but a
pattern that emerged from auditing a carefully built system and then executing it. The failure
modes that survive a competent build are rarely errors of estimation. They are unstated
assumptions about the \emph{environment}, and each can be converted into a quantity that costs
less to measure than its consequence costs to absorb. Table~\ref{tab:tacit} states the seven
found here: four by reading the design, and three by running the validation.
\begin{table}[H]
\centering\small
\caption{Seven assumptions that a careful build leaves tacit, and the executable contrast each
becomes. The upper block was found by auditing the design and the lower by executing the
validation. In every row the test is cheaper than the consequence of leaving the assumption
implicit, which is the property that makes the pattern worth generalizing past
pricing.}\label{tab:tacit}
\begin{tabular}{@{}p{0.26\linewidth}p{0.32\linewidth}p{0.34\linewidth}@{}}
\toprule
Tacit assumption & Executable test & Consequence if left implicit \\
\midrule
\multicolumn{3}{@{}l}{\emph{Found by auditing the design}}\\[2pt]
The controls are pre-treatment
& Declared temporal status per column; estimation aborts on an undeclared one; mediator bracket (Guard 6)
& Conditioning on a mediator inflates $|\theta|$; the engine under-recommends increases and the number stays plausible \\
The instrument acts through one channel
& Exclusion audit: residual correlation of instrument and competitor price against a set carrying an explicit trend (Guard 7)
& A common cost shock moves both estimators together, so the Hausman-type contrast is small \emph{because} the restriction fails \\
The comparison group is inert
& Contamination test on the control group's price shift, against a permutation distribution of false cuts and an economic floor
& A competitor reading regional prices contaminates the donor pool; the measured effect is attributed to the intervention \\
No other optimizer acts on the same unit
& Quota-pressure bracket with a declared source hierarchy (Guard 8)
& End-of-period effort lowers effective price and raises volume together; the engine learns its own firm's commercial policy as demand \\[4pt]
\multicolumn{3}{@{}l}{\emph{Found by executing the validation}}\\[2pt]
A computed interval is an honest interval
& Coverage against a known truth; conditional-versus-unconditional dispersion across designs
& The bootstrap targets dispersion conditional on the realized design, which is not what a guard certifies; three governance switches appeared to operate and did not \\
The unit of decision is the unit of identification
& Aggregate the estimate to brand and category and measure the dispersion across replications
& Design bias averages at the square-root rate, so a category estimate reaches an RMSE of $0.159$ where a presentation estimate reaches $0.571$; publishing per presentation publishes at a grain the evidence does not support \\
A threshold calibrated on one configuration is a threshold
& Freeze the cutoffs and sweep the data-generating configuration
& Gate precision falls from near-perfect on the calibration configuration to $0.000$ once evaluated outside it \\
\bottomrule
\end{tabular}
\end{table}
The symmetry of the table is the argument. Each row names an assumption that nothing in the
estimation stage can detect, because each is a statement about the environment rather than about
the data; each is answered by a contrast the system computes and publishes; and in each case the
test is cheap while the consequence is not. That the last three were found inside the system's
own uncertainty machinery, its choice of unit, and its own thresholds is what makes the pattern
worth stating as a general one. The sixth row deserves one further remark against this paper's
own account: that the same work argues at length for estimating at the week rather than the
month, because the grain of a panel is an identification question, and then publishes per
presentation without asking the same question of the cross-section, is the clearest instance
here of a principle applied on one axis and not on the other. It is also the part of this work
with no close equivalent in the
enterprise deployments taken here as the comparison set \citep{hormby2010marriott,
deng2023alibaba, llenas2026pepsico}, which report what the deployed system achieved rather than
which of its identifying assumptions were checked; the observation is scoped to those three and
is not offered as a survey finding.
Provenance differs across the seven, and stating it plainly matters for judging what this audit
adds against what it borrows. Three operationalize a concern already established in the
causal-inference literature: the exclusion audit (Guard~7) turns the standard instrumental-variable
exogeneity assumption into a residual-correlation contrast; the contamination test adapts
permutation-based inference on the no-interference half of SUTVA \citep{rubin1980randomization}
to a donor-pool interference question; and the threshold-generalization check (row seven) adapts
ordinary out-of-configuration validation to a decision cutoff rather than to a fitted model. The
coverage diagnostic of row five reuses established conformal machinery
\citep{gibbs2021adaptive, barber2023conformal} for the test itself; what that test revealed about
the source of the shortfall is taken up in Section~\ref{subsec:artifact} and developed formally
in companion methodological work \citep{delgado2026acrossdesignuncertaintyshortpricing}. Two rows have no close equivalent in the enterprise-pricing or
causal-inference literatures this paper compares itself against: the quota-pressure bracket
(Guard~8), because interference between independently operated optimizers acting on the same
unit is not a standard identification concern, and the aggregation test of row six, which this
paper develops as its own contribution (Section~\ref{subsec:unit-of-identification}) rather than
importing one. The mediator bracket (Guard~6) sits between the two groups: the underlying
bad-control concern is well known, but declaring temporal status per column and aborting
estimation on an undeclared one is this audit's own enforcement mechanism rather than a borrowed
diagnostic.
\subsection{Where the value of a system like this one lies}\label{subsec:contribution}
In the modern-trade deployments this literature documents, where promotional variation is
abundant and shelf prices are controlled, the reported value rests in the optimization layer
navigating a combinatorial space beyond manual planning. In intermediated retail the binding
constraint shifts: the decision space for a single presentation is small and one-dimensional,
the input slope is the fragile link, and optimization is the \emph{simplest} component of the
system.
The run sharpens that ordering and organizes what follows around the paper's three
contributions, each stated as a question in Section~\ref{sec:intro}. Value does not concentrate
in the estimator, since which of two nuisance learners is better reverses between two generators
that differ only in how pass-through is drawn. What survives is narrower and more durable: the
value lies in \emph{measuring which of the system's own claims hold}, and in publishing the
measurement. That discipline is what the three contributions below operationalize, each in a
different part of the system.
\textbf{The ACT/WAIT framework (RQ1, Section~\ref{subsec:decision-layer}).} Causal identification
isolates price effects from promotion, competitor action and co-trending inflation, and the
recovery experiments establish where it does and does not. The conformal layer is the one
uncertainty component whose measured behaviour matches its claim. Neither is this paper's
contribution on its own; DML, the instrumental-variable contrast and conformal calibration are
supporting tools. The contribution is the gate that reads them together and abstains rather than
publishes a number the evidence does not support, at a family-wise false-veto cost estimated at
$0.053$ on an
identifiable panel ($N=19$; \S\ref{subsec:funnel}) -- the difference between publishing several different wrong answers and
publishing none. Its specific thresholds are calibrated rather than general
(Section~\ref{subsec:gate-generalization}), which is stated as a limitation rather than
qualified away. A WAIT verdict is not a dead end, either: which guard produced it indicates
whether the shortfall is one of identification, calling for independently assigned variation, or
an operational threat that any experiment supplying that variation would have to control for
first (Section~\ref{subsec:artifact}).
\textbf{The admissibility audit (RQ2, Section~\ref{subsec:tacit}).} An engine correct in
isolation yet blind to an adjacent optimizer acting on the same unit will model its own firm's
commercial policy and mistake it for demand; conditioning on a mediator inflates the recovered
effect; an instrument's exclusion restriction can fail silently. Section~\ref{subsec:tacit}
converts seven such assumptions into diagnostics the system computes and publishes before a
recommendation is released, and states which formalize a concern the causal-inference literature
already recognizes and which do not have a close equivalent there.
\textbf{Legitimate aggregation (RQ3, Section~\ref{subsec:unit-of-identification}).} A component
should report the unit at which its evidence identifies rather than the unit at which the
business happens to decide. Design bias averages at the square-root rate across replications, so
a category estimate reaches an RMSE of $0.159$ where a presentation estimate reaches $0.571$;
Section~\ref{subsec:unit-of-identification} supplies the procedure -- aggregate, then measure the
dispersion the aggregation buys -- for finding the level at which a decision is legitimate rather
than assuming it is the level at which the business happens to decide.
The three are complementary rather than substitutable. Causal elasticity without an admissibility
gate is descriptive analytics that cannot say when to trust itself. A gate without the
aggregation result enforces a precision standard at a unit the evidence cannot reach. Either
without the tacit-assumption audit is blind to failure modes that are properties of the
environment rather than of the estimator -- and, as the previous section notes, some of those
failure modes have no close equivalent in the enterprise deployment literature taken here as the
comparison set.
\subsection{Governance, adoption and the recalibration this run forces}\label{subsec:artifact}
The artifacts feed a human-in-the-loop process, not an autopilot: a periodic trigger, the data
system, the recommendation layer with its corridor, commercial validation with a documented
override, a committee with minutes linked to the analysis, and execution with ex-post
measurement. Governance assigns explicit roles and typifies valid overrides (strategic,
commercial, model-based, always documented) against the single invalid one, that this is how it
has always been done. To these is added the cross-system convention under which each alternative
declares its expected cost on adjacent decision systems, so conflicts between models are
adjudicated on the same sheet as the price.
Three implementation challenges dominated, and the third has changed. The first was the internal
coherence of the resampling policy, since one component that shuffles rows suffices for optimism
to leak into everything downstream. The second was the discipline of declaring estimands, since
the difference between a competitor-conditional elasticity and an equilibrium response is
invisible in a report that shows only ``the elasticity''. The third is organizational, and the
conversation a committee must be prepared for is harder than ``the system will sometimes decline
to recommend.'' It is instead ``the system
cannot, on the evidence this channel supplies, publish a per-presentation elasticity with an
interval narrow enough to act on, and here is the level at which it can.'' A system that says the
first earns trust; a system that says the second earns it and changes the operating unit, which
is a larger organizational ask and the honest one.
One governance risk is created by the system's own discipline and should be anticipated rather
than discovered. An engine designed to halt rather than degrade returns WAIT for a substantial
share of a real portfolio, which reads to a committee expecting a price list as a malfunction,
and the pressure that follows is predictable: relax a threshold for one category, waive a guard
for a strategic presentation, publish the midpoint because a number is required. The
countermeasure is not technical. It is to publish an explicit degraded mode (no recommendation,
reason, date of next review) rather than silence, and to accompany every abstention with the
decision arithmetic, so that holding is legible as discipline rather than as incapacity.
The run adds a second and more consequential governance item. Two calibrations, the admissibility
width threshold of $0.6$ and the construction of the interval it reads, were set independently,
and Section~\ref{subsec:impossibility} reports that the interval this implementation currently
produces does not reliably clear that threshold at the presentation level. A committee cannot be
asked to adjudicate that as two switches. It is one decision with three defensible resolutions,
each of which should be presented costed: publish at a coarser unit, where the interval is
narrowest; publish per presentation with an interval that does not cover, and say so on the
artifact; or supplement the observational panel with experimental variation, a path
Section~\ref{subsec:ablations} takes up empirically and companion methodological work \citep{delgado2026acrossdesignuncertaintyshortpricing} develops
further. The first is available now and the second is what the system does today without
declaring it.
A WAIT verdict is itself diagnostic, and which guard produced it indicates what kind of remedy
is called for and not only that one is needed. Two failure modes are structurally different.
When the admissibility gate fails for want of price variation, or Guard~2's robustness value
flags that a plausible unmeasured confounder could overturn the estimate, the shortfall is one
of identification: the observational record does not contain variation an omitted-variable
story cannot also explain, and only variation assigned independently of that story resolves it.
When the trigger is instead Guard~8's quota-pressure bracket or the contamination test of
Appendix~\ref{subsubsec:experimental-layer}, the threat is operational rather than statistical: a
sales team retaining discretion over incentives during a test window, or a competitor able to
observe and react within the assignment's own geography, would undermine an experiment as
readily as it undermines the observational estimate. A finding of the second kind is therefore
not evidence against experimenting; it is a requirement on how the experiment must be run --
incentive discretion frozen or centralized for the test's duration, and treatment and control
drawn at a geographic scale wide enough that a competitor's local reaction cannot cross between
them. Sizing that scale, and the level of aggregation at which such a design becomes
organizationally realistic, is beyond this paper's scope and is taken up in companion
methodological work \citep{delgado2026acrossdesignuncertaintyshortpricing}.
The ex-post loop remains the most valuable asset. The residual uncertainty this validation
isolates behaves like dispersion across designs rather than sampling noise within one -- a
distinction developed formally in companion methodological work \citep{delgado2026acrossdesignuncertaintyshortpricing} -- and if that description is
right, better estimation of the same panel narrows it only partially, while generating new
designs narrows it further. For the pilot that would follow, success criteria are stated in
advance rather than reported as results, staged across process, decision, model and business; the
business target of a $2$--$5\%$ revenue increase is recorded as the sponsoring organization's
commercial ambition and not as a forecast the evidence supports, since the classic evidence
establishes only that price carries high leverage on profit \citep{marn1992managing} and the
documented deployments only that analytics-enabled pricing has delivered material gains at scale
\citep{deng2023alibaba, llenas2026pepsico}. One criterion is added by the run: rather than assume
a holdout of any size is adequate, size it against the experimental layer's own measured
resolution, for which Table~\ref{tab:sdid} gives the committee a concrete range to plan against,
between twelve and thirty independently assigned units in this panel.
Appendix~\ref{subsec:implementation} develops the governance design and the roadmap.
\subsection{Declared limitations}\label{subsec:limitations}
The system is deliberately explicit about what it does not identify. The cross elasticity
inherits the endogeneity of the competitor's price, for which no instrument exists with these
data; the competitive scenarios are reduced-form, one round, with a constant reaction; and the
performativity of the engine has no econometric fix, only an experimental one.
Two limitations attach to the classification of Definition~\ref{def:extended-classification}
specifically, and both are properties of the evidence rather than of the rule. First, the
\textsc{transient} category is empirically empty in everything this paper measures: Guard~5 is
reported as not operative in the validated configuration (Section~\ref{subsec:guard5-power})
and binds in $0.0\%$ of the $300$ instances of the extended-taxonomy study
(Section~\ref{subsec:extended-taxonomy}). The category is
defended on the argument that a recent structural break is neither a precision shortfall nor a
validity violation and that assigning it to either would be a category error; that argument
does not depend on the cell being populated, but the cell is not populated, and a reader is
entitled to treat the third branch as untested. Second, the funnel of
Section~\ref{subsec:funnel} reports per-cause incidence for six of the ten causes and marks
four as pending, because this run's aggregator does not
record separate exits for Guards~5 through~8. A complete per-cause funnel on the same
population as the main run is the obvious next measurement and is not in this paper.
Six further limitations were produced by the validation run itself, and each qualifies a claim
made elsewhere in this paper.
\begin{enumerate}
\setlength{\itemsep}{0pt}\setlength{\parskip}{0pt}
\item \textbf{The interval this implementation currently produces falls short of nominal
coverage at the presentation level, and the shortfall is one of centring rather than width.}
An alternative construction was identified that reaches a nominal-appropriate coverage rate;
it is \emph{not yet adopted}, and every figure in this paper is produced under the
construction the implementation currently runs. The comparison between the two constructions
is developed in companion methodological work \citep{delgado2026acrossdesignuncertaintyshortpricing}.
\item \textbf{No aggregation level tried both covers and clears the admissibility gate at its
current threshold.} Any statement here about the calibration of the three width-dependent
switches is conditional on a threshold that has not yet been reconciled with an interval
construction that covers; Section~\ref{subsec:unit-of-identification} takes up the
reconciliation at the level of aggregation, and companion methodological work \citep{delgado2026acrossdesignuncertaintyshortpricing} takes it up at
the level of interval construction.
\item \textbf{Guard~5 is not operative.} With the asymptotic critical value the break tests
reject at $0.94$ under no break; with a calibrated critical the size returns to range on
every residual path, but power at the largest break injected still does not exceed $0.161$.
Both the recency veto and the
sequential reset that depend on detection are withdrawn, and the currency of the elasticity
rests on the drift alert and the revalidation cadence alone.
\item \textbf{The gate is calibrated, not general.} Precision is $0.000$ once the frozen
thresholds are evaluated on a configuration outside the one they were calibrated on, against
a rule of $\ge 0.80$ (\S\ref{subsec:gate-generalization}).
\item \textbf{The Dual Beta is withdrawn.} Size $0.286$ against power $0.321$ on the nominal
specification, $0.321$ against $0.357$ on the deflated one; the lookup table
loses its per-direction band.
\item \textbf{The conformal guarantee is calibrated within the historical regime.} A band can
be narrow and wrong if the regime shifts, and the system's defense against that lived in the
structural-break guard, which this run reports as not operative. That exclusion is therefore
now open rather than delegated.
\end{enumerate}
Two further items are recorded because they run against
the system rather than for it: the claim that a boosted nuisance learner is dominated by a
regularized one does not hold on the extended generator, where no rival dominates the production
learner at a conventional threshold; and the round-point failure, the sharpest illustration this
literature typically reaches for, does not reproduce on the experiment built to demonstrate it.
One further item belongs with these because it is a property of the run rather than of the
design. The cross-fitting partition this run uses is the unembargoed one
(Section~\ref{subsec:causal-identification}), so a single week can appear on both sides of a
fold boundary through a different region. The embargoed variant is implemented and its coverage
measured; it was not adopted here because it would change the residualizer that two of the
ablations compare. Every figure in this paper is therefore produced under a partition that
carries a declared and unquantified leak.
\paragraph{Limitations of scope.} Three are inherited from the validation design and bound every
number here. The system has not been run on a commercial panel: it is complete and a
single configuration switch redirects it, but the switch has not been thrown, and no result
reported here is an observation of a category. Both generators are stationary apart from an
inflationary drift and inject their confounding at magnitudes the designer chose, so agreement
between them is evidence of robustness to configuration and not of external validity. And
promotional incrementality is outside the system's scope entirely: the engine treats promotion as
a confounder to be removed, not as a lever to be optimized. Layer-specific limitations, including
the likely violation of the exclusion restriction and the inter-area agreement the quota layer
awaits, are recorded in Appendix~\ref{app:thresholds} together with every governance threshold,
all of which must be recalibrated against a real category before being fixed. None of these is
hidden in the artifacts: they are declared, because a decision system that feigns certainty where
it has none is more dangerous than none at all.
\section{Conclusion}\label{sec:conclusion}
This paper set out to ask, of an operationally complete causal pricing system, not what the
elasticity is but when there is enough evidence to act on it and when honest conduct is to wait.
The system was implemented end to end and validated against two data-generating processes with
known ground truth under thirty pre-registered rules. Eleven pass, eighteen fail, and one does
not apply. That the
failures outnumber the passes is the result rather than a caveat on it: a validation designed to
confirm returns confirmations, and the value of this one lies in the places where the system's
own claims did not survive the test written for them. These thirty rules are engineering
verification internal to this run; the paper's scientific claims are the three contributions
below, and a high fail count on the former says nothing against any of the latter.
Section~\ref{subsec:contribution} states the three contributions in full; briefly, each answers
one question posed in Section~\ref{sec:intro}. The ACT/WAIT framework
(Section~\ref{subsec:decision-layer}) answers \emph{when one can act}: a gate that reads causal
identification, conformal calibration and the tacit-assumption audit together and abstains rather
than publish a number the evidence does not support, at a false-veto cost estimated at $0.053$ on an
identifiable panel ($N=19$; \S\ref{subsec:funnel}), with thresholds this paper reports as calibrated on the configurations tested
rather than shown to be general -- and whose abstentions are diagnostic rather than terminal.
Definition~\ref{def:extended-classification} makes that last property exact: every one of the
ten abstention causes maps to exactly one of three verdict types, so a WAIT always carries the
reason it would stop being one. A \textsc{wait-contamination} calls for control over the
offending channel, a \textsc{wait-transient} for accumulated history under the new regime, and
a \textsc{wait-identification} for independently assigned price variation, which makes that
third verdict a formal statement of experimental eligibility rather than a shrug
(Section~\ref{subsec:artifact}). What this paper does not do is design that experiment. It
establishes the epistemic necessity of one, by showing that a resampled interval computed on
observational variation is blind to the variance a deliberate design would control, and it
marks the presentations for which the necessity binds; the optimization problem of acquiring
that variation across a portfolio at least operational cost is a separate contribution and is
left to companion work. The
admissibility audit (Section~\ref{subsec:tacit}) answers
\emph{what has to hold for that gate to be trusted}: seven assumptions a competent build otherwise
leaves tacit, converted into diagnostics the system computes and publishes before it recommends,
some adapting a concern the causal-inference literature already recognizes and some original to
this audit. And the aggregation result (Section~\ref{subsec:unit-of-identification}) answers
\emph{at what level a decision is legitimate}: a category-level estimate reaches a root mean
squared error of $0.159$ against $0.571$ at the presentation level, and the paper supplies the
procedure -- aggregate, then measure what the aggregation buys -- for finding that level rather
than assuming it is the level at which the business happens to decide.
The second and third contributions connect through a single observation. A resample inside one
realized panel cannot represent dispersion that comes from the panel being the particular one it
is; in this run that dispersion behaves like variation \emph{across} designs rather than noise
within one, a distinction Section~\ref{subsec:artifact} motivates conceptually and companion
methodological work develops formally. If that description holds, better estimation of the same
panel would narrow it only partially, which is why aggregation is often the difference between a
decision the evidence supports and one it does not, on the kind of thin, infrequently-moving panel
a intermediated retail channel typically supplies (Section~\ref{subsec:limitations}).
This paper does not resolve that limitation, and does not claim to. What it delivers instead are
the three contributions above, and a system that would rather report a coarser unit, an interval
that declares what it does and does not cover, and an abstention whose cost has been measured,
than publish a number it cannot support. The lesson this case adds to the literature on enterprise pricing
systems is that, in hard-identification environments such as the intermediated retail channel of
emerging markets, the greatest returns lie not in more powerful optimizers but in more honest
inputs, and that honesty has a measurable price which a system should be built to report rather
than to obscure. A second lesson generalizes further: the failures that survive a careful build
are rarely errors of estimation, but unstated assumptions about the environment, each cheaper to
test than to leave implicit. Ranges rather than points, abstention rather than false precision,
and declared assumptions rather than convenient ones are what let a system of this kind be
adopted, audited, and trusted to say when it does not yet know.
\section*{Declarations}\label{sec:declarations}
\addcontentsline{toc}{section}{Declarations}
\paragraph{Data availability.} The results reported in this paper are produced by a synthetic data-generating process that contains no commercial information; it is described in Appendix~\ref{subsubsec:dgp} and published in full with the implementation. The production panel that the same system is designed to consume derives from proprietary commercial sources and is subject to confidentiality restrictions; it is not publicly available, and it was not used to produce any result reported here.
\paragraph{Code availability.} The complete system is published as an executable reference implementation, configured by default against the synthetic generator, so that every figure, table, and threshold in this paper can be reproduced end to end without access to any commercial data. Reproducibility is a property of the implementation rather than a claim about it: the global seed is fixed, the cross-fitting folds are deterministic functions of the temporal ordering, the conformal split is chronological, holdout assignment is a deterministic hash of the cycle identifier, and the sequential state of the pooling layer is persisted and versioned per cycle. Redirecting the system to a production panel is a single configuration change.
\paragraph{Funding.} This research received no financial support from any public, commercial, or non-profit organization. The work was conducted independently, without institutional funding or the use of proprietary data.
\paragraph{Conflicts of interest.} The author declares no competing interests. The research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. No proprietary or commercial data were used in this study.
\paragraph{Author contributions.} \textbf{Pedro Cadahia Delgado:} Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Data curation, Original draft preparation, Review and editing, Visualization.
\paragraph{Ethics approval.} This study involves no human or animal subjects.
It analyses aggregated commercial time series and a synthetic generator, and no
personal data are processed at any stage.
\clearpage