EconBase
← Back to paper

Across-Design Uncertainty in Short Pricing Panels: Evidence from Simulated Price Trajectories

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

59,783 characters

Across-Design Uncertainty in Short Pricing Panels: Evidence from Simulated Price Trajectories


\maketitle

\begin{abstract}
\noindent
Short observational pricing panels can contain many observations while offering only a small number of distinct price movements. This paper studies the inferential consequences of that distinction in a synthetic data-generating process calibrated to a sparse pricing regime. We separate uncertainty conditional on a realised price trajectory from variation in estimation error across alternative trajectories generated by the same pricing process. In the baseline simulations, the latter component accounts for $97.6\%$ of the variance of the estimation error for the gradient-boosted specification. This magnitude is a property of the simulated design, not a general constant. Within-panel resampling and cluster-robust procedures use the information contained in one realised trajectory and therefore do not, without additional modelling or independently realised trajectories, identify this across-design component. The simulations also indicate that the coverage shortfall is associated with design-specific centring error, not only with underestimated conditional variance.

Three results organise the analysis. First, over the simulated grid, across-design dispersion is well described by the empirical relation $\hat\sigma_b \approx 0.182 V^{-0.271}$, where $V=n_{\text{moves}}\times\text{magnitude}^2$. We treat the exponent $-0.271$ as a simulation regularity rather than as a theoretical scaling law. Second, adding regions that share a common price path can reduce outcome noise and improve nuisance estimation, but it does not create additional independent price trajectories. By contrast, averaging estimates across units with approximately independent design-specific errors reduces the standard deviation of that component at the familiar $\sqrt{k}$ rate under the stated covariance conditions. Third, a Paule--Mandel variance component estimated across independently priced units substantially increases empirical coverage in the homogeneous simulation, from $0.469$ to $0.931$ for the posterior-centred variance-augmented construction; when unit-level truths differ, the same component mixes design uncertainty with genuine heterogeneity and should be interpreted accordingly. The broader implication is a shift from asking only how to refine inference on a fixed passive panel to asking how the data-generating design can create additional identifying variation. Controlled regional randomisation is one transparent way to do so, but the numerical gains reported here remain simulation-specific.
\medskip
\noindent\textbf{Keywords:} price elasticity; design uncertainty; Double Machine Learning; coverage; short panels; variance components.

\noindent\textbf{JEL:} C33, C55, L11, M31.
\end{abstract}

\clearpage
\section{Introduction}\label{sec:intro}

Empirical pricing work often reports an elasticity estimate together with a confidence interval constructed from the realised panel. In short panels, however, the number of rows can be a poor description of the amount of treatment variation available for identifying the price coefficient. A manufacturer may observe many region--week cells while changing a common list price only a handful of times. This paper asks what conventional inferential procedures represent in that setting and what additional uncertainty appears when the realised price trajectory itself is treated as one draw from a broader pricing process.

The organizing distinction is between \emph{within-design} and \emph{between-design} uncertainty. Let a design $D$ denote the realised configuration of price movements and other non-shock features of the panel, and let $u$ denote the demand shock. Conditional on $D$, repeated draws of $u$ generate the familiar sampling or shock uncertainty around an estimator. Across alternative realisations of $D$, the estimator can also be displaced because different price trajectories provide different residualised treatment variation and different alignments with confounders. We refer to the dispersion of this design-specific centring error as \emph{between-design uncertainty}. The distinction is conceptually related to the separation between sampling-based and design-based uncertainty in \citet{abadie2020sampling} and to recent design-based analyses of quasi-experiments \citep{rambachan2025design}. Our use of the terminology is narrower: the simulations explicitly generate multiple pricing trajectories so that the two components of the estimation error can be measured separately.

This distinction also clarifies what can and cannot be learned from a single realised panel. A block bootstrap, a cluster-robust covariance estimator, or a multiway-clustered covariance estimator can be appropriate for aspects of dependence and repeated-sampling variation supported by the observed data. None of these procedures is designed, by itself, to nonparametrically recover the distribution of estimator displacement across price histories that were not observed. This is not the claim that such methods are generically invalid. It is a statement about the object they condition on and about the additional assumptions or data required to learn a distribution over alternative designs.

We study this issue in a synthetic data-generating process with $120$ weeks, sparse list-price movements, promotions, competitor prices, calendar variables, weather and inflation. The baseline configuration has ten list-price changes per series with magnitudes between $3.0\%$ and $7.5\%$. These values are chosen to be broadly comparable with empirical moments reported by \citet{nakamura2008five}; the calibration is intended to place the exercise in a plausible pricing regime rather than to establish external validity for any particular category. Because the generator supplies the ground truth and permits repeated shocks within the same price trajectory as well as repeated price trajectories, it makes observable a decomposition that cannot be learned nonparametrically from one realised historical panel.

The baseline simulation is deliberately stark. For the gradient-boosted specification, $97.6\%$ of the variance of the estimation error is associated with variation in the design-specific conditional mean error; for the sieve specification the corresponding share is $99.3\%$. The reported bootstrap standard error is larger than the measured within-design standard deviation but much smaller than the across-design dispersion. Thus the empirical finding is not simply that one bootstrap happens to be too narrow. Rather, in these simulations a substantial part of the coverage shortfall is associated with variation in the conditional centring of the estimator across realised price histories.

\paragraph{Contributions.} The paper makes three related contributions.

\begin{enumerate}
\item \textbf{Within-design versus between-design uncertainty.} We formulate the variance decomposition on the estimation error $\hat\theta-\theta(D)$ and measure its two components by holding the design fixed while resampling shocks and then repeating the exercise across designs. In the baseline DGP, between-design centring dispersion is much larger than within-design shock dispersion. The exercise shows why procedures based only on one realised panel need not represent repeated-design uncertainty, while avoiding the stronger claim that all within-panel inference is invalid for conditional questions.

\item \textbf{More observations are not the same as more identifying trajectories.} Adding region--week rows under a common national list price may improve estimation of outcome-side objects and reduce some conditional noise, but it does not multiply the number of independently realised price paths. More generally, the variance of an average of design-specific errors depends on their covariance. Under approximately independent, mean-zero design-specific errors, averaging $k$ independently priced units reduces their standard deviation by $\sqrt{k}$; under a common design component, the reduction is smaller and can disappear. This provides the formal basis for the paper's practical message that ``adding rows'' is not equivalent to ``adding identification.''

\item \textbf{From inferential correction to data design.} Across the simulated grid, the relation $\hat\sigma_b \approx 0.182V^{-0.271}$ summarizes how design dispersion changes with the heuristic variation index $V=n_{\text{moves}}\times\text{magnitude}^2$. The exponent is reported only as an empirical regularity of these simulations. A variance-component construction across independently priced units improves repeated-design coverage, but it becomes wide and, when true effects differ, mixes design dispersion with genuine heterogeneity. These findings shift the practical emphasis from searching only for a more elaborate post-hoc interval toward generating additional independent treatment variation. Randomized regional pricing is one transparent design option, not a logically necessary one.
\end{enumerate}

\paragraph{Conditional and repeated-design questions.} The paper reports both conditional coverage given a realised design and coverage averaged over generated designs. These answer different questions. Conditional coverage asks how an interval behaves if the observed price trajectory is held fixed and only the shock is repeated. Repeated-design coverage asks how the full estimation procedure behaves when the price trajectory is itself redrawn from the DGP. The second object is useful for methodological evaluation and for a firm contemplating future pricing histories; it should not be conflated with the first.

\paragraph{Relation to weak identification.} Sparse price movement reduces the residualised treatment variation available in the partially linear regression. This is naturally related to the broader weak-identification and weak-signal literature \citep[e.g.,][]{stock2005testing, andrews2019weak}, but the connection is analogical rather than an application of weak-IV asymptotics. The model here does not introduce a weak instrument. The relevant empirical symptom is that the denominator based on residualised price variation can be small or highly design-dependent, making finite-sample centring and dispersion sensitive to the realised trajectory. We therefore use ``limited identifying variation'' as the primary description and reserve ``weak identification'' for the broader conceptual connection.

\paragraph{Variance components across units.} To use multiple realised price histories, we adapt a standard random-effects variance-component estimator from meta-analysis \citep{paule1982consensus, dersimonian1986meta, hartung2001refined}. The Paule--Mandel estimator is not a mathematical contribution of the paper. Its role is to estimate excess between-unit dispersion under a working model in which units provide approximately independent design realisations, within-unit variances are adequately measured, and the design-specific displacement is exchangeable with mean zero. When the true unit-level elasticities are common, this excess dispersion can be interpreted as design uncertainty under those assumptions. When true elasticities differ, the between-unit component generally combines genuine heterogeneity with design uncertainty, so it no longer identifies the latter without additional structure.

\paragraph{Relation to pricing and inference practice.} Limited and endogenous price variation is a long-standing problem in empirical demand estimation \citep{villas1999endogeneity, bijmolt2005new}, and recent pricing work emphasizes the value of controlled variation for learning demand and treatment heterogeneity \citep{ferreira2016analytics, dube2023personalized}. Our contribution is narrower: we use the simulation environment to distinguish additional rows observed under essentially the same price path from additional, weakly dependent price trajectories, and to quantify how that distinction affects repeated-design error. The centring issue is also related to bias-aware inference \citep{armstrong2018simple}, although the displacement studied here is generated by the realised pricing design rather than by a generic smoothing-bias bound. Multiway clustering and related covariance corrections remain relevant for dependence within the observed panel \citep{cameron2011robust, chiang2022multiway}; they answer a different question from learning the distribution of errors across unrealised pricing designs.

\paragraph{Scope of the evidence.} Proposition~\ref{prop:agg} below is an algebraic variance identity. All quantitative magnitudes---including the $97.6\%$ decomposition, the coverage results, the estimated exponent $-0.271$, and the variance-component performance---are simulation results from the stated DGP and grid. We do not interpret the exponent as a general rate, the baseline variance share as a population constant, or the failure of the eight evaluated intervals as an impossibility theorem. The simulations are intended to isolate a mechanism and to show conditions under which data design, rather than row count alone, materially changes inference.

\section{Setup}\label{sec:setup}

\subsection{The estimand}

Let $i$ index product units, $r$ regions and $t$ weeks. The outcome is
$Y_{irt} = \log Q_{irt}$, log volume; the treatment is $T_{irt} = \log
P^{\text{si}}_{irt}$, the log list price the manufacturer sets. We work with the
partially linear model \citep{robinson1988root}
\begin{equation}
Y = \theta\, T + g(W) + \varepsilon, \qquad
T = m(W) + \eta, \qquad
\mathbb{E}[\varepsilon \mid W,\eta]=0,\ \ \mathbb{E}[\eta\mid W]=0,
\label{eq:plr}
\end{equation}
with $W$ a vector of observed confounders (promotional flag and intensity, log
competitor price, weather, holidays, working days, week of year, log inflation).
The parameter $\theta$ is estimated with a cross-fitted residual-on-residual procedure motivated by Double Machine Learning \citep{chernozhukov2018double}. Nuisance functions $\hat\ell = \hat{\mathbb{E}}[Y\mid W]$ and $\hat m = \hat{\mathbb{E}}[T \mid W]$ are fit on $K=5$ contiguous temporal blocks, with all regional observations of a week assigned to the same block. The implementation aggregates fold-specific coefficients by their median, as in \eqref{eq:dml}; because this is not the canonical pooled DML2 formula, all claims below should be read as applying to the estimator actually implemented here rather than to DML estimators as a class:
\begin{equation}
\hat\theta \;=\; \operatorname*{median}_{k}
\frac{\sum_{i \in I_k}\tilde T_i \tilde Y_i}{\sum_{i \in I_k}\tilde T_i^2},
\qquad
\tilde Y_i = Y_i - \hat\ell^{(-k)}(W_i),\quad
\tilde T_i = T_i - \hat m^{(-k)}(W_i).
\label{eq:dml}
\end{equation}

\begin{remark}[the estimand under split pricing]\label{rem:split}
In the channel that motivates this work the manufacturer sets the list price
while a shopkeeper sets the shelf price, so volume responds to a price the
manufacturer does not control. Writing $\rho = \partial \log P^{\text{so}} /
\partial \log P^{\text{si}}$ for the pass-through and $\varepsilon_c$ for the
consumer elasticity against the shelf price, the quantity identified by
\eqref{eq:dml} is the convolution $\theta = \varepsilon_c\,\rho$ and not
$\varepsilon_c$. Everything below concerns $\theta$; the separation of the two
is a distinct problem and is not our subject. We note the point only so that the
ground truth against which we measure is unambiguous: it is $\varepsilon_c \rho$
as realised in each generated panel, read from the generator rather than from its
configuration.
\end{remark}

\subsection{The generator}\label{sec:generator}

Panels are drawn from a data-generating process described in
Appendix~\ref{app:dgp}. Volume is generated jointly with prices, costs,
promotions, calendar effects, weather and an inflation index, so that
confounding is present by construction. The baseline configuration is $W=120$
weeks, six regions, list prices moving ten times per series by $3.0$--$7.5\%$,
pass-through $\rho = 0.85$, a cross-price elasticity of $+0.30$ on the log
competitor price, and a demand shock of standard deviation $0.05$ in logs.

To ensure that this data-generating process reflects realistic market dynamics rather than a pathological edge case, these parameters are calibrated against the empirical moments in \citet{nakamura2008five}. In US microdata, regular price changes exhibit a median monthly frequency of $9$--$12\%$---corresponding to $8$--$11$ months of median duration, or roughly $2.5$--$3.3$ price movements over a $120$-week horizon. The median absolute size of regular price adjustments is $8.5\%$, whereas temporary sales are more than three times larger at a median of $29.5\%$. The baseline simulation regime directly mirrors these benchmark frequencies and magnitudes.

A \emph{design} is a realisation of everything except the demand shock; a
\emph{replication} of a design is a fresh draw of that shock alone. This
separation is what makes Section~\ref{sec:decomposition} possible and is
implemented by giving the shock its own random stream.

Two learners recur throughout. \textsf{gbr} is a gradient-boosted regressor
(depth $3$, $200$ trees, learning rate $0.05$), the specification most common in
applied DML. \textsf{sieve} is a ridge on an explicit basis: the confounders, a
quadratic time trend and three annual harmonic pairs. They differ sharply in
bias and variance, and the contrast turns out to matter for the aggregation
result of Section~\ref{sec:aggregation}.

\section{Within- and between-design uncertainty}\label{sec:decomposition}

\begin{definition}[within- and between-design error dispersion]
Let $D$ denote a design, $u$ a shock realisation, and $\theta(D)$ the design-specific ground truth used by the generator. Define the estimation error
\begin{equation}
e(D,u)=\hat\theta(D,u)-\theta(D),
\end{equation}
and its design-specific conditional mean $b(D)=\mathbb{E}_u[e(D,u)\mid D]$. The \emph{within-design} variance is
\begin{equation}
\sigma_w^2=\mathbb{E}_D\!\left[\operatorname{Var}_u(e(D,u)\mid D)\right],
\end{equation}
and the \emph{between-design centring variance} is
\begin{equation}
\sigma_b^2=\operatorname{Var}_D\!\left(b(D)\right).
\end{equation}
By the law of total variance,
\begin{equation}
\operatorname{Var}_{D,u}(e(D,u))=\sigma_w^2+\sigma_b^2.
\end{equation}
If $\bar b=\mathbb{E}_D[b(D)]$ is nonzero, the corresponding mean-squared error is $\sigma_w^2+\sigma_b^2+\bar b^2$.
\end{definition}

\begin{definition}[conditional and repeated-design coverage]
Let $CI_{1-\alpha}(D,u)$ denote an interval computed from one realised panel. Its \emph{conditional coverage} given $D$ is
\begin{equation}
\Pr_u\!\left(\theta(D)\in CI_{1-\alpha}(D,u)\mid D\right).
\end{equation}
Its \emph{repeated-design coverage} under the simulation DGP is
\begin{equation}
\Pr_{D,u}\!\left(\theta(D)\in CI_{1-\alpha}(D,u)\right)
=\mathbb{E}_D\!\left[\Pr_u\!\left(\theta(D)\in CI_{1-\alpha}(D,u)\mid D\right)\right].
\end{equation}
The equality is a probability identity; the substantive distinction is which source of randomness is treated as relevant to the inferential question.
\end{definition}

A within-panel procedure is constructed from one realised design. Depending on the resampling scheme and dependence structure, it may approximate aspects of the conditional distribution more or less well, but it does not by itself provide nonparametric information about the distribution of $b(D)$ over price trajectories that were not realised. We therefore do not equate the bootstrap standard error mechanically with $\sigma_w$. Instead, we measure $\sigma_w$ and $\sigma_b$ directly from the generator. We draw $12$ designs and $40$ shock replications within each, estimate $\theta$ on all $480$ panels, and compute the error decomposition relative to each design's ground truth.

\begin{table}[t]
\centering
\caption{Decomposition of the dispersion of the estimation error. Twelve designs, forty
shock replications each ($n=480$). $\sigma_w$ is measured within design,
$\sigma_b$ is the standard deviation of the design-specific conditional mean error $b(D)$; $\widehat{\text{se}}$ is the weekly moving-block bootstrap standard error from the baseline specification. Coverage is averaged over the design-specific conditional coverage probabilities and is measured against each design's own ground truth, at a nominal $0.95$.}
\label{tab:decomp}
\begin{tabular}{lrrrrrr}
\toprule
Learner & Bias & $\sigma_w$ & $\sigma_b$ & $\widehat{\text{se}}$
        & $\sigma_b^2/(\sigma_w^2{+}\sigma_b^2)$ & Coverage \\
\midrule
\textsf{gbr}   & $0.150$ & $0.0756$ & $0.4815$ & $0.1595$ & $97.6\%$ & $0.565$ \\
\textsf{sieve} & $0.016$ & $0.0589$ & $0.7187$ & $0.1553$ & $99.3\%$ & $0.242$ \\
\bottomrule
\end{tabular}
\end{table}

Table~\ref{tab:decomp} reports the result. Three features deserve comment.

\paragraph{Between-design centring dispersion is large in the baseline DGP.} It accounts for $97.6\%$ of the variance of the estimation error for \textsf{gbr} and $99.3\%$ for \textsf{sieve}. Conditional on a design, the estimator is relatively concentrated---$\sigma_w=0.0756$ for \textsf{gbr}---but its conditional mean error varies substantially across designs. These percentages characterize the simulated generator and should not be read as universal variance shares for short pricing panels.

\paragraph{The bootstrap standard error is not an estimate of the full decomposition.} For \textsf{gbr}, the mean bootstrap standard error, $0.1595$, is about $2.1$ times the measured within-design standard deviation and about one third of $\sigma_b$. Thus it is not accurate to describe this bootstrap as simply estimating $\sigma_w$ in the finite sample considered here. The more limited conclusion is that a procedure built from one realised panel does not reproduce the across-design distribution measured by repeated generation of price trajectories.

\paragraph{The simulations point to a centring component.} The observed coverage shortfall is difficult to reconcile with interval width alone. For \textsf{gbr}, a typical half-width of about $0.305$ is large relative to $\sigma_w=0.0756$, yet mean conditional coverage across designs is $0.565$. In the generator, the design-specific conditional mean error $b(D)$ varies substantially across designs. A rough calculation based on the dispersion of $b(D)$ yields coverage of the same order as the simulation. We interpret this as evidence that design-dependent centring is an important mechanism in this DGP, not as a proof that every resampling method fails whenever treatment variation is sparse.

\begin{remark}[which target is the right one]
For a firm interpreting one realised historical trajectory, conditional coverage is a natural target. For evaluating a procedure that will be reused across future pricing histories, repeated-design coverage is also informative. Neither target dominates by definition; the relevant one depends on the decision problem. The simulations report both so that a conditional inferential question is not inadvertently replaced by an unconditional one.
\end{remark}

\section{No evaluated within-panel construction achieves nominal coverage}\label{sec:eight}

To assess whether familiar within-panel procedures are sufficient in this DGP, we compare eight constructions on identical fits over $200$ independently drawn panels: the baseline moving-block bootstrap over weeks; the same at $600$ draws; a hierarchical block bootstrap that also
resamples cells within replicated weeks; the analytic i.i.d.\ standard error;
cluster-robust by week; multiway cluster-robust by week and region
\citep{cameron2011robust, chiang2022multiway}; a Newey--West long-run variance
on the weekly score \citep{newey1987simple}; and a wild cluster bootstrap-$t$
with Rademacher weights \citep{cameron2008bootstrap}.

\begin{table}[t]
\centering
\caption{Eight interval constructions, identical point estimator and identical
fits, over $200$ panels. Nominal level $0.95$. \emph{Reference repeated-design width} is $3.92\,\hat\sigma$ with $\hat\sigma$ the realised across-panel standard deviation of the estimation error. This is a descriptive benchmark, not a confidence interval with a formal finite-sample coverage guarantee.}
\label{tab:eight}
\begin{tabular}{lrrrr}
\toprule
& \multicolumn{2}{c}{\textsf{gbr}} & \multicolumn{2}{c}{\textsf{sieve}} \\
\cmidrule(lr){2-3}\cmidrule(lr){4-5}
Construction & Coverage & Width & Coverage & Width \\
\midrule
Multiway cluster (week $\times$ region) & $0.730$ & $1.062$ & $0.660$ & $1.167$ \\
Block bootstrap, cells resampled       & $0.595$ & $0.712$ & $0.435$ & $0.723$ \\
Newey--West on the weekly score        & $0.580$ & $0.798$ & $0.230$ & $0.451$ \\
Block bootstrap, $B=600$               & $0.545$ & $0.591$ & $0.300$ & $0.581$ \\
Block bootstrap, $B=250$ (baseline)  & $0.525$ & $0.597$ & $0.310$ & $0.566$ \\
Wild cluster bootstrap-$t$             & $0.350$ & $0.451$ & $0.105$ & $0.258$ \\
Cluster-robust by week                 & $0.345$ & $0.445$ & $0.105$ & $0.254$ \\
Analytic i.i.d.                        & $0.285$ & $0.338$ & $0.155$ & $0.308$ \\
\midrule
\emph{Reference repeated-design width} $3.92\,\hat\sigma$ & \multicolumn{2}{c}{$1.680$}
                                       & \multicolumn{2}{c}{$2.212$} \\
\bottomrule
\end{tabular}
\end{table}

None of the eight evaluated constructions reaches $0.95$ coverage in this experiment (Table~\ref{tab:eight}). Coverage is generally higher for wider intervals, and multiway clustering has the highest measured coverage for both learners among the procedures considered. This ranking should not be read as an impossibility result for all within-panel methods; it is a finite simulation comparison of these eight constructions. The table also emphasizes that the empirical shortfall is not resolved simply by choosing the widest evaluated interval.

For the decision exercise used in this paper, we pre-specify an application-specific width threshold of $W\le 0.6$. This threshold is not a general standard for pricing studies. Under that criterion, the higher-coverage intervals in Table~\ref{tab:eight} remain too wide to satisfy the paper's operational precision requirement. The resulting coverage--width trade-off motivates the across-unit variance-component exercise in Section~\ref{sec:hierarchical}.

\section{An empirical dispersion pattern over the simulation grid}\label{sec:frontier}

If design variance is what binds, the natural question is how much identifying
variation would be needed to make it small. We sweep the number of price moves
$n \in \{2,4,6,8,12,16\}$, their magnitude (four ranges from $0.4$--$0.9\%$ to
$3.0$--$7.5\%$), the degree of confounding with promotion
($\{0, 0.5, 0.9\}$) and the pass-through ($\{0, 0.30, 0.60, 0.85\}$): $288$ grid
points, each estimated with both learners and replicated eight times, for $576$
configurations and $4{,}608$ fits in total. At each point we measure the realised
dispersion of $\hat\theta$ across replications rather than its reported standard
error.

\begin{table}[t]
\centering
\caption{A selected column of the simulation grid: no confounding with
promotion, $\rho = 0.85$, moves averaging $5.25\%$, learner \textsf{gbr}. True
$\theta = -1.114$; eight replications per grid point, so $\sigma$ is itself
estimated with visible noise --- which is why the fit \eqref{eq:scaling} is taken
over the whole column rather than read off adjacent rows. $\operatorname{Var}
(\tilde T)/\operatorname{Var}(T)$ is the share of the log list price surviving
residualisation, i.e.\ the identifying variation the estimator actually uses. \textit{Filtered ratio} indicates the proportion of simulated configurations satisfying the pre-estimation identification criterion (a treatment coefficient of variation $CV > 0.03$ and at least four discrete price movements).}
\label{tab:frontier}
\begin{tabular}{rrrrr}
\toprule
Moves & Bias & $\sigma$ (design)
      & $\operatorname{Var}(\tilde T)/\operatorname{Var}(T)$ & Filtered ratio \\
\midrule
$2$  & $0.254$ & $0.681$ & $0.352$ & $0.000$ \\
$4$  & $0.243$ & $0.869$ & $0.407$ & $0.375$ \\
$6$  & $0.471$ & $0.429$ & $0.530$ & $0.688$ \\
$8$  & $0.290$ & $0.444$ & $0.491$ & $0.729$ \\
$12$ & $0.290$ & $0.394$ & $0.569$ & $0.896$ \\
$16$ & $0.145$ & $0.382$ & $0.547$ & $0.958$ \\
\bottomrule
\end{tabular}
\end{table}

Let $V=n\cdot\text{magnitude}^2$ summarize the amount and size of list-price movement. The index is a deliberately simple design summary; it is not asserted to be Fisher information or a sufficient statistic for identification. Over the $24$ points with no promotion confounding and full pass-through---six move counts $\times$ four magnitude ranges for learner \textsf{gbr}---a log--log regression gives
\begin{equation}
\hat\sigma_b \;=\; 0.182\;V^{-0.271},\qquad
\text{s.e.}(-0.271)=0.023,\qquad R^2=0.86.
\label{eq:scaling}
\end{equation}

Equation~\eqref{eq:scaling} is an empirical regularity of this simulation column. The exponent $-0.271$ should not be interpreted as a theoretical convergence rate for DML, for short panels, or for observational pricing designs in general. A $-1/2$ slope can serve as a heuristic reference in stylized settings where information grows proportionally with a sum of squared treatment movements, but $V$ is not shown here to satisfy that information equivalence. The comparison is therefore descriptive rather than an asymptotic test.

Within the simulated grid, larger $V$ is associated with lower across-design dispersion, but the decline is slow enough that the explored configurations retain substantial error dispersion. Of the $576$ learner--configuration combinations, none simultaneously attains absolute bias below $0.15$ and design dispersion below $0.20$. This is a statement about the enumerated grid and its finite replication count, not an impossibility result outside the grid. The best cells should likewise be interpreted cautiously because each grid point has only eight replications.

\begin{remark}[on extrapolating \eqref{eq:scaling}]\label{rem:extrap}
If one mechanically extrapolates the fitted relation, halving $\sigma_b$ corresponds to multiplying $V$ by approximately $2^{1/0.271}\approx13$. Such calculations extend beyond the simulated support and are included only to convey the curvature of the fitted relationship. They are not forecasts of how a real category would respond to thirteen times more price variation.
\end{remark}

\paragraph{Implication for precision calculations.} On this grid, precision measures based on the reported within-panel standard errors can be much smaller than the realised repeated-design error dispersion. The median ratio of repeated-design dispersion to the reported standard-error scale is $2.26$ for \textsf{gbr} and $8.03$ for \textsf{sieve}. These ratios are simulation diagnostics, not generic correction factors for minimum detectable effects.

\section{Common versus independent price paths under aggregation}\label{sec:aggregation}

The key distinction is not panel width by itself but the covariance structure of the design-specific estimation errors contributed by the added units.

\begin{proposition}[variance of averaged design-specific centring errors]
\label{prop:agg}
Let $b_j$ be the design-specific conditional mean error for unit $j$, and let $\bar b_k=k^{-1}\sum_{j=1}^k b_j$. Then
\begin{equation}
\operatorname{Var}(\bar b_k)
=\frac{1}{k^2}\left(\sum_{j=1}^k\operatorname{Var}(b_j)
+2\sum_{j<\ell}\operatorname{Cov}(b_j,b_\ell)\right).
\label{eq:aggvar}
\end{equation}
If $\operatorname{Var}(b_j)=\sigma_b^2$ and all pairwise correlations equal $\rho$, then
\begin{equation}
\operatorname{Var}(\bar b_k)=\sigma_b^2\frac{1+(k-1)\rho}{k}.
\label{eq:equicorr}
\end{equation}
Hence $\rho=0$ gives a standard deviation of $\sigma_b/\sqrt{k}$, while $\rho=1$ gives no reduction. Aggregation does not remove a nonzero common mean bias $\mathbb{E}[b_j]$.
\end{proposition}

The proposition is an algebraic identity, not a claim that shared prices automatically imply $\rho=1$ or that distinct price paths automatically imply $\rho=0$. In the simulation, regions exposed to the same national list-price trajectory inherit a large common design component, whereas product units are generated with distinct price trajectories and approximately weak cross-unit dependence in that component. The empirical exercise below asks whether the observed covariance structure is close enough to the independent benchmark to produce the $\sqrt{k}$ pattern.

\begin{table}[t]
\centering
\caption{Estimating the same object at three levels of aggregation, over $200$
panels of eight product units in two brands of one category. Residualisation is
always per unit. The Frisch--Waugh--Lovell logic \citep{lovell1963seasonal} relates the pooled residual-on-residual coefficient to a specification with unit effects; the exercise is designed so that aggregation combines information from multiple unit-specific price histories. \emph{Units} is the number of units pooled per estimate.
$\sigma$ (design) is the dispersion across replications, and \emph{Width} is the mean width of the baseline block-bootstrap interval at that level.}
\label{tab:levels}
\begin{tabular}{llrrrrrrr}
\toprule
Learner & Level & Units & Bias & $\sigma$ (design) & RMSE & Coverage & Width
        & $\sigma$ ratio \\
\midrule
\textsf{gbr}   & unit     & $1$ & $0.315$ & $0.4603$ & $0.559$ & $0.474$ & $0.693$ & --- \\
\textsf{gbr}   & brand    & $4$ & $0.293$ & $0.2219$ & $0.368$ & $0.370$ & $0.414$ & $2.07$ \\
\textsf{gbr}   & category & $8$ & $0.291$ & $0.1538$ & $0.327$ & $0.280$ & $0.364$ & $2.99$ \\
\addlinespace
\textsf{sieve} & unit     & $1$ & $0.021$ & $0.4737$ & $0.475$ & $0.282$ & $0.422$ & --- \\
\textsf{sieve} & brand    & $4$ & $0.021$ & $0.2267$ & $0.227$ & $0.282$ & $0.191$ & $2.09$ \\
\textsf{sieve} & category & $8$ & $0.022$ & $0.1585$ & $0.157$ & $0.240$ & $0.134$ & $2.99$ \\
\midrule
\multicolumn{6}{l}{\emph{Predicted by Proposition~\ref{prop:agg}}}
& \multicolumn{3}{r}{$\sqrt{4}=2.00$, \ $\sqrt{8}=2.83$} \\
\bottomrule
\end{tabular}
\end{table}

Table~\ref{tab:levels} is close to the independent-design benchmark in this simulation: the observed reduction factors are $2.07$ and $2.09$ against $\sqrt{4}=2.00$, and $2.99$ against $\sqrt{8}=2.83$, for both learners. This is evidence that the simulated cross-unit design-specific errors have low enough covariance for the $\sqrt{k}$ approximation to be useful here; it is not evidence that product-level price histories are independent in empirical applications.

Two further readings of the table matter. First, aggregation improves RMSE substantially in the reported simulations, but the baseline interval does not gain coverage because its width contracts along with the empirical error dispersion. Second, the two learners respond differently. The \textsf{gbr} mean bias remains near $0.30$ across aggregation levels, which is consistent with a persistent common centring component; its coverage falls as the interval tightens. The \textsf{sieve} mean bias remains near $0.02$, so the reduction in weakly correlated dispersion translates more directly into lower RMSE. These are empirical patterns of the simulated portfolio, not generic properties of boosted trees or sieve learners.

\begin{table}[t]
\centering
\caption{Six nuisance learners at unit level, $120$ panels. \emph{Low-frequency
share} is the fraction of the residual's spectral mass below one twentieth of the
series length. No learner minimises bias, maximises coverage and minimises RMSE
simultaneously.}
\label{tab:learners}
\begin{tabular}{lrrrr}
\toprule
Learner & Bias & RMSE & Coverage & Low-frequency share \\
\midrule
\textsf{lasso}  & $-0.007$ & $0.652$ & $0.258$ & $0.508$ \\
\textsf{ridge}  & $-0.008$ & $0.651$ & $0.250$ & $0.510$ \\
\textsf{sieve}  & $-0.009$ & $0.652$ & $0.258$ & $0.505$ \\
\textsf{rbf}    & $ 0.055$ & $0.534$ & $0.533$ & $0.277$ \\
\textsf{histgb} & $ 0.153$ & $0.517$ & $0.383$ & $0.128$ \\
\textsf{gbr}    & $ 0.176$ & $0.501$ & $0.383$ & $0.144$ \\
\bottomrule
\end{tabular}
\end{table}

Table~\ref{tab:learners} places this in context. At unit level the choice of
nuisance learner is a genuine trade-off: the regularised learners are nearly
unbiased and have the worst RMSE and coverage, the boosted learners are the
reverse, and no learner wins on all three. The ranking therefore changes with the aggregation level in these simulations. This illustrates, rather than proves generally, that nuisance-learner comparisons can depend on the level at which the estimand is reported, because aggregation attenuates weakly correlated dispersion but not a persistent common mean bias.

\section{A variance-component interval estimated across realised designs}\label{sec:hierarchical}

A single realised price trajectory does not nonparametrically identify the distribution of $b(D)$ across counterfactual trajectories. Multiple units with independently or weakly dependently realised pricing designs can supply information about that distribution, but only under additional structure. We use the working model
\begin{equation}
\hat\theta_j=\theta_j+B_j+\varepsilon_j,\qquad
\mathbb{E}[B_j]=0,\quad \operatorname{Var}(B_j)=\tau_B^2,\quad
\mathbb{E}[\varepsilon_j]\approx0,\quad \operatorname{Var}(\varepsilon_j)\approx s_j^2,
\label{eq:random_bias}
\end{equation}
with $B_j$ exchangeable across units and approximately independent of the within-unit error $\varepsilon_j$. The interpretation of $\tau_B^2$ as design variance additionally requires that the units' pricing trajectories act as independent or sufficiently weakly dependent design draws and that $s_j^2$ adequately represents the within-unit component.

Let $\hat\theta_j$ and $s_j^2$ be the estimate and bootstrap variance of unit $j$, $j=1,\dots,J$. We estimate an excess between-unit variance with the Paule--Mandel moment equation \citep{paule1982consensus}
\begin{equation}
\sum_j\frac{(\hat\theta_j-\hat\mu)^2}{s_j^2+\tau^2}=J-1,
\qquad
\hat\mu=\frac{\sum_j(s_j^2+\tau^2)^{-1}\hat\theta_j}{\sum_j(s_j^2+\tau^2)^{-1}}.
\label{eq:pm}
\end{equation}
Paule--Mandel is a standard between-study variance estimator; the contribution here is its use as a working estimator of excess dispersion across independently realised pricing designs, not a new variance-component method. We compare the bootstrap percentile interval; a conjugate normal--normal posterior interval based on a leave-one-out branch prior; a variance-augmented interval $\hat\theta_j\pm q\sqrt{s_j^2+\hat\tau^2}$; and its posterior-centred analogue $\hat\theta_j^{\mathrm{post}}\pm q\sqrt{v_j^{\mathrm{post}}+\hat\tau^2}$, where $q=z_{1-\alpha/2}$.

The homogeneous and heterogeneous scenarios have different interpretations. In the homogeneous scenario all eight units share the same true $\theta$ and have independently generated price paths. Under model \eqref{eq:random_bias}, excess between-unit dispersion can then be attributed to the design component up to estimation error in $s_j^2$ and $\hat\tau^2$. In the heterogeneous scenario the true $\theta_j$ differ, so the same between-unit component generally combines true effect heterogeneity and design-specific displacement. It should therefore be viewed as a conservative mixture rather than as a clean estimate of design variance.

\begin{table}[t]
\centering
\caption{Four intervals on identical fits, $200$ panels of eight units per
scenario ($n=1{,}600$ unit-panels). Nominal level $0.95$. In the homogeneous
scenario the eight units share the same true $\theta$, so $\hat\tau$ is interpretable as excess design-related dispersion under the working-model assumptions described in the text.}
\label{tab:hier}
\begin{tabular}{llrr}
\toprule
Scenario & Interval & Coverage & Width \\
\midrule
Homogeneous & Posterior-centred variance-augmented $\sqrt{v^{\text{post}}+\hat\tau^2}$ & $0.931$ & $1.480$ \\
Homogeneous & Variance-augmented $\sqrt{s^2+\hat\tau^2}$                      & $0.906$ & $1.516$ \\
Homogeneous & Bootstrap percentile                                     & $0.469$ & $0.549$ \\
Homogeneous & Posterior $\sqrt{v^{\text{post}}}$                       & $0.426$ & $0.501$ \\
\addlinespace
Heterogeneous & Posterior-centred variance-augmented                    & $0.943$ & $1.623$ \\
Heterogeneous & Variance-augmented                                     & $0.931$ & $1.650$ \\
Heterogeneous & Bootstrap percentile                                   & $0.475$ & $0.539$ \\
Heterogeneous & Posterior                                              & $0.426$ & $0.502$ \\
\bottomrule
\end{tabular}
\end{table}

Table~\ref{tab:hier} reports the outcome. In the homogeneous scenario, adding the estimated between-unit component raises empirical coverage from $0.469$ for the bootstrap percentile interval to $0.931$ for the posterior-centred variance-augmented interval. This is a substantial movement toward the nominal $0.95$ level, but it is not exact nominal coverage. The estimated $\hat\tau=0.372$ is much larger than the mean bootstrap standard error of $0.144$, which is consistent with substantial excess dispersion across independently generated unit-level designs in this simulation.

\paragraph{Shrinkage alone does not address the simulated centring dispersion.} The posterior interval has coverage $0.426$, below the bootstrap percentile interval. In this DGP, shrinkage reduces the conditional variance while leaving design-specific displacement largely unmodelled. The result is consistent with the centring mechanism documented in Section~\ref{sec:decomposition}; it should not be generalized to hierarchical Bayes procedures that explicitly model the design process or bias component.

\paragraph{The coverage improvement is purchased with width.} The two variance-augmented constructions are $2.70$ and $2.76$ times as wide as the bootstrap in the homogeneous simulation, at widths of $1.480$ and $1.516$. Under the paper's pre-specified operational threshold, these intervals are too wide to support the intended unit-level pricing decision. The result illustrates the trade-off in this DGP: representing excess across-design dispersion improves coverage but can materially reduce decision precision.

\section{The variance-component interval at higher aggregation levels}\label{sec:level}

Sections~\ref{sec:aggregation} and \ref{sec:hierarchical} suggest a natural combined exercise. In the simulated portfolio, aggregation is close to the independent-design $\sqrt{k}$ benchmark, while the variance-component construction produces a wider interval at unit level. Combining them is
therefore the natural question, and it is the one on which the practical value
of the whole argument turns: does the variance-component interval, evaluated at a higher
level of aggregation, become narrow enough to support a decision?

Arithmetic from Tables~\ref{tab:levels} and \ref{tab:hier} suggests it should.
At category level the design dispersion is $0.1585$ and the bootstrap standard
error $0.034$, so the variance-augmented interval would have width $2 \times 1.96 \times \sqrt{0.034^2 + 0.1585^2} = 0.636$, against the $1.516$ that the same construction attains at unit level in Table~\ref{tab:hier} --- moving from uninformative toward the target precision bound. The arithmetic is not decisive,
however, because $\hat\tau$ is re-estimated among the units of whatever level is
being published, and its magnitude at that level is not the design dispersion of
Table~\ref{tab:levels}. We therefore run the construction directly, on a
portfolio of $24$ units --- six categories, two brands each, two units per brand
--- estimating $\theta$ at each of the three levels and solving \eqref{eq:pm}
between the units of that level: $24$ units at unit level, $12$ at brand, $6$ at
category, over $100$ replications per scenario.

\begin{table}[t]
\centering
\caption{The variance-augmented interval $\hat\theta \pm q\sqrt{s^2+\hat\tau^2}$ at three
levels of aggregation, $100$ replications per scenario on a portfolio of $24$
units ($6$ categories $\times$ $2$ brands $\times$ $2$ units). $\hat\tau$ is
estimated between the units of each level, so its $\sqrt{k}$ column is the
prediction of Proposition~\ref{prop:agg} applied to $k$ = units pooled per
estimate. \emph{Ratio to threshold} indicates interval width expressed as a multiple
of the $0.6$ operational precision bound; no combination falls below $1.0$.}

\label{tab:nivel}
\begin{tabular}{llrrrrrr}
\toprule
Scenario & Learner & Level & $k$ & $\hat\tau$ & $\sqrt{k}$ pred.
          & Coverage & Rel. Width ($W/0.6$) \\
\midrule
Homogeneous & \textsf{gbr}   & unit     & $1$ & $0.4052$ & ---      & $0.899$ & $1.725$ \ ($2.88$) \\
Homogeneous & \textsf{gbr}   & brand    & $2$ & $0.2779$ & $0.2865$ & $0.830$ & $1.191$ \ ($1.99$) \\
Homogeneous & \textsf{gbr}   & category & $4$ & $0.1908$ & $0.2026$ & $0.665$ & $0.829$ \ ($1.38$) \\
\addlinespace
Homogeneous & \textsf{sieve} & unit     & $1$ & $0.4569$ & ---      & $0.953$ & $1.816$ \ ($3.03$) \\
Homogeneous & \textsf{sieve} & brand    & $2$ & $0.3104$ & $0.3231$ & $0.958$ & $1.221$ \ ($2.04$) \\
Homogeneous & \textsf{sieve} & category & $4$ & $0.2268$ & $0.2285$ & $0.962$ & $0.861$ \ ($1.43$) \\
\addlinespace
Heterogeneous & \textsf{gbr}   & unit     & $1$ & $0.4471$ & ---      & $0.938$ & $1.877$ \ ($3.13$) \\
Heterogeneous & \textsf{gbr}   & brand    & $2$ & $0.3449$ & $0.3161$ & $0.930$ & $1.432$ \ ($2.39$) \\
Heterogeneous & \textsf{gbr}   & category & $4$ & $0.2850$ & $0.2236$ & $0.893$ & $1.166$ \ ($1.94$) \\
\addlinespace
Heterogeneous & \textsf{sieve} & unit     & $1$ & $0.5039$ & ---      & $0.974$ & $2.003$ \ ($3.34$) \\
Heterogeneous & \textsf{sieve} & brand    & $2$ & $0.3832$ & $0.3563$ & $0.983$ & $1.506$ \ ($2.51$) \\
Heterogeneous & \textsf{sieve} & category & $4$ & $0.3249$ & $0.2520$ & $0.993$ & $1.250$ \ ($2.08$) \\
\bottomrule
\end{tabular}
\end{table}

For the evaluated portfolio and threshold, the result is negative (Table~\ref{tab:nivel}). Three readings are useful.

\paragraph{Coverage remains near or above nominal for one learner in this experiment.}
\textsf{sieve} covers $0.953$, $0.958$ and $0.962$ at the three levels of the
homogeneous scenario, and $0.974$ to $0.993$ in the heterogeneous one. Its
$\hat\tau$ declines close to the independent-design benchmark ($0.4569\to0.3104\to0.2268$ versus $0.3231$ and $0.2285$), providing a second simulation check of the covariance pattern in Section~\ref{sec:aggregation}. For \textsf{gbr}, the mean bias remains persistent as aggregation increases and coverage declines from $0.899$ to $0.665$ in the homogeneous scenario. In the heterogeneous scenario the decline is milder, while $\hat\tau$ also includes genuine effect heterogeneity.

\paragraph{The width remains above the pre-specified threshold.} The narrowest interval that covers
at all is \textsf{sieve} at category level, $0.861$, still $1.43$ times the
operational precision threshold. Of the twelve combinations none covers $0.90$ within a width of $0.6$; the closest approach in coverage, \textsf{sieve} at category
level in the heterogeneous scenario, covers $0.993$ at a width of $1.250$. The
arithmetic prediction of $0.636$ was optimistic for the reason anticipated
above: $\hat\tau$ among the six category-level units of this portfolio is
$0.2268$, not the $0.1585$ design dispersion of Table~\ref{tab:levels}, whose
category aggregate pools eight units rather than four.

\paragraph{In the heterogeneous scenario the interval is wider.} With truths that differ across categories, $\hat\tau$ falls more slowly
than $\sqrt{k}$ ($0.5039 \to 0.3832 \to 0.3249$ against $0.3563$ and $0.2520$),
because part of the between-unit component is genuine heterogeneity rather than design error. The resulting interval has coverage $0.993$ in the reported cell. Under the working model this is consistent with a conservative mixture of effect heterogeneity and design dispersion; it should not be interpreted as calibrated estimation of either component separately.

\begin{remark}[what would fit]
Extrapolating the $\sqrt{k}$ rate from the homogeneous \textsf{sieve} row, and
holding $s \approx 0.05$, a width of $0.6$ requires $\sqrt{s^2+\hat\tau^2} \le
0.153$, hence $\hat\tau \le 0.145$ and $k \approx (0.4569/0.145)^2 \approx 10$
units with independently realised price paths per published estimate. This
portfolio offers four. Whether an aggregate spanning roughly ten independently priced units remains aligned with the operational decision unit is application-specific. The extrapolation carries the caveat of
Remark~\ref{rem:extrap} --- it extends an observed rate one doubling beyond the
range measured.
\end{remark}

For the particular portfolio, learners, scenarios and width threshold studied here, no evaluated aggregation level simultaneously achieves the paper's coverage and precision criteria. This motivates examining the assignment mechanism itself, while stopping short of an impossibility claim for other estimators, portfolios or data-generating processes.

\section{Implications for data design}\label{sec:implications}

\paragraph{Adding rows is not equivalent to adding independent treatment variation.} When all regions face the same national list-price path, additional regions can still reduce outcome noise, help estimate nuisance functions, and improve precision conditional on that path. What they do not provide is an additional independent realisation of the price trajectory. Proposition~\ref{prop:agg} makes the relevant distinction explicit: the gain from aggregation is governed by the covariance of design-specific errors, not by the raw number of rows.

\paragraph{Longer panels help only through the variation they actually add.} In the simulation grid, larger $V=n_{\text{moves}}\times\text{magnitude}^2$ is associated with lower across-design dispersion. If the fitted relation \eqref{eq:scaling} were used descriptively, doubling $V$ would be associated with a reduction in $\sigma_b$ by a factor of about $2^{0.271}=1.21$. Because the exponent is a simulation regularity and $V$ is only a heuristic index, this arithmetic should not be used as a general forecast for extending a real panel.

\paragraph{Independent pricing paths can be more valuable than replicated exposure to one path.} In the simulated portfolio, product units with separately generated price trajectories display an approximately $\sqrt{k}$ reduction in design-specific dispersion. The key empirical requirement is low covariance of the relevant design errors; distinct product labels alone do not guarantee it. In an application, the covariance of price-setting rules, common cost shocks and synchronized promotions would need to be assessed before treating product histories as independent design draws.

\paragraph{Randomisation is one transparent way to create and justify independence.} Controlled regional price assignment can generate treatment variation whose assignment mechanism is known, thereby supporting design-based reasoning in a way that passive histories generally cannot. Under the illustrative assumptions of Proposition~\ref{prop:agg}, taking the baseline $\sigma_b=0.4815$ and imposing independent, mean-zero regional design errors would imply $0.4815/\sqrt{30}\approx0.088$ for an average across thirty regions. This is an algebraic illustration, not a simulated result for a thirty-region experiment. Natural experiments, staggered policy changes, or other sources of plausibly exogenous and weakly dependent price variation could play a similar identifying role. The central recommendation is therefore to create or exploit independent identifying variation, not that randomisation is the only admissible design.

\section{Limitations}\label{sec:limitations}

\paragraph{Synthetic evidence.} Every quantitative result is measured on a generator written for this study. This is useful because repeated price trajectories and design-specific ground truth are observable by construction, but it sharply limits external validity. The simulations establish that the proposed mechanism can be quantitatively important under the stated DGP; they do not establish that between-design dispersion dominates for every short pricing panel or every estimator.

\paragraph{The error decomposition must match the implemented ground truth.} The formal decomposition in Section~\ref{sec:decomposition} is defined for $e(D,u)=\hat\theta(D,u)-\theta(D)$. This matters because the generator allows the design-specific truth to vary through pass-through. Any implementation that computes $\sigma_b$ from $\operatorname{Var}_D\{\mathbb{E}_u[\hat\theta\mid D]\}$ without subtracting $\theta(D)$ would combine variation in the estimand with variation in estimation error. Reproducibility code should therefore report explicitly which object is used in the decomposition.

\paragraph{A heuristic variation index, not an information measure.} The index $V=n_{\text{moves}}\times\text{magnitude}^2$ is a useful low-dimensional summary of the simulation grid, but the paper does not derive it as Fisher information or prove that estimator dispersion must scale as a power of $V$. Consequently, the fitted exponent $-0.271$ is a descriptive simulation coefficient and extrapolations from \eqref{eq:scaling} are illustrative only.

\paragraph{Variance-component assumptions.} Interpreting Paule--Mandel's $\tau^2$ as design variance requires approximately independent or weakly dependent design draws, adequate within-unit variance estimates, exchangeability of the design-specific displacement and, for a clean interpretation, common unit-level truths. When true elasticities vary, $\hat\tau^2$ generally mixes genuine heterogeneity with design uncertainty. With only eight units in the main homogeneous exercise, uncertainty in $\hat\tau^2$ itself can also be material.

\paragraph{One estimator implementation.} The simulations use the estimator in \eqref{eq:dml}, including median aggregation of fold-specific residual-on-residual coefficients. This is DML-motivated but is not the canonical pooled DML2 estimator. The canonical DML theory also relies on regularity and dependence conditions that are not established here for the short serially dependent panel. The paper should therefore avoid attributing the simulation findings to Double Machine Learning as a class. Whether analogous between-design dispersion arises for instrumental-variable, structural, Bayesian, or alternative orthogonal-score estimators remains an empirical and theoretical question.

\paragraph{Calibration and decision thresholds.} The $W\le0.6$ precision threshold is a pre-specified operational criterion for the simulation exercise, not a universal standard for pricing decisions. Likewise, the identification filter is calibrated on the same generator and has limited transportability across configurations. These thresholds are useful for organizing the experiment but should not be interpreted as externally validated cutoffs.

\paragraph{External validation.} A natural next step is to compute the proposed variation summaries and dependence diagnostics on real price histories, ideally in settings where multiple quasi-independent or randomized price paths exist. Such evidence would be needed before interpreting the simulation magnitudes as representative of empirical pricing panels.

\section{Conclusion}\label{sec:conclusion}

This paper uses a synthetic pricing environment to separate two sources of inferential uncertainty that are easy to conflate in short panels. The first is variation conditional on one realised price trajectory. The second is variation in the estimator's conditional mean error across alternative trajectories generated by the pricing process. In the baseline DGP, the second component is quantitatively large: for the gradient-boosted specification it accounts for $97.6\%$ of the variance of the estimation error. The number is simulation-specific, but the decomposition clarifies why an interval based on one realised panel need not represent repeated-design performance.

The simulation evidence also refines the diagnosis of undercoverage. The moving-block bootstrap standard error is not simply too small relative to within-design dispersion; in fact it exceeds the measured $\sigma_w$ in the baseline decomposition. The more important feature is that the estimator's conditional centre varies across realised price histories. None of the eight within-panel constructions evaluated here reaches nominal coverage, although this finite comparison is not an impossibility theorem. The result motivates explicitly modelling or sampling the distribution of designs rather than treating a wider conventional standard error as a complete solution.

Two constructive findings follow. First, the gain from aggregation depends on the covariance of design-specific errors. Proposition~\ref{prop:agg} shows exactly why independent design draws produce a $\sqrt{k}$ reduction in standard deviation and why a common component limits that gain. The simulated product units are close to the independent benchmark, whereas additional exposure to a common price path does not create the same source of identifying variation. This is the precise sense in which adding observations is not equivalent to adding identification.

Second, a standard between-unit variance component can be useful when multiple independently realised price paths are available. In the homogeneous simulation, the Paule--Mandel-based variance augmentation raises empirical coverage from $0.469$ to $0.931$ for the posterior-centred construction, at the cost of much wider intervals. This exercise does not establish a new meta-analytic estimator. It shows that, under a random-bias working model, information across designs can represent a component that a single realised trajectory cannot identify nonparametrically. When true unit effects differ, the between-unit component mixes heterogeneity with design uncertainty and must be interpreted more cautiously.

The fitted relation $\hat\sigma_b\approx0.182V^{-0.271}$ provides a compact summary of the simulation grid, not a scaling law. Its role is descriptive: within the explored DGP, substantially more independent price movement is associated with lower across-design dispersion, but the decline is slow over the observed range. The broader empirical lesson is therefore about data design. If the inferential target requires performance across plausible future price histories, collecting more rows under essentially the same treatment path may be less valuable than generating additional, credibly independent price variation. Controlled regional randomisation is one transparent way to create such variation; other quasi-experimental sources may serve the same purpose when their assignment and dependence structure are defensible.

The paper's contribution is consequently best read as a shift in emphasis from ``which standard error should be attached to this one passive panel?'' toward ``what variation would make the elasticity identifiable and its uncertainty learnable at the decision level?'' The simulations suggest that, in sparse pricing regimes, better inference and better data design are closely linked. Establishing how far that conclusion travels beyond the present generator is the next empirical task.

\paragraph{Reproducibility.} The generator, the estimators, the thirty-five experiments and the pre-registered decision rules that produced every figure in this paper are published as an executable notebook. All results are obtained from synthetic data and no commercial information is used.


\clearpage