Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
101,807 characters · 22 sections · 47 citation commands
Calibrated Horizon-Weighted Local Projection Designs for Markov Switchbacks
\noindentKeywords: dynamic experimental design; local projections; switchback experiments; Markov assignment; optimum input design; demand response.
\noindentJEL classification: C18; C26; C44; C90; Q41.
Many economic experiments are dynamic. A demand-response signal can reduce electricity use immediately, trigger rebound later in the day, and reshape the load profile over subsequent periods. A platform pricing experiment can affect current participation, future search behavior, and congestion in later periods. A marketing intervention can have delayed effects that are invisible in a contemporaneous average treatment effect. In these settings, the object of interest is often a dynamic response curve rather than a single scalar effect.
Local projections provide a natural empirical language for reporting such responses. Researchers estimate horizon-specific effects $(g_0,\ldots,g_H)$, cumulative effects such as $\sum_{h=0}^H g_h$, or shape contrasts that distinguish immediate reductions from later rebound. Recent work has clarified the causal interpretation of local projections and related time-series objects: potential-system frameworks provide nonparametric content for time-series experiments, local projections, impulse responses, and SVARs CarlsonShephard2026, and policy-path local projection frameworks define dynamic counterfactuals under sequential interventions Wang2026.
We ask a complementary design question. If the researcher plans to report a dynamic response curve or a horizon-weighted functional of it, how should the treatment assignment path be designed? Existing switchback and time-series experimental-design papers largely target average effects, carryover robustness, covariate balance, or generic mean squared error. BojinovSimchiLeviZhao2023, for example, study optimal switchback designs under assumptions on carryover order and provide randomization-based inference. Recent time-series A/B testing work considers full-history adaptive designs optimized for generic treatment-effect MSE WuEtAl2026. These are natural objectives for many experiments, but they are not the same as optimizing precision for the multi-horizon local-projection object that the researcher ultimately reports.
We make the reported dynamic causal object the design target. We introduce the horizon-weighted local projection (HW-LP) criterion. Let $g=(g_0,\ldots,g_H)'$ be a design-invariant response curve and let $\theta_c=c'g$ be the scalar object of interest. The design risk is the asymptotic variance of $c'\hat g$ under the assignment mechanism. More generally, for a positive semidefinite matrix $W$, the risk is \[ \mathcal R_W(\pi)=\operatorname{tr}\{W V_\pi\}, \] where $V_\pi$ is the design-dependent covariance matrix of the estimated response curve. This criterion covers isolated horizon effects, cumulative effects, delayed-response windows, rebound contrasts, and full-curve reporting.
The main analytical results are developed for balanced Markov switchback designs. Let $A_t\in\{0,1\}$, $\Pr(A_t=1)=1/2$, and $\Pr(A_t=A_{t-1})=s$. Write $r=2s-1$, so that $r=0$ is iid randomization, $r>0$ is persistent switching, and $r<0$ is alternating switching. For $X_t=(A_t-1/2,\ldots,A_{t-H}-1/2)'$, the population information matrix is \[ Q_H(r)=\frac14 (r^{|i-j|})_{i,j=0}^H. \] In the homoskedastic finite-memory model, $V(r)=\sigma^2 Q_H(r)^{-1}$. Because the inverse of this AR(1)-Toeplitz matrix is tridiagonal, the HW-LP risk for the contrast $c'g$ has the closed form \[ \mathcal R_c(r)=4\sigma^2\frac{a+br^2-2dr}{1-r^2}, \] where $a=\sum_{j=0}^H c_j^2$, $b=\sum_{j=1}^{H-1}c_j^2$, and $d=\sum_{j=1}^{H}c_{j-1}c_j$. When $0<|d|<(a+b)/2$, the interior stationary point has the closed form \[ r_c^{\mathrm{int}}=\frac{(a+b)-\sqrt{(a+b)^2-4d^2}}{2d}. \] If $d=0$, iid assignment is optimal in the benchmark class; if $|d|=(a+b)/2$, the variance-only infimum is approached at the boundary and the feasible optimum is the nearest allowed endpoint.
The formulas imply that optimal switchback persistence is target-specific: isolated immediate effects favor iid assignment, smooth cumulative targets favor persistence, and field deployment should replace the closed-form benchmark with calibrated covariance selection when the benchmark assumptions fail. Within the benchmark Markov class, iid assignment is optimal for isolated horizon effects. Persistent treatment episodes are variance-favored for smooth cumulative effects. Alternating or moderate-persistence designs can be optimal for oscillating and rebound contrasts. The feasible set matters: at half-hour frequency, for example, $r=0.95$ implies an expected episode length of $2/(1-r)=40$ half-hours, while $r=0.7$ implies 6.67 half-hours. When carryover may extend beyond the estimated horizon, excessive persistence can also increase omitted-lag bias. We therefore add a bias-augmented sensitivity criterion that trades off target variance against worst-case omitted-carryover bias without claiming to be an exact finite-sample MSE under every omitted-tail data-generating process.
Figure (ref) summarizes how the reporting target changes the variance-minimizing persistence.
We make three contributions. First, we formulate a horizon-weighted LP design criterion for a pre-specified reduced-form reporting object and show why path adjustment fixes that object while naive LP coefficients are design-dependent projections. Second, we derive a closed-form Markov benchmark for binary switchbacks and explain its frequency-domain interpretation, sparse active-share variants, and scope limits. The AR(1)-Toeplitz precision calculation is inherited from optimal input design; our contribution is the LP reporting-object translation, the binary Markov feasible class, and the operational mapping from target shape to persistence. Third, we turn the benchmark into a calibrated implementation rule by adding HAC and residualized covariance selection, pilot-regret bounds, local-bias carryover sensitivity, randomization-first near-boundary inference, and semi-synthetic covariance stress tests. The closed form is used as an interpretable benchmark for the persistence margin; field recommendations are based on calibrated covariance and randomization diagnostics when residual dependence, calendar controls, or feasibility constraints depart from the balanced homoskedastic Markov benchmark.
Section (ref) defines the dynamic response targets and path-adjusted local projections. Section (ref) derives the HW-LP criterion and closed-form optimal persistence, with extensions in Appendix (ref). Section (ref) studies omitted-carryover misspecification. Section (ref) presents estimation, design selection, and inference, with details in Appendix (ref). Section (ref) reports the empirical calibration, and Section (ref) discusses interpretation and external validity.
The analysis relates to six literatures. The first is the literature on local projections and dynamic causal response objects. Local projections were introduced by Jorda2005 as a flexible way to estimate impulse responses. Subsequent work clarified that LPs and VARs share the same population impulse-response estimands under unrestricted lag structures, while LP inference can remain robust in persistent settings and at long horizons PlagborgMollerWolf2021,MontielOleaPlagborgMoller2021. Other work studies the efficiency and finite-sample behavior of LP estimators, including comparisons with VAR-based impulse responses, smoothing across horizons, and practical inference procedures KilianKim2011,BarnichonBrownlees2019,InoueJordaKuersteiner2026. Recent causal foundations show when time-series regressions, local projections, and policy-path responses can be interpreted as dynamic causal effects rather than purely predictive objects RambachanShephard2025,CarlsonShephard2026,Wang2026. Related work on nonlinear environments shows that linear LPs with observed shocks or proxies can still identify weighted averages of causal effects, which is useful for interpreting response curves outside fully linear models KolesarPlagborgMoller2025. Macroeconomic applications using LP-style response curves further illustrate why researchers often care about cumulative, delayed, and state-dependent dynamic effects rather than a single contemporaneous coefficient Ramey2016,AuerbachGorodnichenko2013,RameyZubairy2018.
The second literature studies randomized experiments and temporal assignment. General potential-outcome treatments of randomized experiments emphasize that the estimand and assignment mechanism should be specified jointly ImbensRubin2015,AtheyImbens2017. BojinovShephard2019 extend potential-outcome reasoning to single time-series experiments and exact randomization tests, while BojinovRambachanShephard2021 develop finite-population dynamic causal effects for panel experiments. Switchback designs with carryover have been analyzed by BojinovSimchiLeviZhao2023, while minimax designs under habituation and clustered switchbacks under spatiotemporal interference are studied by BasseDingToulis2023 and JiaKallusYu2023. More recent work connects regression-based and design-based inference, balances lagged outcomes and covariates through sequential rerandomization, and uses full-history reinforcement learning to optimize generic treatment-effect MSE LinDing2025,ZengEtAl2026,WuEtAl2026. The design objective differs from contemporaneous average-effect, balance-score, and global-MSE criteria: it optimizes the covariance of the horizon-weighted response object that the researcher plans to report.
A related set of papers studies interference and exposure mappings in randomized experiments. Temporal carryover in this paper can be viewed as within-unit interference over time, so the exposure-mapping perspective is conceptually useful. Work on two-stage randomized designs, general interference, and unknown interference formalizes how treatment assignment can affect outcomes through exposure paths rather than through a single contemporaneous treatment indicator HudgensHalloran2008,AronowSamii2017,SavjeAronowHudgens2021. Randomization tests under network interference also show how design-based inference changes when the null hypothesis is not sharp under the original assignment mechanism AtheyEcklesImbens2018. The framework here is narrower: the exposure path is temporal, the response object is a local-projection target, and the main design margin is assignment persistence.
The third literature is optimal input design and optimal experimental design for dynamic systems. Classical optimal design gives general criteria for choosing informative experiments, including equivalence and optimality ideas that underlie many information-matrix calculations Kiefer1959,Fedorov1972,Pukelsheim2006. Dynamic input-design work already shows that the input spectrum and autocorrelation shape the covariance of dynamic-parameter estimators; early and textbook treatments include Mehra1974, GoodwinPayne1977, Zarrop1979, and Ljung1999. The AR(1)-Toeplitz precision calculation used below is also a standard Gaussian-Markov precision fact, familiar from state-space and GMRF treatments RueHeld2005. These precision-matrix results are standard. The contribution here is narrower: we translate the OID logic into a binary switchback problem whose target is a pre-specified LP reporting object rather than a transfer-function parameter; whose feasible class is constrained by run length, active share, and verifiability; and whose field recommendation can replace the homoskedastic information matrix by the residualized covariance of the estimator actually used. Thus the closed-form calculation below is an inherited precision-matrix benchmark, while the HW-LP contribution is the switchback-specific reporting, feasibility, and calibrated-implementation layer.
Applied design texts and Bayesian or nonlinear-design reviews emphasize that the target, loss function, and prior or model class determine the optimal design AtkinsonDonevTobias2007,ChalonerVerdinelli1995,PronzatoPazman2013. Our calculations are adjacent to that literature, but the reported LP contrast is itself the object: it can be defined without specifying a full dynamic outcome model and remains meaningful when the design covariance is estimated by HAC, pilot, realized-schedule, or residual-bootstrap methods. This distinction is why the Toeplitz calculation is framed as a transparent benchmark and not as a claim of global dominance over continuous-valued, multisine, adaptive, finite-sequence, or fully optimized input designs.
The fourth literature is target-specific and data-driven switchback design. Empirical-Bayes and prior-data approaches study how carryover, seasonality, autocorrelation, and simultaneous experiments affect switchback MSE XiongChinTaylor2024. Sequentially rerandomized switchbacks use lagged outcomes and covariates to improve balance over time ZengEtAl2026. Online experimentation work highlights the operational importance of logging, randomization, long-run metrics, and platform constraints in large-scale A/B tests KohaviLongbothamSommerfieldHenne2009,KohaviTangXu2020,BakshyEcklesBernstein2014. Long-term metric design in online platforms is especially close in spirit to the idea that the reported object should shape the experiment HohnholdOBrienTang2015. The HW-LP criterion uses the same broad principle, but the tuning target is a pre-specified response curve, cumulative effect, delayed window, or rebound contrast.
The fifth literature is residential electricity demand response and dynamic pricing. The Low Carbon London data were produced by the United Kingdom's first residential dynamic time-of-use electricity-pricing trial and include smart-meter readings and tariff information StrbacEtAl2024,Schofield2015. Reviews and field experiments show that households respond to time-varying prices, but the magnitude and timing of responses depend on tariff design, enabling technologies, and event duration FaruquiSergici2010,NewshamBowker2010,Wolak2011. Behavioral and informational interventions also generate dynamic response patterns, including persistence, habituation, and rebound-like behavior Allcott2011,AllcottRogers2014,ItoIdaTanaka2018. Additional studies of marginal-price perception, time-of-use tariff participation, and distributional effects emphasize that load shifting is shaped by information, salience, and household constraints Ito2014,Ozaki2018,YunusovTorriti2021. The Low Carbon London exercise is not a causal estimate of the observed tariff; it uses realistic high-frequency load dynamics to evaluate how different assignment paths perform when the response path is known.
Finally, we contribute to the broader literature on experiment planning under multiple possible reporting targets. Researchers often enter a dynamic experiment knowing that they may report an immediate effect, cumulative effect, delayed window, or full response curve. The target-menu criterion introduced below provides a simple way to pre-specify this uncertainty without changing the estimand after observing the data. In that sense, the contribution is not another post-estimation adjustment for local projections, but a design-stage rule that ties the randomization path to the dynamic objects that will be reported.
Potential-outcome notation defines the causal object before the LP design criterion is introduced, because the assignment path is dynamic and temporal carryover is a form of within-unit interference over time.
Let $a_{1:T}$ denote a candidate binary assignment path and let $Y_t(a_{1:T})$ be the potential outcome at time $t$. Our main target is an aggregate or single-series dynamic response, not a household-level heterogeneous-treatment-effect parameter. In applications with many households, $Y_t(a_{1:T})$ is the aggregate potential outcome under the common tariff or platform policy path. This convention rules out cross-sectional interference only if the target is an individual-level effect. If network, feeder, or congestion spillovers are part of the aggregate load response, they are included in the aggregate potential outcome; we then design for that aggregate response rather than for a no-spillover household estimand. This is a scope choice, not a proof that household-level interference is irrelevant. If the scientific object is an individual effect, a spillover effect, or an exposure-specific effect, the design should be expanded to a two-stage or clustered switchback design and an explicit spatial-temporal exposure mapping HudgensHalloran2008,AronowSamii2017,SavjeAronowHudgens2021,Leung2022.
A further aggregation condition is required for design invariance. If the observed aggregate is a fixed linear average of household potential outcomes under the common policy path, then the aggregate response vector is the corresponding average of household response paths and does not change with the Markov persistence used to estimate it. If compliance, opt-out behavior, appliance automation, or missingness makes the set of effective contributors depend on the assignment path, the aggregate estimand becomes design-weighted. The HW-LP criterion still optimizes precision for the reported aggregate object, but the object should then be described as a policy-path response for the induced compliance regime rather than as a simple population average of household elasticities. This is the Jensen-style aggregation caveat: nonlinear participation and heterogeneous compliance can move the estimand when the design changes.
Assumption (ref) separates identification from design. The potential-outcome restrictions define the response vector. The Markov assignment mechanism then determines the information matrix for estimating that vector. The response vector in the aggregate design is an aggregate reduced-form path response: it is the response of the chosen aggregation unit to a common assigned path. It is not an individual treatment effect, a structural price elasticity, a welfare primitive, or a household-level spillover estimand. If the scientific target is household elasticity, network spillover, feeder-level exposure, or structural welfare, the exposure mapping and design must be changed before applying the HW-LP criterion. The main asymptotic theory is written in a stationary super-population language because it yields compact covariance formulas. A finite-population or fixed-schedule reading is also possible: condition on the potential-outcome schedule and evaluate the randomization distribution induced by the assignment path. The realized-schedule diagnostic in Proposition (ref) uses exactly this finite-design reading. The Low Carbon London exercise is semi-synthetic: it fixes the observed baseline load path, injects known response paths, and randomizes assignments to evaluate design risk. It therefore validates the variance and covariance arithmetic under realistic baseline dynamics, not causal recovery of an unknown historical tariff response.
Consider a single time-series experiment with binary assignment $A_t\in\{0,1\}$. Under Assumption (ref), the initial analytical setting can be written as the finite-memory additive response model
The vector $g=(g_0,\ldots,g_H)'$ is the dynamic response curve. Throughout the analytical section, deterministic controls, intercepts, calendar effects, and pre-treatment covariates are treated as partialled out. Equivalently, all regressions below can be read after applying a Frisch--Waugh--Lovell residualization step to the outcome and to the lagged assignment vector. A researcher may care about the full curve or about a functional
The contrast $c=e_0$ targets the contemporaneous effect; $c=\mathbf 1$ targets the cumulative effect; weights concentrated on medium-run horizons target delayed effects; sign-changing weights target rebound or shape contrasts.
The notation below uses $Q_H$ for the assignment information matrix, $V_H$ for the covariance of the path-adjusted LP estimator, $\mathcal R$ for the feasible persistence set, $p_0$ for an active-share budget, and $r$ for assignment persistence.
The key distinction is between the causal target and the assignment design. The target $g$ or $c'g$ is fixed. The assignment mechanism changes the information available about that target.
The terminology “local projection target” refers to the response object the researcher plans to report: a vector of horizon-specific effects or a weighted functional of that vector. In the finite-memory additive model, the path-adjusted local projection and the distributed-lag representation estimate the same design-invariant response components. We use the distributed-lag form in the analytical benchmark because it exposes the assignment information matrix and yields closed-form design rules. In more general nonlinear, state-dependent, or heteroskedastic environments, the same design principle applies to the covariance matrix of the chosen local-projection estimating equations: \[ \hat r_W=\arg\min_{r\in\mathcal R}\operatorname{tr}\{W\hat V(r)\}, \] where $\hat V(r)$ can be estimated from a pilot experiment, residual bootstrap, or calibrated baseline model. Thus the finite-memory results are not meant to exhaust all local-projection settings; they provide a transparent benchmark and comparative statics for choosing assignment persistence. The reported scalar object is $\theta_c=c'g$, where $c$ encodes the pre-specified horizon weights. Naive LP coefficients and path-adjusted coefficients coincide only under finite-memory restrictions that remove future-assignment mixtures; outside that benchmark, the calibrated implementation replaces the closed-form covariance by $\widehat V(r)$ for the estimator actually used.
The first requirement for a design criterion is that the estimand does not change when the assignment persistence changes. In the finite-memory benchmark, the path-adjusted LP satisfies this requirement. Let $X_t=(\widetilde A_t,\ldots,\widetilde A_{t-H})'$ denote the centered lag vector and let $Q_X=E[X_tX_t']$.
A naive local projection of $Y_{t+h}$ on $A_t$, by contrast, can mix the effect of the initial assignment with the effects of subsequent assignments when the design induces persistence. The following proposition makes the source of the problem explicit.
The proposition does not imply that the naive LP coefficient is inconsistent for its projected estimand. Rather, the projected estimand itself changes with assignment persistence. This leads to a simple reporting rule. If the pre-analysis plan defines a design-invariant dynamic response component, cumulative effect, or shape contrast, the analysis should use the path-adjusted LP or the equivalent distributed-lag regression. If the scientific object is instead the projection induced by the implemented policy path--for example, the average response to the whole bundled tariff path rather than to a single initial assignment--then the naive LP can be the appropriate estimand, but the estimand should be described as design-specific. Our design problem is the first case: the response object is fixed before choosing assignment persistence.
Proposition (ref) and Lemma (ref) are the reason the design target is defined through path-adjusted local projections. The goal is not to optimize precision for a response object that changes with the assignment persistence, but to fix the dynamic response component and then choose the assignment path that estimates it well. To isolate dynamic response components, we use the path-adjusted local projection, equivalently the distributed-lag representation
In richer models, the horizon-$h$ regression can be written as
where future assignments and history controls adjust for design-induced path correlation. The finite-memory representation in (ref) is the workhorse for the closed-form design results below.
We focus first on balanced Markov switchback designs:
Define $r=2s-1$ and $\widetilde A_t=A_t-1/2$. The parameter $r$ is the assignment persistence: $r=0$ is iid Bernoulli assignment, $r>0$ favors longer treatment episodes, and $r<0$ favors alternation. Designs are chosen from a feasible interval $\mathcal R=[\underline r,r_{\max}]\subset(-1,1)$, which can encode minimum dwell-time, fatigue, safety, switching-cost, or operational constraints. Since $s=(1+r)/2$, the expected treatment-episode length is \[ E[\text{run length}]=\frac{1}{1-s}=\frac{2}{1-r}. \] This translation is useful in applications: at half-hour frequency, $r=0.7$ corresponds to about 3.33 hours, while $r=0.95$ corresponds to about 20 hours. Let $X_t=(\widetilde A_t,\widetilde A_{t-1},\ldots,\widetilde A_{t-H})'$.
This unbalanced extension is important in demand-response settings, where high-price or event signals are often sparse. Sparse assignment shares do not change the Toeplitz persistence logic, but they reduce information by the factor $p(1-p)$ and make active-run-length constraints explicit.
For sparse demand-response signals, strong negative persistence may be infeasible because the chain cannot alternate frequently while keeping a very small active share. The lower bound $r_{\min}(p_0)$ is therefore not a technicality: it rules out assignment patterns that would violate the event budget. This matters most for sign-changing or oscillating targets, for which the unconstrained benchmark may prefer negative persistence. If the active share is $p_0=0.10$, for example, the Markov feasibility lower bound is only $r_{\min}=-0.111$, so strongly alternating designs are unavailable regardless of their variance ranking in the balanced benchmark. Positive persistence is usually feasible, but active- and inactive-spell caps may impose upper bounds below one.
Table (ref) reports the information scale and active-run constraints for representative active-share budgets.
Population information scaling does not by itself describe the finite-path risk of a sparse experiment. Appendix Table (ref) therefore simulates the cumulative-target design at the three-hour active-spell cap using the exact realized lagged-assignment Gram matrix. With $T=2{,}000$ and $p_0=0.05$, the median risk ratio is 1.034 and the p90 ratio is 1.817, while the p10 path contains only 12 active spells; at $T=17{,}520$ the p90 ratio falls to 1.197. None of the simulated Gram matrices is singular. Thus sparse short experiments should screen candidate randomization paths before launch even when the population Markov information matrix is nonsingular.
Proof. See Appendix (ref).
The clean covariance in Proposition (ref) is a benchmark. In applications with deterministic calendar controls or other pre-specified residualization variables $Z$, the relevant finite-sample information is not the raw Toeplitz matrix but the residualized matrix \[ Q_{H,Z}(r)=\operatorname*{plim}_{T\to\infty}T^{-1}X_H(r)'M_ZX_H(r),\qquad M_Z=I-Z(Z'Z)^{-1}Z'. \] The closed-form $Q_H(r)$ rule is exact for the no-control homoskedastic benchmark and remains useful as a persistence diagnostic. The calibrated implementation below instead evaluates the residualized design matrix, or its bootstrap/HAC analogue, for the realized or pilot-calibrated schedule. Thus calendar residualization is not treated as innocuous: it is precisely one reason we recommend switching from the analytical rule to the calibrated selector when finite-sample covariance differs materially from the Toeplitz benchmark.
The following proposition records the primitive conditions used for the calibrated implementation. It is stated for a compact persistence set and a standard Bartlett/Newey--West estimator; other kernels can be used under the usual moment and bandwidth restrictions familiar from data-dependent HAC procedures Andrews1991,AndrewsMonahan1992.
The proposition provides the baseline form of the calibrated rule used below. For a finite grid of candidate persistences, the same conclusion requires only pointwise HAC consistency for each grid point. For a continuous persistence interval, the Lipschitz and mixing conditions give the stochastic equicontinuity needed to pass from pointwise consistency to uniform consistency. The uniform statement is for a fixed compact interior set; it is not a claim that the same bandwidth works as $r_{\max}$ drifts to one. If a sequence of feasible sets has $r_{\max,T} \uparrow 1$, the truncation lag must cover the growing dependence length, for example by requiring $b_T(1-r_{\max,T})\to\infty$ and $b_T/T\to0$. This also makes clear why the effective sample size is closer to $T_{\mathrm{eff}}(r)=T(1-r)/(1+r)$ than to $T$ when $r$ is large.
In implementation we therefore use a candidate-specific bandwidth diagnostic rather than a single iid-oriented bandwidth. The starting value is \[ b_{0T}(r)=\left\lceil 1.3 T^{1/3}\frac{1+|r|}{1-|r|}\right\rceil, \] and the finite-sample default rounds the capped value to the nearest integer according to \[ b_T(r)=\max\!\left\{1,\left\lfloor \min\{b_{0T}(r),\,T/4,\,0.25T_{\mathrm{eff}}(r)\}+\frac12 \right\rfloor\right\}, \] with sensitivity checks at one half and twice the capped value. If the cap binds, the report treats the design as a high-persistence finite-sample warning rather than as evidence that a very long HAC window is reliable. This formula is not an optimal-bandwidth theorem; it is a conservative operational rule that expands the lag window as assignment persistence approaches the boundary while keeping the truncation lag close to one quarter of the effective sample or below. The effective-sample-size column is likewise a scalar AR(1) variance-inflation diagnostic, not an exact sample size for every vector LP contrast; its role is to flag when persistence makes iid-oriented HAC choices unreliable. Figure (ref) reports the scale for the LCL sample length.
The LCL calibrated implementation also examines sensitivity to the HAC truncation lag. Table (ref) recomputes the residualized aggregate cumulative-target selector under bandwidths equal to one quarter, one half, one, two, and four times the capped candidate-specific rule. The selected persistence remains at the upper fine-grid endpoint in all five cases, while the relative risk of iid assignment rises from 3.43 to 4.36 as the longer bandwidths emphasize persistent low-frequency residual variation. This small diagnostic supports the operational interpretation of the bandwidth rule: the cap is a finite-sample safeguard, and the substantive calibrated recommendation is not driven by a single truncation lag.
The design criterion for a contrast $c'g$ is $\mathcal R_c(r)=c'V_H(r)c$. In the homoskedastic benchmark,
Define $a=\sum_{j=0}^{H}c_j^2$, $b=\sum_{j=1}^{H-1}c_j^2$, and $d=\sum_{j=1}^{H}c_{j-1}c_j$. Then
The boundary case in Proposition (ref) is substantive rather than pathological. For cumulative targets, the variance-only benchmark will always choose the largest admissible persistence, so the definition of $r_{\max}$ is part of the design. We therefore set $\mathcal R$ before evaluating the target rule by translating operational constraints into persistence bounds. For balanced designs, an upper bound on expected run length $L_{\max}$ gives $r_{\max}\le1-2/L_{\max}$. For unbalanced event designs, active and inactive caps give $r_{\max}\le1-\{(1-p_0)L_{1,\max}\}^{-1}$ and $r_{\max}\le1-\{p_0L_{0,\max}\}^{-1}$. A minimum expected number of active spells $N_{1,\min}$ over a planned length $T$ gives the additional diagnostic constraint \[ T p_0(1-p_0)(1-r)\ge N_{1,\min}. \] When a target lands on the boundary, we report the endpoint choice and treat sensitivity to a tighter endpoint as a design diagnostic, not as a new estimand.
The scalar contrast criterion is the leading design object, but the same covariance logic applies to matrix-weighted curve losses, transformed response objects, and finite candidate menus. Appendix (ref) develops these extensions in detail. It gives the frequency-domain interpretation, operational thresholds, numerical-stability triggers, higher-order and non-Markov benchmarks, finite-sequence coordinate-exchange designs, target-menu robust rules, and links to classical information-matrix criteria. The central implementation distinction is unchanged: the Toeplitz formula explains the target-shape margin, whereas field deployment should compare the covariance of the estimator and reporting object over the feasible design menu.
The appendix also quantifies when the first-order Markov benchmark is adequate. Finite-sequence diagnostics compare population and realized information matrices, target-menu calculations report the cost of using a common persistence across several reporting objects, and spectral diagnostics identify high-persistence, large-horizon configurations that require regularization or a narrower target menu. These diagnostics qualify the benchmark rather than alter the core closed-form result.
A persistent design can be variance-efficient for a smooth cumulative target, but persistence also correlates nearby treatment lags. This creates a second risk: if the analyst estimates a model with maximum horizon $H$ but the true carryover extends beyond $H$, omitted treatment lags can bias the estimated dynamic response. In geometric-tail benchmarks, this omitted-carryover bias is governed by the buffer horizon $K$, not by the assignment persistence: the cumulative omitted-tail bias is essentially invariant to $r$ over the persistence range, while persistence reallocates variance across horizons and the buffer horizon controls bias. The buffer controls how much unmodeled tail can enter the reported contrast.
Let $X_t=(\widetilde A_t,\ldots,\widetilde A_{t-H})'$ collect the included lags and let $Z_t=(\widetilde A_{t-H-1},\ldots,\widetilde A_{t-H-K})'$ collect $K$ omitted lags. Suppose \[ Y_t=X_t'g+Z_t'\gamma+\varepsilon_t, \] but the analyst estimates the truncated model using only $X_t$.
The formula has two interpretations. If the omitted tail is fixed, the expression in Proposition (ref) is a population bias diagnostic: it measures how sensitive the reported contrast is to carryover beyond the estimated horizon. If the omitted tail is local, $\gamma=\delta/\sqrt T$ with $\|\delta\|_2\le M_0$, then it yields a dimensionally consistent scaled-MSE design criterion,
Equivalently, $T\,E\{c'(\hat g-g)\}^2$ is approximated by (ref) under the local-bias sequence. The local sequence is used because it puts variance and misspecification sensitivity on the same asymptotic scale; a fixed nonlocal tail would dominate the variance term and turn the design problem into pure worst-case bias avoidance. In applications $M_0$ should be pre-specified. We use three defaults: $M_0=0$ for the variance-only benchmark, a pilot-calibrated value based on the Euclidean norm of the last observed or fitted response-tail block, and a sensitivity grid such as $M_0\in\{0,0.5,1,2\}$ times that pilot norm. If no pilot tail is credible, the recommended report is the frontier over $M_0$ rather than a single robust design. For a fixed nonlocal tail, we report the second term as a bias sensitivity rather than as a finite-sample MSE component. The closed-form tail expression shows which targets are most vulnerable: in the Markov benchmark, omitted carryover beyond the estimated horizon affects the reported contrast only through the last included reporting weight $c_H$.
Corollary (ref) gives a practical robustness rule: estimate beyond the horizons that will be reported. Buffer horizons can absorb residual carryover while leaving the reported contrast unchanged. They are not free: increasing the estimated horizon enlarges the nuisance part of the regression and can raise finite-sample variance through the inverse design matrix. The default procedure is to start with $H=K_0+2$, compute the calibrated target variance and the local-bias term, and then increase $H$ until the bias reduction is small relative to the induced variance increase. Under these empirical defaults, the search stops when the target variance rises by more than 10 percent or when the condition number of the residualized assignment information matrix exceeds the pre-specified cap. When $M_0=0$, (ref) reduces to the variance-only HW-LP criterion. When $M_0>0$, high persistence can become less attractive because it increases the correlation between included and omitted treatment lags.
Figure (ref) reports the resulting bias--variance tradeoff.
Implementation and inference require a reporting target, feasible design menu, and estimator class to be specified in advance. Pointwise inference, simultaneous full-curve reporting, post-selection inference, and estimator choice address distinct objects: fixed-target coverage, curve-level coverage, and design-selection uncertainty. The benchmark comparisons evaluate the design criterion without introducing a separate estimator. The benchmark DGP is \[ Y_t=\sum_{\ell=0}^{H}g_\ell A_{t-\ell}+u_t, \] with $H=8$. For scalar targets, the reported metric is the benchmark risk for $c'g$ under the relevant assignment persistence. This avoids comparing the same design under independent simulation draws: whenever the HW-LP rule selects $r=0$, it is displayed as the same iid design.
The analytical risk curves and benchmark tables are deterministic functions of the design covariance formulas. The Low Carbon London calibration uses 10,000 Monte Carlo replications per scenario-design pair and common random numbers across designs within each scenario. The replication count was fixed ex ante.
Table (ref) reports relative target risk for five scalar response objects. The table focuses on scalar targets and excludes full-curve reporting, which is a matrix-weighted objective rather than a scalar contrast. Full-curve reporting is evaluated separately below with $W=I$.
The benchmark results align with the theory. Isolated horizon effects favor balanced iid assignment. Smooth cumulative targets favor high persistence. Delayed and rebound-shaped targets favor intermediate or target-specific persistence. Oscillating contrasts favor negative persistence.
The closed-form Markov results use a homoskedastic benchmark to expose the assignment information matrix. In applications, serial dependence and calendar controls should be handled through the estimated covariance matrix of the residualized path-adjusted local-projection moments. The feasible implementation is therefore to compute $\hat V(r)$ from a pilot, HAC calculation, residual bootstrap, or calibrated baseline model and choose $r$ by minimizing $\operatorname{tr}\{W\hat V(r)\}$ over the feasible design class. The Low Carbon London evaluation below has two layers. The main scalar-target table reports the retained closed-form benchmark comparator and finite-candidate performance under realistic load dynamics. The delayed-reduction case, where the closed-form rule is not the retained best design, is treated as evidence that calibrated covariance selection is needed when residualized finite-sample covariance differs from the homoskedastic Toeplitz benchmark. The empirical evaluation retains the full response vector in each Monte Carlo draw, estimates design-specific covariance matrices, computes paired Monte Carlo standard errors, and evaluates observed tariff active indicators as fixed assignment schedules.
The design criterion is an ex ante precision criterion, but it maps directly into standard inference for the reported target. Once the experiment is run under a pre-specified design, the contrast estimator satisfies \[ \sqrt T\{c'(\hat g-g)\}\Rightarrow N(0,c'V_\pi c), \] under the same score-covariance conditions used for the design calculation. A confidence interval for $c'g$ is therefore \[ c'\hat g \pm z_{1-\alpha/2}\sqrt{c'\widehat V_\pi c/T}, \] with $\widehat V_\pi$ computed using the residualized HAC, bootstrap, or pilot-calibrated covariance appropriate for the realized assignment.
The variance interpretation should match the empirical design. The theoretical $V_\pi$ is a super-population long-run covariance for the stationary assignment-and-score process. A fixed-schedule diagnostic, by contrast, conditions on the realized assignment path and measures information in that finite design matrix. The Low Carbon London exercise is intermediate: the baseline load path is fixed from the observed data, while randomized assignments and injected responses are used to compare designs. In all cases the HW-LP rule optimizes the covariance of the estimator that will be used for inference, but the source of randomness should be stated in the pre-analysis plan and in the table notes. The design problem is thus separated from the inference problem: HW-LP chooses the assignment path to improve the covariance of estimators that are then reported with standard confidence intervals, bands, or design-based diagnostics.
Randomization-based inference is the primary finite-sample route in near-boundary regimes when the design produces few assignment episodes. Because the researcher controls the assignment law, a sharp-null Fisher test can be implemented by repeatedly drawing $A_{1:T}^{\ast}$ from the locked switchback design, recomputing the chosen statistic $\widehat{\theta}_c(A^{\ast},Y^{\mathrm{obs}})$, and comparing it with the observed statistic BojinovShephard2019,AtheyEcklesImbens2018. Under the sharp null and the pre-specified exposure mapping, this test is exact conditional on the observed outcomes. Neymanian finite-population intervals or conservative randomization variances can also be reported when the target is an average over a fixed schedule of potential outcomes BojinovRambachanShephard2021,LinDing2025. For $r>0.95$, a binding HAC cap, or an active-spell lower tail below the pre-specified threshold, the pre-analysis plan should make the Fisher or Neymanian randomization analysis primary and treat normal/HAC intervals as descriptive. Finite-sample covariate balance should also be logged for persistent designs. A simple diagnostic is the maximum standardized treated-control imbalance over pre-specified calendar covariates, such as hour-of-day, weekday, weekend, holiday, and season indicators; sequential rerandomization or rejection sampling can be added as a feasibility constraint when this imbalance exceeds a pre-specified threshold.
The inference route depends on the realized design diagnostics. For interior Markov designs with many active spells and a non-binding HAC cap, fixed-design HAC intervals or simultaneous bands are the default. When persistence exceeds 0.95, the HAC cap binds, or the active-spell lower tail falls below the pre-specified threshold, the primary finite-sample report should be a Fisher randomization test or a Neymanian randomization interval under the locked assignment law. Pilot-selected designs use fixed-design inference conditional on the selected design, with pilot uncertainty reported separately. If the same outcomes are used for tuning and inference, the protocol should use sample splitting, a simultaneous multiverse band, or selection-adjusted intervals. The pre-analysis plan should record the candidate grid, bandwidth rule, randomization draws, balance diagnostics, model menu, and family-wise error rule for the chosen route.
The fixed-$r$ theory does not provide local-to-unity inference guarantees. If the selected persistence is close to one, the effective number of independent assignment episodes can be small even when $T$ is large. This problem is analogous in spirit to persistent-regressor inference concerns in time-series LPs, although here the persistence comes from the assignment process rather than from the outcome dynamics Phillips1998,Mikusheva2007,PesaventoRossi2007. Table (ref) reports a finite-sample diagnostic under the homoskedastic Markov benchmark with $T=17{,}520$, $H=8$, cumulative target weights, and 2,000 Monte Carlo assignment draws. Under this ideal benchmark, conditional-assignment and Markov plug-in intervals remain close to nominal coverage, but the active-spell count falls sharply: at $r=0.99$ the fifth percentile is only 36 active spell starts. The operational implication is therefore not a new asymptotic correction, but a pre-analysis requirement: if $r>0.95$ or the capped bandwidth binds, the protocol should report a design-specific coverage, balance, and active-spell simulation before field deployment.
A scalar target can use the pointwise covariance formula above, but a reported response curve requires a simultaneous band. We use a Gaussian or residual-multiplier sup-$t$ critical value based on the estimated joint covariance, with Bonferroni as a conservative fallback when that covariance is unstable. Appendix (ref) gives the construction, finite-sample coverage diagnostics, power and minimum-detectable-effect calculations, sequential-monitoring cautions, and post-selection conditions.
The confirmatory design, target, horizon, covariance estimator, and tuning rules should be fixed using pilot information or a separated training sample. When the same confirmatory outcomes are used to select among designs or tuning parameters, valid reporting requires sample splitting, simultaneous inference over the pre-specified menu, or an explicit post-selection procedure. The appendix states these requirements and separates them from the estimator-efficiency comparison that follows.
The HW-LP criterion optimizes the design for a specified estimator and reporting object. The closed-form formulas use the path-adjusted least-squares LP because it is transparent, design-invariant under Assumption (ref), and robust to misspecifying a full dynamic outcome model. They are not a semiparametric Cram\'{e}r--Rao bound over all possible estimators.
If the researcher is willing to specify more structure, the same design logic applies with a different estimator-specific covariance matrix. With a known serial-error covariance $\Sigma$, for example, the feasible GLS distributed-lag estimator has covariance proportional to $(X'\Sigma^{-1}X)^{-1}$ rather than the OLS/HAC covariance $Q^{-1}\Omega Q^{-1}$. With a correctly specified VAR or transfer-function model, system estimation can dominate unrestricted LPs in variance, while unrestricted LPs can have lower specification bias when the model is wrong PlagborgMollerWolf2021,LiPlagborgMollerWolf2024. Smooth, Bayesian, ridge, or other regularized LP estimators similarly replace $\widehat V(r)$ by the covariance or posterior risk of the regularized estimator BarnichonBrownlees2019,FerreiraMirandaAgrippinoRicco2025. When a ridge-regularized estimator is used to stabilize an ill-conditioned design, the relevant object is the regularized target MSE, not variance alone: $\operatorname{tr}\{W\operatorname{MSE}(r,\lambda)\}=\operatorname{tr}\{WV(r,\lambda)\}+b(r,\lambda)'Wb(r,\lambda)$, where $b(r,\lambda)$ is the regularization bias for the reported object. The HW-LP selector should then minimize this pre-specified $(r,\lambda)$ MSE grid; the closed-form variance rule is the $\lambda=0$ benchmark. Thus a better estimator under iid assignment could dominate path-adjusted OLS under a persistent assignment for some data-generating processes. The contribution here is narrower: once the reporting object and estimator class are specified, HW-LP gives the assignment rule that minimizes that estimator's target risk within the candidate design class.
Table (ref) gives a compact numerical illustration. With known AR(1) residual correlation $\rho=0.5$, the OLS and GLS criteria agree for isolated and smooth cumulative targets, but they differ for the delayed-window contrast: the OLS benchmark selects $r=0.38$, while the feasible-GLS criterion selects $r=0.55$. Choosing the OLS persistence for the GLS estimator raises the GLS target risk by about nine percent. The stylized example makes the estimator-specific point concrete: persistence is not a property of the assignment class alone; it is a property of the reporting object and the planned estimator.
Operationally, this means the design stage should not hide estimator choice. A protocol that plans to report GLS, VAR-based IRFs, smooth LPs, or Bayesian LPs should compute $\operatorname{tr}\{W\widehat V_{\mathcal E}(d)\}$ for that estimator $\mathcal E$, not reuse the homoskedastic OLS closed form. The OLS formulas remain useful as a benchmark because they isolate the assignment-spectrum channel and provide a diagnostic for when persistence itself is the dominant design margin.
The same estimator-specific interpretation covers GMM. A stacked path-adjusted LP can be written as a moment problem with moment vector $m_t(\theta)=(Z_t u_{t,h}(\theta))_{h=0}^H$, where $Z_t$ contains the residualized current and lagged assignment terms and any controls. Ordinary path-adjusted least squares corresponds to a particular identity or equation-by-equation weighting of these moments. Optimally weighted GMM, continuously updated GMM, or empirical likelihood would instead use the long-run covariance of the stacked moments, and the resulting design risk is the GMM sandwich risk rather than the OLS HW-LP risk Hansen1982,HansenHeatonYaron1996,Owen1988. These estimators can dominate identity-weighted LP when the moment covariance is well specified, and their overidentifying-restriction diagnostics have design-dependent power. In such an analysis, the switchback persistence should be chosen by minimizing the planned GMM risk; the closed-form OLS persistence rule is the benchmark for identity-weighted path-adjusted LP, not an efficiency bound for all moment estimators.
A Bayesian decision-theoretic design is another valid specialization of the same principle. If the analyst has credible prior information about smoothness, sign restrictions, monotone decay, or cross-platform shrinkage, the design objective should be the prior or posterior-predictive expected loss, for example \[ d_\pi^\star\in\arg\min_{d\in\mathcal D} E_\pi\{L(\widehat{\theta}_d,\theta)\}, \] where $L$ may be squared error for a reported contrast, a threshold loss for a regulatory decision, or a utility loss for a platform decision. Bayesian $D$- or $A$-optimal design criteria and empirical-Bayes switchback rules are natural counterparts ChalonerVerdinelli1995,AtkinsonDonevTobias2007,XiongChinTaylor2024. Such priors can change the optimal persistence because they shrink or constrain the response curve across horizons. The main criterion remains frequentist and reporting-object based, but the implementation rule is modular: replace $\widehat V(d)$ or the loss function by the posterior risk implied by the planned Bayesian estimator.
Robust LP estimators fit into the same modular template. Half-hourly load, revenue, and clickstream outcomes can have heavy tails, outages, holidays, or equipment failures, so the $4+\delta$ moment condition in Assumption (ref) is an assumption to be checked rather than a free empirical fact. If the pre-analysis plan specifies median, quantile, Huber, or trimmed LP estimation, the design criterion should use the corresponding influence-function or sandwich covariance instead of the OLS covariance Huber1964,KoenkerBassett1978,HampelEtAl1986. Likewise, pilot preprocessing rules---outage flags, holiday flags, winsorization thresholds, or trimming fractions---should be fixed before the confirmatory experiment and repeated in sensitivity checks. The closed-form OLS rule then becomes a benchmark for clean finite-variance designs, while robust or distributionally conservative designs replace $\widehat V(d)$ by the risk of the robust estimator or by a worst-case covariance over a declared ambiguity set.
We study average response targets. If the analyst instead reports a quantile local projection at a fixed quantile $\tau$ and the standard sparsity condition holds, the leading sandwich variance has the form $\tau(1-\tau)f_\varepsilon\{F_\varepsilon^{-1}(\tau)\}^{-2} Q_H(r)^{-1}$ under the same lagged-assignment design. Thus, for a fixed horizon contrast $c$, this homoskedastic quantile-LP benchmark selects the same persistence as the OLS HW-LP rule; the scalar density factor changes scale, not the optimal $r$. This equivalence should not be overread. Peak-load, tail-risk, distributional-welfare, or conditional quantile treatment-effect reporting objects generally change the target vector, aggregation level, subgroup weights, or covariance estimator, and should be designed by replacing $V(d)$ with the covariance or loss for that distributional estimand rather than by reusing the mean-LP rule mechanically.
Applied researchers often report the full response curve rather than a single scalar contrast. This is a different design objective from the scalar cumulative target. If equal marginal precision across horizons is desired, the benchmark weight is $W=I$ and the criterion is $\operatorname{tr}\{V(r)\}$. Under the homoskedastic Markov benchmark this criterion favors iid assignment, not cumulative-target persistence. The numerical full-curve comparison uses the same candidate design menu, while Section (ref) gives the inference rule and Appendix (ref) gives the simultaneous-band construction.
We evaluate the design rules in a high-frequency residential electricity setting calibrated to the Low Carbon London smart-meter data StrbacEtAl2024. The exercise is semi-synthetic: it is not an estimate of the causal effect of the observed dynamic time-of-use tariff. Instead, the observed data provide realistic half-hourly residential-load dynamics, and known dynamic response paths are injected into those dynamics so that different assignment designs can be compared against known targets.
The empirical analysis distinguishes three design questions. A balanced-persistence comparison isolates the persistence margin in the same finite design class as the analytical benchmark. Sparse active-share designs show how target-specific persistence changes under an event budget. Historical tariff schedules are evaluated as fixed design paths using the same information criterion. The distinction separates the persistence choice for a given target, its modification under a sparse budget, and the informativeness of a previously implemented schedule for that target. Only the first two blocks are directly comparable in semi-synthetic target-MSE units; the historical-schedule block is a fixed-schedule information diagnostic rather than a replay-based loss comparison.
The calibrated evaluation uses the full 2013 trial year and the normal-tariff households as the baseline source. The outcome is the group-average half-hourly electricity consumption series in kWh per half-hour. The aggregate panel contains 17,520 half-hour observations. The normal-tariff group has roughly 4,000--4,400 available household readings for most half-hours, and the dynamic-tariff group has roughly 1,000--1,100 available readings for most half-hours. Aggregation is by timestamp and tariff group; available readings are averaged within each group and timestamp. The calibrated normal-tariff baseline series has mean 0.214 kWh per half-hour and standard deviation 0.076.
The pre-specified Monte Carlo design uses 10,000 replications for each scenario-design pair and common random numbers across designs within each scenario. The horizon is $H=8$, corresponding to four hours at half-hour frequency. The four response scenarios are immediate reduction, persistent reduction, delayed reduction, and rebound. Scenario names describe the injected response path used for evaluation; design selection uses only the reporting weights $c$, the feasible persistence set, and the assignment covariance, not the injected magnitudes of $g$.
The evaluation fixes 10,000 replications per scenario-design pair, common random numbers across designs within each scenario, the horizon $H=8$, the balanced candidate set $r\in\{-0.7,0,0.3,0.7\}$ plus target-specific HW-LP candidates, sparse active shares $p_0\in\{0.05,0.10,0.14\}$, and the observed tariff indicators used only for fixed-schedule diagnostics. The generic-MSE selector summarized in Table (ref) uses the same retained finite design menu as the target-specific selector and changes only the objective from $c'\widehat V_d c$ to $E\|\hat g-g\|^2$. The scenario menu was fixed before the full run to span four qualitatively distinct response shapes--front-loaded, smooth cumulative, delayed, and sign-changing rebound--rather than being selected from observed treatment effects. Common random numbers are implemented by using the same innovation stream within a scenario and changing only the design transformation; this is why paired comparisons have much smaller Monte Carlo noise than unpaired comparisons. The replication count was fixed ex ante.
The primary comparison keeps the balanced Markov benchmark to isolate the persistence margin. We compare iid assignment, negatively persistent assignment with $r=-0.7$, a moderate Markov switchback with $r=0.3$, a persistent Markov switchback with $r=0.7$, and the application-feasible closed-form HW-LP rule. The main metric is target MSE normalized so that the best design in each scenario has value one. Because common random numbers are used, the table also reports paired Monte Carlo standard errors for the HW-LP design relative to the best design and a 95 percent tie flag.
Table (ref) evaluates each pre-specified design under known injected response paths; selection results are reported separately below. The delayed-reduction scenario is reported as a stress test for the closed-form rule. The homoskedastic Markov benchmark selects $r=0.382$, while the retained calibrated loss is minimized by the persistent design $r=0.7$. This is a substantive negative finding for the closed-form benchmark as a field rule: the analytical formula is useful for understanding the persistence mechanism, but it should not be used mechanically when residualized finite-sample covariance, serial dependence, or calendar dynamics differ from the stylized Toeplitz benchmark. The calibrated implementation in Proposition (ref) is designed for exactly this case: the selected persistence should be based on $c'\widehat V_d c$ over the retained feasible designs.
The remaining scenarios are closer to the analytical benchmark. For the immediate-reduction target, the HW-LP rule selects $r=0$, which coincides with iid assignment and delivers the lowest target MSE. For the persistent-reduction target, the application-feasible HW-LP rule selects the upper application endpoint, $r=0.7$, and this is also the best design among the retained comparisons. For the rebound scenario, the closed-form HW-LP rule selects moderate persistence and is statistically tied with the fixed moderate design under the paired Monte Carlo comparison.
The balanced comparison is the primary evidence for scalar targets because it most closely matches the closed-form theory. The sparse-budget and historical-schedule analyses extend the same design logic to additional constraints and fixed schedules rather than replacing the balanced benchmark.
The full-covariance evaluation retains the full response vector $\hat g$ in each Monte Carlo draw and computes the empirical covariance of the path-adjusted LP estimator for each retained design. This makes it possible to evaluate the finite-candidate calibrated selector \[ \widehat d_c\in\arg\min_{d\in\mathcal D} c'\widehat V_d c, \] where $\mathcal D$ is the finite set of retained designs. The selector uses the covariance of the estimator, not the realized squared error or the injected magnitudes of the response path. Table (ref) reports the closed-form benchmark, the calibrated selector, and the best target-MSE design side by side. The best target-MSE design is observable only in this semi-synthetic exercise; it is included to show whether the feasible covariance selector points in the same direction as the known target loss.
The calibrated selector agrees with the closed-form benchmark for the immediate, persistent, and rebound targets, but it replaces the delayed-window benchmark with the persistent design. The generic full-curve-MSE selector is efficient for persistent and delayed targets, yet it over-persistently designs the immediate and rebound scalar targets. This divergence reflects the target-specific design criterion: persistence should be chosen for the reporting object, not for a generic curve loss.
The same target-MSE comparison yields a field decision. The benchmark and calibrated candidates coincide for the immediate and persistent targets. For the delayed target, the retained benchmark proxy has relative target MSE 2.224 and should be replaced by the calibrated persistent design. For the rebound target, the benchmark proxy has relative target MSE 1.025 and remains within the pre-specified 5 percent tolerance. The known semi-synthetic target loss is used only to evaluate the rule; field selection itself uses the covariance criterion.
The paired standard errors in Table (ref) are small in raw target-MSE units because common random numbers remove the shared baseline-load component. They are not evidence that the exercise is deterministic. The relative paired standard-error scale is economically interpretable: it is near zero when HW-LP and the best retained design are the same collapsed design, about 3.4 percent in the delayed stress test, and about 2.0 percent in the rebound comparison.
The balanced comparison isolates the persistence margin, but demand-response programs usually impose an active-share budget. Table (ref) evaluates sparse unbalanced Markov HW-LP designs with active shares $p_0\in\{0.05,0.10,0.14\}$. These designs use the same target-specific persistence logic but operate under smaller information scale $p_0(1-p_0)$. The sparse block is both a limitation check and an implementation exercise. When $p_0=0.05$, the delayed-target relative MSE is 12.09: changing persistence cannot compensate for an event budget that leaves too few active half-hours. In such settings the practical conclusion is not that one should search harder within the first-order Markov class, but that the experiment needs a larger active share, a longer trial, a stronger prior/model-based estimator, or a more modest reporting target. The HW-LP rule allocates the available information across horizons; it does not create information that the active-share budget removes.
Appendix (ref) reports five diagnostics that qualify the main comparison without changing its target-MSE interpretation: calendar-state sensitivity, historical tariff schedules evaluated as fixed assignment paths, finite design-matrix conditioning, injected-path scale, and pilot-window stability. The calendar-state exercise shows that a common persistence recommendation can range from iid assignment to about $r=0.69$ when reporting weights change across weekday and weekend states. The historical schedule diagnostic confirms that the observed tariff signals are operationally realistic but sparse and persistent, which makes some multi-horizon contrasts information-poor. These are information diagnostics rather than causal estimates of the historical tariff.
The pilot-window checks support the calibrated-selector implementation. Three-month or longer aggregate baseline windows produce stable recommendations for all four targets, whereas one-month windows can under-select the cumulative target's boundary persistence. Appendix Tables (ref), (ref), and (ref) report the numerical evidence. The exercise remains a semi-synthetic covariance stress test: it evaluates known injected response paths under realistic half-hourly load dynamics and does not estimate household-level elasticities or forecast contemporary demand-response magnitudes.
The results are benchmark and implementation tools for a pre-specified dynamic reporting object. The scope is narrower than the set of all dynamic-design problems. Three boundaries matter.
The main target is an assigned-path, reduced-form LP response for a discrete-time treatment path. Structural price elasticities, welfare primitives, market-clearing counterfactuals, mechanism decompositions, distributional welfare, household-level heterogeneous effects, individual treatment effects, and continuous-horizon response curves require different estimands or additional assumptions. Similarly, GLS, VAR or transfer-function estimation, GMM, Bayesian or smooth LPs, Huber or quantile LPs, and regularized large-horizon LPs should use their own estimator-specific covariance or posterior risk rather than mechanically reusing the OLS/HAC benchmark. The finite-memory benchmark treats the lag coefficients $g_\ell$ as time-invariant over the experiment. If $g_\ell$ varies smoothly with season, temperature, or adoption trends, the path-adjusted LP targets a time-averaged response under the chosen weights; abrupt regime changes instead require stratified analysis or time-varying-parameter estimators before applying the calibrated selector. The HW-LP logic is modular in this respect: specify the reporting object, estimator, and loss weights first, then evaluate the corresponding design risk.
First-order Markov switchbacks are attractive because they are operationally transparent, easy to implement, and directly control expected run length and active-spell counts. They are not globally optimal over higher-order Markov, multisine, adaptive, finite-sequence, business-loss-aware, or continuous-time designs. Multi-arm or dose-response switchbacks also fall outside the binary closed form because their transition structure is higher-dimensional and the information matrix is generally block-Toeplitz. Appendix Figure (ref) shows the distinction in a three-arm diagnostic: symmetric switching retains a separable scalar structure after contrast coding, while directional cycling produces matrix-valued lag blocks. The calibrated covariance selector remains applicable once the realized or model-implied block information matrix is computed. Richer finite-sequence algorithms can improve risk when the exact calendar grid and balance constraints are fixed, as Table (ref) illustrates. If the assignment follows a continuous-time switching process and is sampled at interval $\Delta$, the discrete persistence should be interpreted through $r\simeq \exp(-\lambda\Delta)$; changing the analysis frequency changes the persistence scale and requires a new design calculation. Irregularly sampled outcomes--event-driven trades, jittered impressions, or sparse session-level logs--require preprocessing to a regular analysis grid or a continuous-time HW-LP formulation; the discrete Toeplitz benchmark assumes fixed-interval observations. The reporting horizon $H$ is a pre-specified reporting resolution, not an estimate of the true biological or economic memory. If a deployment requires 12--48 hour rebound targets or smooth response curves, the analysis should jointly pre-specify the target horizon, regularization, and inference adjustment rather than treating the $H=8$ benchmark as universal. Cross-sectional interference, two-stage clustered switchbacks, state-dependent persistence, online learning, group-sequential stopping, and platform bandits are compatible with the broad idea of target-specific design, but they replace the stationary Toeplitz matrix with a realized, time-inhomogeneous, or policy-dependent information matrix.
Field use also imposes constraints that are not statistical optimality conditions: participant burden, business loss, settlement windows, capacity obligations, maximum bill-risk protections, consent and withdrawal rules, privacy budgets, secure aggregation, audit logs, and regulator-facing disclosure. These constraints should enter the feasible design set or the loss function before the experiment begins. The Low Carbon London analysis is a semi-synthetic aggregate covariance stress test based on one public demand-response setting, not a transportability claim for other populations, climates, technologies, or household subgroups. For new deployments, the recommended procedure is to choose the reporting object and weight matrix, select a pilot-free or pilot-calibrated assignment rule, use randomization-first inference in high-persistence regimes, and document the realized assignment, exclusions, numerical diagnostics, and finite-sample balance reports.
Dynamic experiments should be designed for the dynamic objects researchers plan to report, and the field recommendation should be calibrated to the estimator and covariance environment actually used. We develop a horizon-weighted local projection criterion that maps a target response curve or contrast into an assignment design problem. In balanced Markov switchback designs, the criterion is analytically tractable: the information matrix is AR(1)-Toeplitz, its inverse is tridiagonal, and the optimal assignment persistence has a closed form. That formula is the benchmark, not the universal prescription.
The benchmark results show that there is no universally optimal assignment persistence. Within the benchmark Markov class, iid assignment is appropriate for isolated horizon effects. Persistent assignment can be efficient for smooth cumulative effects. Moderate or alternating assignment can be useful for rebound and other sign-changing contrasts. The omitted-carryover analysis further cautions against choosing extreme persistence solely because it lowers variance for a smooth target: persistence is the variance-allocation margin, while buffer horizons are the main analysis-design tool for protecting reported contrasts from omitted tails.
The results should be interpreted as design guidance for experiments that can choose the temporal pattern of assignment after pre-specifying a dynamic reporting object. The empirical results separate three design tasks: use the analytical benchmark when the main issue is persistence itself, use the sparse-budget comparison when the design is constrained by event frequency or fatigue, and use the historical-schedule diagnostic when the relevant question is whether an already deployed tariff path was informative for the target of interest. These three evidence blocks answer related but distinct design questions and should not be collapsed into a single ranking. The closed-form rules are benchmark formulas; in field settings with heterogeneous residual variance, calendar controls, or serial dependence, the same objective is implemented using pilot, HAC, residual-bootstrap, or calibrated-baseline covariance estimates. The semi-synthetic Low Carbon London exercise therefore evaluates designs under realistic load dynamics with known injected response paths, rather than estimating the causal effect of the observed tariff schedule. The LCL results calibrate design risk under realistic half-hourly covariance; they are not a forecast of contemporary demand-response magnitudes.
The benchmark comparisons and Low Carbon London calibrated evaluation illustrate that these differences are empirically meaningful. Mismatching the assignment design to the reported dynamic object can substantially inflate target MSE. In sparse demand-response settings, the active-share budget should be separated from the persistence choice: the budget determines information scale, while persistence determines how information is allocated across horizons. Extending the closed-form rule to multi-arm or dose-response switchbacks is a natural next step. The three-arm diagnostic in Appendix Figure (ref) shows why the extension is nontrivial: the information matrix is block-Toeplitz and directional switching can invalidate a scalar persistence summary, although the calibrated selector still applies once that block matrix is formed. The implication is that the experimental design should be chosen after specifying the local-projection target, not before.