Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
79,197 characters · 27 sections · 71 citation commands
Bounds for Standard Errors in Combined Data
Researchers frequently combine empirical moments from multiple, and potentially interdependent, data sources. Examples include integrating survey data with administrative records, combining cross-sectional data with time-series data, and using time-series observations recorded at different frequencies. In applications such as meta-analysis and calibration, the focus is often placed more on parameter estimation than on statistical inference. This emphasis may reflect the fact that conventional inference methods are typically infeasible, because the covariances of empirical moments across samples are often unknown, rendering standard errors for the parameter of interest impossible to compute.
When a covariance is not directly observable, it can still be bounded using marginal variances. Recent work by mikkel derives a sharp upper bound on the covariance from marginal variances via the Cauchy--Schwarz inequality, and vohra2025inference extends mikkel along several dimensions. We complement this literature by characterizing the full range of standard errors that are feasible when only marginal variances are known. In addition to the sharp worst-case standard error, we derive the sharp best-case standard error and show that it is governed by a geometric cancellation condition. The resulting interval describes exactly how much uncertainty about the standard error is induced by the absence of covariance information.
The diagonal-only case provides the least-informative benchmark. The upper bound corresponds to maximal alignment of the sampling errors across empirical moments. The lower bound corresponds to maximal cancellation. This lower bound is often zero, but this should not be interpreted as a defect of the bound. Rather, it reveals when the unknown covariance structure can, in principle, offset the marginal sampling uncertainty in the combined estimator. Our characterization makes this possibility precise. With two moments, exact cancellation requires a knife-edge balance between the two components. With three or more moments, the condition has a geometric interpretation: the marginal contributions must be capable of closing into a polygon. Thus, the sharp lower bound describes the covariance configurations under which the combined estimator can have a very small standard error despite nontrivial uncertainty in its individual components.
The diagonal-only bounds are useful not because researchers should necessarily base inference on either extreme, but because they describe the full range of standard errors consistent with the available information. Figure (ref) outlines how applied researchers can use the bounds we derive and the existing knowledge of the data as a diagnostic for adjusting their empirical design. If both the lower and upper bounds are large, the data are uninformative for the parameter of interest regardless of the unknown covariance structure. If the bounds are far apart, the missing covariance information is consequential: different admissible covariance structures lead to substantively different inferential conclusions. In that case, the bounds may help decide whether it is worth collecting, estimating, or justifying additional information about the covariance structure.
In many combined-data applications, researchers may know more than the marginal variances. The design of the data combination may imply that some components are estimated from independent samples, that two samples partially overlap, that dependence is confined within blocks, or that some pairwise correlations have known signs or plausible magnitude restrictions. We show how such information sharpens the standard error bounds. Information that rules out strong negative effective dependence raises the lower bound by limiting cancellation. Information that rules out strong positive effective dependence lowers the upper bound by limiting alignment. Exact independence, block independence, bounded sample overlap, and sign restrictions therefore have transparent effects on the range of feasible standard errors.
We develop two complementary approaches. First, we provide analytic characterizations for several empirically relevant forms of covariance information. For example, when moments can be partitioned into independent blocks, the bounds decompose block by block. When two components are estimated from partially overlapping samples, the overlap structure can place direct limits on the magnitude of their correlation. These cases show how features of the data-combination design translate into sharper inference. Second, for more general forms of partial covariance information, we formulate the problem as a semidefinite program (SDP). This computational approach allows researchers to incorporate arbitrary linear or convex restrictions on the unknown covariance matrix and obtain the corresponding standard-error bounds using standard numerical solvers.
Our framework applies to a wide range of empirical settings. In two-sample instrumental variables, conventional inference often relies on an independence assumption between the sample used to estimate the first stage and the sample used to estimate the reduced form. Our approach nests that assumption but also allows researchers to examine weaker alternatives, such as bounded cross-sample dependence. Similar issues arise in empirical macroeconomics, where constructed or generated shock series are often estimated in one sample and then used in downstream analyses. Classical generated-regressor arguments, such as those associated with pagan1984econometric, justify standard inference when the relevant covariance structure is negligible or consistently estimable. But when a constructed shock is made externally and then used in a different estimating sample, the relevant joint covariance may not be directly recoverable. Moreover, calibration, the most common setting in macroeconomics applicable to our framework kaplan&violante:2014, mckayetal:2016, nakamura&steinsson:2018, kydland&prescott:1982, gourinchas&parker:2003, arellano:2008, eatonetal:2011, faces an analogous problem under an unknown covariance structure among moments used for matching. Our bounds then quantify the range of standard errors consistent with the available marginal information.
The paper contributes to the broader literature on data combination, where much of the existing work focuses on identification and estimation rather than inference; see handbook. It also contributes to recent work on inference when covariance information is incomplete, especially mikkel and vohra2025inference. Relative to this literature, our contribution is threefold. We derive sharp lower as well as upper bounds on standard errors under diagonal-only information. We characterize the geometry of the lower bound and the conditions under which exact cancellation is possible. Finally, we show how partial information about the covariance structure moves the identified interval from the diagonal-only benchmark toward the usual full-information standard error.
The remainder of the paper is organized as follows. Section (ref) introduces the framework. Section (ref) derives the explicit sharp bounds. Section (ref) characterizes the set of achievable standard errors and how additional information on the covariances of the moments can tighten the bounds on this set. Section (ref) presents the SDP approach. Section (ref) illustrates empirical applications. Section (ref) concludes. All mathematical proofs are collected in the appendix.
Suppose that we are interested in a scalar parameter $\varphi(\theta)$, where $\varphi : \mathbb{R}^k \to \mathbb{R}$ is a known function of a $k$-dimensional vector of structural parameters $\theta$. The structural parameters, in turn, determine a $p$-dimensional vector of reduced-form parameters $\mu = h(\theta)$ through a known injection $h : \mathbb{R}^k \to \mathbb{R}^p$.
If estimates $\hat{\mu} = (\hat{\mu}_1, \cdots, \hat{\mu}_p)'$ of the reduced-form parameters $\mu$ are available, then standard inference about $\varphi(\theta)$ is based on the linear approximation
The results below rely only on the linear representation in (ref). The structural mapping $h$ provides one leading way to obtain the loading vector $\ell$, but the same framework applies to any estimator whose scalar target admits such a first-order expansion, including moment-matching and two-step estimators discussed below.
In the exactly identified case in which $\hat\theta=h^{-1}(\hat\mu)$, the loading vector is \[ \ell = \frac{\partial \varphi}{\partial \theta'}(\theta) \cdot \frac{\partial h^{-1}}{\partial \mu}(\mu). \]
In particular, if we have the standard convergence in distribution
then the delta method yields
This asymptotic distribution facilitates inference about $\varphi(\theta)$, provided that the asymptotic variance $\sigma^2$ is estimable.
However, if the reduced-form estimates $\hat{\mu}_1, \cdots, \hat{\mu}_p$ are obtained from different samples, then not all elements of the asymptotic covariance matrix $\Sigma$ are estimable. To fix ideas, suppose that $\hat{\mu}_j$ is computed from a sample $X_j = (X_{j,1}, \cdots, X_{j,n_j})$, where each $X_j$ is observed separately for $j = 1, \cdots, p$. In this setting, only the diagonal elements of $\Sigma$ are estimable: \[ s_j^2 := \Sigma_{jj}, \quad \text{for } j = 1, \cdots, p. \] With the off-diagonal elements of $\Sigma$ unknown, however, it is not possible to estimate the asymptotic variance $\sigma^2 = \ell' \Sigma \ell$ of the parameter $\varphi(\theta)$.
In the remainder of the paper, we propose methods to obtain sharp lower bounds for the asymptotic variance $\sigma^{2}$ of interest, using the marginal variances $s_{j}^{2}$ for $j = 1, \dots, p$, together with no or only partial information about the covariances.
If the researcher has information only about the diagonal elements of $\Sigma$, then explicit sharp bounds for $\sigma$ can be derived. The following theorem formally states this result.
Proof of this statement is found in Appendix (ref).
The sharp upper bound coincides with the result of mikkel; this component of Theorem (ref) is therefore not novel, and we include it solely for completeness and the convenience of the reader. In contrast, the sharp lower bound established in Theorem (ref) is new to the literature.
By the continuous mapping theorem, these sharp bounds can be consistently estimated under the following assumption.
The analytical bounds can be interpreted as the following. The upper bound is attained when the estimation errors associated with all moments are perfectly aligned in the direction that increases the variance of the scalar target. The lower bound is attained when these components cancel as much as possible. With two moments, exact cancellation is possible if and only if $|\ell_{1}|s_{1} = |\ell_{2}|s_{2}$. More generally, the lower bound is zero if and only if the largest $|\ell_j| s_j$ is no larger than the sum of the remaining $|\ell_k|s_k$'s. This is the polygon condition: the lengths $|\ell_1|s_1,\dots,|\ell_p|s_p$ can be arranged as vectors whose sum is zero. Thus, a zero lower bound reflects feasible covariance cancellation under diagonal-only information, not a failure of the bound.
The analytical lower bound provided by Theorem (ref) and Proposition (ref) allows all correlation structures consistent with positive semidefiniteness. In many applications, however, the researcher may have some information about the signs, magnitudes, or block structure of the unknown correlations. Such information restricts the covariance structures that can generate maximal cancellation or maximal alignment. The subsequent sections formalize this idea.
This section introduces empirically plausible covariance information and how it may tighten the bounds in Section (ref). We show that the lower bound is driven by the extent to which the unknown covariance structure can generate cancellation across empirical moments, whereas the upper bound is driven by the extent to which it can generate alignment. This representation also clarifies how additional information about the covariance matrix tightens the bounds.
Before investigating the role of covariance information, we first characterize the set of achievable standard errors without information restriction. Furthermore, we study where the best-case and worst-case correlation structures are located in the set of feasible correlation matrices.
Proposition (ref) establishes that the set of achievable standard errors takes the form of a simple interval, with no gaps between its endpoints. Consequently, the full-information standard errors, those computed with knowledge of the off-diagonal elements of $\hat{\Sigma}$, can be expressed as a weighted average of the best-case and worst-case standard errors.
Proof of this statement is found in Appendix (ref).
We can also generally characterize where the minimal and maximal correlation structures lie within the set of feasible correlation matrices. In the special case of two moments, the best-case correlation occurs when the two moments are either perfectly positively or perfectly negatively correlated.
Proof of this statement is found in Appendix (ref).
In the $p = 2$ case of Proposition (ref), the dependence on the sign of the product of the entries of $\ell$ highlights the mechanism that drives the best-case correlation away from the worst-case correlation.
Let $D = \operatorname{diag}(s_1,\ldots,s_p)$ where $s_j^2=\Sigma_{jj}$ for $j=1,\cdots,p$. Write $\Sigma = D R D$, where $R$ is a correlation matrix. Define $z_j = |\ell_j|s_j$ for each $j = 1,\cdots,p$, and collect them as the vector $z=(z_1,\ldots,z_p)'$. Let $S$ be the diagonal matrix with $j$-th diagonal entry equal to $\operatorname{sign}(\ell_j)$, with an arbitrary sign convention when $\ell_j=0$, and define the sign-adjusted correlation matrix $T = SRS$. Then, $T$ is also a correlation matrix, and the asymptotic variance can be written as
This representation is useful because the sign of $T_{ij}$ has a direct interpretation. When $T_{ij}>0$, the estimation errors associated with moments $i$ and $j$ are effectively aligned in the direction relevant for $\varphi(\theta)$, increasing the variance. When $T_{ij}<0$, they are effectively offsetting each other, decreasing the variance. Thus, restrictions that prevent $T_{ij}$ from being very negative limit cancellation and tend to raise the lower bound, while restrictions that prevent $T_{ij}$ from being very positive limit alignment and tend to lower the upper bound.
This distinction is important because the raw correlation $R_{ij}$ need not have the same interpretation as the effective correlation $T_{ij}$. If $\ell_i$ and $\ell_j$ have the same sign, then $T_{ij}=R_{ij}$. If they have opposite signs, then $T_{ij}=-R_{ij}$. Thus, a positive raw correlation may induce cancellation when the two moments enter the scalar target with opposite signs.
Let $\mathcal T$ denote a nonempty set of admissible sign-adjusted correlation matrices. For example, $\mathcal T$ may encode diagonal-only information, known zero correlations, sign restrictions, magnitude restrictions, block restrictions, or any combination of these. Define the lower and upper feasible standard errors associated with $\mathcal T$ by
respectively. When $\mathcal T$ is the full set of correlation matrices, these endpoints coincide with the diagonal-only bounds in Theorem (ref). This notation makes the role of covariance information explicit. If $\mathcal T_2\subseteq\mathcal T_1$, then
That is, additional covariance information weakly contracts the feasible range of standard errors. The rest of this section studies which empirical forms of covariance information shrink $\mathcal T$ in ways that are informative for the lower bound, the upper bound, or both.
The representation in (ref) makes clear which types of covariance information affect which endpoint of the range. Negative values of $T_{ij}$ generate cancellation, while positive values generate alignment. Consequently, restrictions that bound $T_{ij}$ away from $-1$ limit cancellation and can raise the lower bound. Restrictions that bound $T_{ij}$ away from $1$ limit alignment and can lower the upper bound. A useful general form of pairwise covariance information is
where the lower endpoint $\underline{\rho}_{ij}$ restricts the amount of cancellation available between moments $i$ and $j$, while the upper endpoint $\overline{\rho}_{ij}$ restricts the amount of alignment available between them. Thus, a one-sided lower restriction is primarily informative for the best-case standard error, a one-sided upper restriction is primarily informative for the worst-case standard error, and a two-sided restriction may tighten both.
We discuss several empirically motivated restrictions.
Sign restrictions are an important special case. If $T_{ij}\ge0$, then the pair $(i,j)$ cannot contribute to cancellation, although it may still contribute to alignment. Such information is therefore informative for the lower bound. If $T_{ij}\le0$, then the pair cannot contribute to alignment, although it may still contribute to cancellation; this is informative for the upper bound. If $T_{ij}=0$, then the pair contributes neither to cancellation nor to alignment, and both endpoints may tighten. These restrictions are stated in terms of the sign-adjusted correlation $T$, not the raw correlation $R$; the two coincide only when $\ell_i$ and $\ell_j$ have the same sign.
These restrictions often arise from the design of a combined-data problem. If two moments are estimated from independent samples, the corresponding covariance is zero and hence $T_{ij}=0$. If moments can be partitioned into independent blocks, then all cross-block entries of $T$ are zero, while dependence may remain unrestricted within each block. If moments are estimated from partially overlapping samples, then the overlap structure may imply a magnitude bound of the form
where $q_{ij}<1$ summarizes the maximal cross-source dependence allowed by the sampling design. More generally, weak dependence across markets, cohorts, time periods, or data sources can be represented by similar bounds on cross-block entries of $T$.
In a few special cases, these restrictions yield closed-form bounds. For example, with two moments and $T_{12}\in[\underline{\rho},\overline{\rho}]$, we have
This formula nests independence, sign restrictions, and bounded-dependence restrictions. For instance, independence corresponds to $\underline{\rho}=\overline{\rho}=0$, while $|T_{12}|\le q$ corresponds to $[\underline{\rho},\overline{\rho}]=[-q,q]$.
Another tractable case is block independence. Suppose that the moments are partitioned into blocks, cross-block correlations are zero, and within each block no off-diagonal information is available. Then, the diagonal-only formulas of Theorem (ref) apply within each block, and the resulting block-level variances add across blocks. Thus, block information rules out cross-block cancellation and alignment while preserving the diagonal-only benchmark within blocks.
The preceding discussion concerns the full range of feasible standard errors. It is also useful to visualize when additional covariance information rules out exact cancellation, that is, when it rules out a zero lower bound. Consider the case of three moments and normalize the marginal contributions by $q_j=z_j/(z_1+z_2+z_3)$. Each point in the simplex represents the relative importance of the three marginal contributions. Under diagonal-only information, exact cancellation is feasible if and only if $\max_j q_j\le1/2$. If exact cancellation occurs, then the sign-adjusted correlations must satisfy $Tq=0$, where $$T = \left(
\right) \qquadand\qquad q=\left(
\right).$$ Solving this linear system gives the following pairwise correlations:
If the researcher imposes $T_{ij}\ge-\gamma$ for all $i,j$, then exact cancellation requires
As $\gamma$ decreases, this zero-cancellation region shrinks. Figure (ref) plots this region for different values of $\gamma$. The simplex describes the relative sizes of the three marginal contributions.
Shaded points are those for which the available covariance information still permits exact cancellation. Unshaded points are those for which the zero lower bound is no longer feasible.
These examples illustrate the general mechanism. Lower-bound informativeness comes from ruling out the negative effective correlations required for cancellation. Upper-bound informativeness comes from ruling out the positive effective correlations required for alignment. The next section describes how these restrictions can be imposed computationally when closed-form expressions are not available.
The previous section shows that a researcher can obtain explicit sharp bounds when the values of only the diagonal elements of $\Sigma$ are known. In practice, however, there are situations where the researcher also possesses partial information about the off-diagonal elements of $\Sigma$. For example, block-diagonal elements of $\Sigma$ are known when multiple moment conditions are constructed from each data set.
We show that, in such general settings, the problem of obtaining bounds for the standard error can be reformulated as an equivalent semidefinite program (SDP), which can be readily implemented using existing numerical solvers. Since mikkel analyzes worst-case standard errors, we focus hereafter on the best-case standard errors.
Let $\mathcal{I} \subset \mathbb{N}^2$ denote the set of row--column index pairs of the variance estimate $\hat{\Sigma}$ for which the researcher possesses information. For example, if the researcher knows only the diagonal elements of $\hat{\Sigma}$, then $\mathcal{I} = \{(1,1), \cdots, (p,p)\}$. If, in addition, the researcher knows certain off-diagonal elements $\hat{\Sigma}_{jk}$ and $\hat{\Sigma}_{kj}$, then the corresponding index pairs $(j,k)$ and $(k,j)$ are also included in $\mathcal{I}$. Define $S(\hat{\Sigma}, \mathcal{I})$ as the set of all symmetric, positive semidefinite matrices that share with $\hat{\Sigma}$ the entries corresponding to the index pairs in $\mathcal{I}$, that is, \[ S(\hat{\Sigma}, \mathcal{I}) \equiv \left\{ \Tilde{\Sigma} \succeq 0 \ :\ \Tilde{\Sigma} = \Tilde{\Sigma}', \ \Tilde{\Sigma}_{ij} = \hat{\Sigma}_{ij} \text{ for all } (i,j) \in \mathcal{I} \right\}, \] where $A \succeq 0$ indicates that the matrix $A$ is positive semidefinite.
With the above notation, our problem now concerns the computation of $n^{-1/2}$ times
To formally motivate this minimization problem, we state the following assumption.
Part (i) requires that all diagonal elements of $\hat{\Sigma}$ are known, consistent with the setting in the previous section. Part (ii) parallels Assumption (ref) and is extended to the current framework.
The following proposition formally justifies (ref) as the main object of interest.
Proof of this statement is found in Appendix (ref).
Furthermore, the following proposition establishes that the constrained optimization problem (ref) can be reformulated as a numerically feasible SDP.
Proof of this statement is found in Appendix (ref).
With Proposition (ref), we can obtain the best-case standard error (ref) by solving the SDP (ref)--(ref) after computing $\hat\ell$, $\hat\Sigma$, and $\hat{\rho}_{ij}$ for every pair of moments $i$ and $j$ whose correlation we can estimate.
Although reliable SDP solvers are available in virtually every computing environment,\footnote{For this paper, the preferred language was Python, and the SDP solver employed was the function provided in the package CVXOPT.} numerical optimization may occasionally yield undesired results. In particular, depending on the numerical properties of the problem or the available computational resources, applying an SDP solver to (ref)--(ref) may produce an $\Omega$ that is not exactly a correlation matrix. For instance, the diagonal elements of $\Omega$ may deviate from $1$, or the off-diagonal elements $\Omega_{j,k}$ may fall outside the interval $[-1,1]$.
Nonetheless, such violations occur only at a negligible scale,\footnote{For example, one of the diagonal entries may take the value $\Omega_{j,j} = 0.9999999$, as observed in the initial computations of the best-case standard errors for the parameter $n$ of the alvarez&lippi:2014 menu cost model.} and are more likely attributable to limitations of the software than to any substantive theoretical flaw. To address this issue, a researcher may round the entries of the resulting correlation matrix, which yields virtually identical standard errors for the estimated parameters. Rather than directly rounding the entries of the $\Omega$ produced by the SDP solver, the researcher can proceed as follows: first, run the SDP; then, if the solution does not correspond to a valid correlation matrix, apply the following alternative method to construct a correlation matrix that satisfies (ref)--(ref):
The above steps address the core numerical optimization issue: the initial correlation matrix output is not exactly positive semidefinite because it contains negative eigenvalues. Since positive semidefiniteness requires all eigenvalues to be non-negative, these negative eigenvalues can be rounded to zero, as they are effectively negligible. Such rounding does not affect the standard errors implied by the correlation matrix. Steps 3, 4, and 5 then provide an additional robustness check to ensure that, after decomposition and rounding, the resulting correlation matrix remains valid and optimal.
As the initial outputs of SDP solvers may require adjustment to yield valid correlation matrices, one might worry whether the adjusted correlation matrix obtained via the implementation guidelines in the previous subsection indeed corresponds to the true lower bound of the standard errors.
A natural way to verify the robustness of this approach is to check whether the Karush--Kuhn--Tucker (KKT) conditions are satisfied by the adjusted correlation matrix. The KKT conditions generalize the method of Lagrange multipliers by providing optimality criteria for problems with both equality and inequality constraints. Commonly employed in nonlinear programming, the KKT conditions involve solutions to both the primal problem (the original formulation of the optimization problem) and the dual problem (an alternative but equivalent formulation) to determine whether a global optimum has been attained. When the primal problem is convex, the KKT conditions are both necessary and sufficient for primal and dual optimality, thereby guaranteeing that the solution is globally optimal boyd.
By checking the KKT conditions, researchers can confirm that the implied standard errors of the adjusted best-case correlation matrices for particular model parameters are indeed the smallest feasible standard errors. Since most SDP solvers simultaneously compute both primal and dual solutions, verifying the KKT conditions requires only straightforward computations involving these solutions, the known marginal variances of the empirical moments, and the Jacobian of the structural model evaluated at the point estimates of the chosen moments. Proposition (ref) below specifies the conditions that can be used to verify the optimality of adjusted correlation matrices.
Proof of this statement is found in Appendix (ref).
The Primal Feasibility condition requires that the adjusted correlation matrix $\Omega$ satisfies the generalized inequality and equality constraints of the original SDP, as specified in (ref) and (ref). The Dual Feasibility condition imposes that the dual variable associated with the positive semidefiniteness constraint is itself positive semidefinite, ensuring that the resulting matrix $\Omega$ is indeed positive semidefinite. The Complementary Slackness condition requires that, for each constraint, exactly one of the following holds: (i) the constraint binds with equality, or (ii) the corresponding Lagrange multiplier equals zero. Finally, the Stationarity condition, which is analogous to a first-order condition in unconstrained optimization, requires that the gradient of the Lagrangian with respect to $\Omega$ vanishes at the optimal $\Omega$.
In this section, we apply the SDP method described in Section (ref) to compute the best-case standard errors of three empirical applications: two calibrated macro models and an IV study of public-housing effects. We verify these estimates as optimal using the KKT conditions in Proposition (ref), and we compare the best-case standard errors to the worst-case standard errors.
The first application is based on alvarez&lippi:2014 and concerns the model of a multiproduct firm facing a price-setting decision in the context of sticky prices. Price rigidity stems from the need to pay a fixed “menu” cost prior to price adjustment. While the firm's desired prices change throughout time, it must weigh the benefits of changing all the prices of its products simultaneously against the loss it will incur in paying the menu cost. The $k = 3$ structural parameters of the model are the number of products of the firm ($m$), the volatility of the firm's desired prices ($\nu$), and the scaled menu cost relative to the curvature of the firm's profit function ($\sqrt{\psi/B}$).\footnote{$\psi$ denotes the menu cost while $B$ represents the curvature of the profit function.} These parameters are calibrated using $p = 4$ empirical moments consisting of the frequency of price changes, and the first, second, and fourth moments of price changes.
A firm sells $m$ products and sets the prices of each product simultaneously. The set prices remain fixed until the firm pays the menu cost $\psi$, and the firm's desired log prices change in continuous time following the process of $m$ independent random walks without drift and with volatility $\nu$. The price gap is defined as the difference between the firm's desired prices for their $m$ products and the actual prices of the products. Each period, the firm's profits are proportional to the sum of the squared price gaps. $B$ is a proportionality constant representing the loss of charging a price different from the desired level. Thus, $B$ measures the curvature of the profit function. The firm must then minimize the expected discounted cost of selling at prices that deviate from the desired levels and the loss from paying the fixed menu cost. A nonlinear transformation of the structural model parameters imply the observed empirical moments of the number of price adjustments per unit of time ($N_{a}$), the first moment of the absolute log price changes ($\mathbb{E}[|\Delta p_{i}|]$), the second moment ($\mathbb{E}[(\Delta p_{i})^{2}]$), and the fourth moment ($\mathbb{E}[(\Delta p_{i})^{4}]$). Below, the implied moment-matching relationship is rederived using Propositions 4 and 6, and Equation 10 of alvarez&lippi:2014:
Following the data cleaning procedure of alvarezetal:2016, we fit the calibrated model in Equation \hyperref[eq:13]{(ref)} using scanner data from a single store of the supermarket chain Dominick's for the years 1989 to 1994. For this application, the four empirical moments are computed with the movement data set for beer products.\footnote{This data set, entitled wber.csv, can be found on the Chicago Booth website (\url{https://www.chicagobooth.edu/research/kilts/datasets/dominicks}).} As in alvarezetal:2016, we follow a four-step process to clean the data prior to estimation: (1) We drop the data on all stores except for store 122; (2) Any observations that correspond to prices below 20 cents or above 25 dollars are discarded; (3) Observations with absolute price changes less than one percent are set to zero; (4) The largest 1% of absolute log price changes are also discarded. We use only data on beer products to simplify the inference exercise and make the results of the application easier to interpret.
This process yields a data set of sample size $N = 37,916$ observations on weekly prices of 499 beer products. On average, each beer product has observations of around 76 weeks. Price changes are treated as i.i.d. processes across beer products and time in the calculation of standard errors.
For the standard error computations, we use the four estimated empirical moments shown in Table 3 of mikkel. We first perform full-information inference (i.e. using the readily estimable correlation structure computed by accessing the full data from a single source). Then, we find the worst-case and best-case standard errors simulating the scenario where we only have the diagonal elements or marginal variances of the estimated covariance matrix. Additionally, we calculate estimates under the erroneous yet commonly used in practice assumption that the empirical moments are mutually independent of one another (i.e. the off-diagonal elements of the covariance matrix are set to zero).
We estimate the structural model parameters using all four moments. The full-information estimates were computed using conventional full-information one-step estimation, and the limited-information best-case and worst-case estimates are calculated with the same efficient weighting matrix suggested by the Section 3.2 algorithm of mikkel. The estimates and standard errors implied by a procedure assuming mutually independent moments rely on an identity weighting matrix. The point estimates\footnote{The point estimates differ slightly across the full-information, limited-information, and independent methods since we use different weighting matrices for each approach. These minor discrepancies are discussed in more detail in Footnote 18 of mikkel.} for the number of products, $\hat{m}$, the volatility of the desired prices of the multiproduct firm, $\hat{\nu}$, and the scaled menu cost with respect to the curvature of the firm's profit function $\widehat{\sqrt{\psi/B}}$ and their corresponding standard errors are reported for each framework in \hyperref[table:menucost]{Table (ref)}.
The Full Information, Independent, and Worst-case $\bar{\sigma}_{\theta}^{CP}/\sqrt{n}$ entries of \hyperref[table:menucost]{Table (ref)} replicate the results of mikkel. Our study's main contribution is the computation of the standard errors under the best-case framework, $\underline{\sigma}_{\theta}/\sqrt{n}$, and under the best-case framework with additional information on the correlation structure, $\underline{\sigma}_{\theta}^{\rho_{23}}/\sqrt{n}$. For the latter column, we invoke Proposition (ref) and make use of the fact that we can compute the correlation between the second and third moments of the data. This allows us to incorporate an additional constraint into the SDP and leverage information that tightens our computation of the best-case standard errors.
The best-case standard errors are substantially smaller than the worst-case standard errors. While the worst-case standard errors are already small, implying that the point estimates of the alvarez&lippi:2014 structural model parameters are statistically significant, we find that the best-case standard errors are notably smaller and effectively zero. However, when we make use of the estimable correlation between the third and fourth moments\footnote{In this setting, as we have access to the full data, we could have used the correlations between other pairings of moments. All other pairings gave similar results with the lower bounds of the standard errors for $\hat{\nu}$ and $\widehat{\sqrt{\psi/B}}$ not differing significantly from the baseline case. But, when using information on the correlation between the second and fourth moments, the best-case standard error for $\hat{m}$ takes on the value of 0.00394. On the other hand, utilizing the correlation between the third and fourth moments yields a lower bound of 0.0182 for the standard error of $\hat{m}$.}, the best-case standard errors for the parameter $\hat{m}$ increase to about 0.8% of the point estimate from being effectively zero.
A few interesting comparisons can also be made with the other estimation frameworks. The next smallest set of standard errors can be found under the full information approach. The first row of \hyperref[table:menucost]{Table (ref)} shows that using the full underlying correlation structure for minimum distance inference leads to standard errors that are much closer to the best-case standard errors. This tells us that, for the number of products and Menu Cost parameters, the best-case standard errors consist of over half of the total weight of the full-information standard errors. Moreover, on average, the distance between the full information standard errors and the best-case standard errors is smaller than the distance between the full information standard errors and the worst-case standard errors by approximately 0.02.
For the next application, we perform inference on the general equilibrium HANK macro model of auclertetal:2021 and mckayetal:2016 using impulse response function estimates from two separate studies: changetal:2014 and agrippino&ricco:2021. The latter two papers provide independent empirical micro cross-sectional and macro time series moments. Consequently, the covariances among the moments computed from these two papers is not immediately available for full information minimum distance estimation.
The aim of this application is to study the impacts of productivity and Monetary Policy shocks on aggregate outcomes. The HANK model relates the interactions among a continuum of heterogeneous households, monopolistically competitive firms facing a price adjustment cost, a government that gives out lump sum payments, and a central bank that sets the nominal interest rate. To characterize these interactions, this HANK inference setting concentrates on estimating $k = 7$ structural model parameters: the Taylor Rule coefficient on inflation that determines the central bank's nominal interest rate policy, the slope of the New Keynesian Phillips Curve, the two autoregressive coefficients of the AR(2) process for TFP growth, the standard deviation coefficient for the TFP growth process, and the two autoregressive coefficients of the AR(2) process for the Monetary Policy shock. In this application, $p = 23$ empirical moments are taken from changetal:2014 and agrippino&ricco:2021 for the calibration procedure.
We follow a modified version of the one-asset HANK model discussed in Appendix B.2 of auclertetal:2021. In this model, the economy consists of heterogeneous households that choose the amount they consume, the quantity they save in a nominal Treasury bond, and the hours they work. They receive lump sum payments from the government and profits from firms that they own. The firms are monopolistically competitive firms that set prices according to a quadratic adjustment cost, implying a New Keynesian Phillips Curve. The government issues one-period nominal bonds and sets taxes to balance its periodic budget. Following a standard Taylor Rule, the central bank determines the nominal interest rate on bonds. Markets clear when the total goods produced by the firms equals private consumption, public consumption, and the price adjustment costs; labor demand equals labor supply in efficiency units; and aggregate household savings equals government bonds.
To solve this model, we use the first-order linearization method of auclertetal:2021. We only estimate the structural model parameters that do not determine the steady state of the model. While it is feasible to estimate parameters that influence the steady-state, this simplification allows us to avoid recomputing the model, evading computational issues. Hence, identical to the approach taken by mikkel, we fix the steady-state parameters at the values in Table B.2 of auclertetal:2021.
changetal:2014 and agrippino&ricco:2021 provide two Structural Vector Autoregressions with point estimates and standard errors of impulse responses with respect to the identified shocks. These estimates allow us to study the impact of productivity (TFP) and Monetary Policy shocks.
From changetal:2014, we directly obtain the estimates corresponding to the TFP shock. In particular, we use the reported values of the blue solid lines of Figures 9 and 11 of changetal:2014 that represent distributional innovations or productivity shocks. The three impulse responses of (1) GDP or aggregate output (in the model), (2) TFP, and (3) the fraction of people with earnings smaller than $2/3$ of GDP per capita are used for this application. The last response variable is taken from cross-sectional data of the Current Population Survey (CPS). Although the CPS has a limited sample size, as noted in changetal:2014, the estimation approach for these impulse responses already takes into consideration the uncertainty caused by the small CPS samples.
On the other hand, from agrippino&ricco:2021, the impulse response estimates used for this application are induced by Monetary Policy shocks. We use the reported impulse response estimates from the solid blue lines of Figure 3 of agrippino&ricco:2021 for the responses of aggregate output (referred to as “industrial production” in the original paper), price level (the “consumer price index”), and the model's annualized nominal interest rate (“1-year Treasury rate”). To match the quarterly structural model, we use the end-of-quarter impulse response estimates.
For the moment-matching procedure, we concentrate on a few impulse horizons and follow mikkel. First, we only study four impulse response horizons: the impact horizon, the 1-quarter horizon, the 2-quarter horizon, and the 8-quarter horizon. Second, we account for the fact that agrippino&ricco:2021 impulse responses are normalized such that an impact or shock will generate a 100 basis point increase in the Treasury rate, while changetal:2014 impulse responses are with respect to a shock of one standard deviation. Finally, following the Bernstein-von Mises theorem, we interpret the reported point estimates of the paper as posterior medians and the standard errors as ones implied by credible intervals with a normal approximation. This is done because both changetal:2014 and agrippino&ricco:2021 report Bayesian posterior quantiles. Accounting for all this, our application makes use of $p = 23$ empirical moments.
We estimate the $k = 7$ structural model parameters with the $p = 23$ moments using a diagonal weighting matrix of the inverse marginal variances of each moment: $\hat{W}_{j,j} = 1/\hat{\sigma}_{j}^{2}$. Each minimum distance estimate reported in Table (ref) is an overall optimum resulting from gradient-based numerical optimization with 100 different starting values. Details of the procedure in generating the starting values can be found in Footnote 27 of mikkel.
The estimation approach is typical in the HANK setting since models of this type often imply nonlinear and non-convex objective functions for the minimum distance framework. These objective functions are normally difficult to manipulate analytically (leading to the need for numerical approximation). Moreover, the sometimes non-convex nature of these functions can lead to estimates that only reflect local optima rather than global optima.
In Table (ref), under the worst-case standard errors, the two autoregressive coefficients for the TFP shock process and the second autoregressive coefficient for the monetary shock process are implied to be statistically insignificant. However, the gap between the best- and worst-case standard errors for the three parameters is rather large as the best-case standard errors are either effectively 0 or a few decimal places smaller in magnitude compared to the worst-case standard errors. Although the first autoregressive parameter for the monetary shock is unambiguously positive in the estimate of the upper bound, the worst-case approach is uninformative about the impact of TFP impulse responses, an equally vital component of the model. The best-case standard errors then offer an alternative perspective on the TFP parameter estimates and highlight how the conjunctive use of the best- and worst-case standard errors can be of value to researchers.
In this application, the best-case and worst-case standard errors disagree by a substantial amount. As argued in Section (ref), this signals that it may be worthwhile to estimate the underlying correlation structure among the moments used in changetal:2014 and agrippino&ricco:2021. Given the large gap between the best- and worst-case standard errors, three out of the seven structural parameters of the HANK model have confidence intervals that may or may not contain zero. As there is a lot of uncertainty over the estimates and interpretations of the results of the minimum distance approach, exam ining the correlation structure of the empirical moments becomes a reasonable next step.
We now apply our method to the two-sample instrumental variables (TSIV) setting originally proposed by klevmarken_missing_1982 and popularized by angrist1992effect. In this framework, the researcher observes two samples that do not contain the full set of variables required for standard IV estimation, but share a common set of instruments. For example, one sample contains the outcome and instruments, and the other contains the endogenous regressor and instruments. We focus on the just-identified version of the two-sample 2SLS (TS2SLS) estimator, which is commonly used for its computational simplicity and statistical efficiency (inoue_two-sample_2010). TS2SLS has been widely used in the literature on intergenerational mobility (olivetti_name_2015; barone_intergenerational_2021), in public-policy studies (currie_are_2000; deng_early-life_2022), and in credit-supply shocks in household and firm contexts (crossley_house_2024; renkin_credit_2024).
Let sample 1 consist of observations $\{(y_{1i}, z_{1i})\}_{i=1}^{n_1}$, and sample 2 consist of $\{(x_{2i}, z_{2i})\}_{i=1}^{n_2}$, where $y_{1i} \in \mathbb{R}$ is the scalar outcome, $x_{2i} \in \mathbb{R}^k$ is the endogenous regressor, and $z_{1i}, z_{2i} \in \mathbb{R}^q$ are instruments. Let $Y_1$, $X_2$, $Z_1$, and $Z_2$ denote the matrices stacking these vectors across $i$. For simplicity, consider a just-identified case where $q=k$. The TS2SLS estimator is then given by:
This expression shows that $\hat{\theta}$ is a smooth function of two empirical moments
Hence, we can write $\hat{\theta} = h(\hat{\mu}_1,\hat{\mu}_2)$ for a differentiable map $h$, and apply the Delta method to obtain the inference for $\hat{\theta}$. Under the usual two-sample assumption, the two moment vectors are treated as independent; the covariance block linking $\hat{\mu}_1$ and $\hat{\mu}_2$ is set to zero. This restriction is stronger than requiring the two datasets to contain different individuals: it requires the sampling errors in the estimated first-stage and reduced-form moments to be uncorrelated. Our method relaxes that restriction and delivers sharp lower and upper bounds for the asymptotic variance of $\hat{\theta}$. In the context of currie_are_2000, this relaxation is empirically relevant. Although the CPS and the 1990 Census do not share individual observations, they cover overlapping cohorts and labor markets and are subject to common aggregate conditions and policy environments. These shared features plausibly induce correlation in sampling error across the estimated first-stage and reduced-form moments.
We illustrate this using the influential empirical application in currie_are_2000, which studies the effect of public housing on children’s outcomes. The endogenous variable is whether a household lives in public housing, and the instrument is an indicator of whether the household has children of different genders. This instrument exploits eligibility rules from the Department of Housing and Urban Development, under which families with mixed-gender children are eligible for larger housing units and thus more likely to enter public housing. The outcome variables are drawn from the 1990–1995 CPS March Supplement (sample 2), while the household covariates come from the 1990 Census (sample 1).\footnote{ We use the cleaned CPS and Census extracts provided in the replication package of choi2018weak.} The analysis focuses on four outcomes, as in Table 4 of currie_are_2000: monthly rent payments, overcrowding (fewer than three rooms), residence in large buildings with more than 50 units, and whether any child has been held back in school.
Estimation results are reported in Table (ref). The first column reproduces the point estimates in currie_are_2000. The next two columns report conventional standard errors from currie_are_2000 and inoue_two-sample_2010, respectively. The remaining columns present explicit sharp bounds on the standard errors derived in Section (ref) under diagonal and block-diagonal dependence structures. Since the block-diagonal covariance structure is fully known and the analytic formulation is feasible in this setting, there is no need to solve the equivalent semidefinite program numerically. Accordingly, we compute the bounds directly using the closed-form expressions in Section (ref). The upper bound for standard errors in each case coincides with the standard error of the worst-case proposed in mikkel.
We first observe that the conventional standard errors lie between our lower and upper bounds. The diagonal-only case is largely uninformative, with the feasible set spanning from near zero to relatively wide upper bounds. In contrast, imposing the block-diagonal structure dramatically tightens the range, yielding informative and economically meaningful uncertainty estimates. These results suggest that accounting for within-block correlation can substantially reduce the identified set for standard errors without relying on independence assumptions. Because there are only two moments in this application, the block-diagonal constraint is especially informative.
In this paper, we contribute to the literature on bounding the standard error of parameters estimated from moment conditions across different samples. First, we derive an explicit sharp lower bound in settings where no information about cross-sample correlations is available. Second, we develop computationally feasible sharp bounds based on a semidefinite program, which delivers the best-case standard error in more general settings. Empirical illustrations in both macroeconomic and microeconomic contexts demonstrate the practical usefulness of our approach. Together with mikkel, this paper provides a comprehensive set of tools for computing sharp lower and upper bounds in a wide range of situations. While the existing literature focuses primarily on calibrated parameters in macroeconometric settings, we expand the scope of applications to include two-step estimators that are also relevant to microeconometric settings.
A key takeaway from the empirical applications is how the informativeness of the bounds depends on the structure of the reduced-form moments and on the amount of auxiliary information available. As the analytic expressions in Section (ref) illustrate, the bounds can become extremely wide when the number of reduced-form moments increases: the lower bound often collapses to zero, while the upper bound can grow very large, yielding limited guidance for inference. However, as shown in Sections (ref) and (ref), incorporating additional information about the correlation matrix can substantially sharpen the bounds, bringing the lower and upper bounds closer and thereby producing a more informative range for the plausible standard errors. In the menu-cost application, even a single strong piece of information leads to a dramatic tightening of the bounds. In the TS2SLS case, with only two reduced-form moments and a block-diagonal restriction on the correlation matrix, the resulting bounds are highly informative, offering practical guidance for empirical researchers while remaining agnostic about arbitrary cross-dataset dependence.