Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
154,556 characters · 37 sections · 71 citation commands
Nonparametric Identification and Estimation of Production Functions Invariant to Productivity Dynamics
\IfFileExists{Table/Empirical/cross_val_macros.tex}
Estimated production functions underpin the measurement of market power, allocative efficiency, and the effects of policy on firm performance. The ratio of the materials elasticity to the materials revenue share gives the firm-level markup deloecker2012markups; the dispersion of the productivity residual measures resource misallocation hsieh2009misallocation; and the productivity level itself serves as the outcome variable in studies of trade liberalization deloecker2013detecting, R&D investment doraszelski2013rdand, and disaster recovery. These downstream analyses inherit the production function estimate: if the materials elasticity is biased, so is the markup, the misallocation measure, and the treatment effect. The recent finding that markups have risen across the global economy deloecker2020rise relies on such estimates, making the consistency of the underlying production function a first-order concern. This paper asks whether the production function can be identified without restricting how productivity evolves over time, and documents the consequences when this restriction is removed.
Since olley1996thedynamics, every major production function estimator has relied on the same structural restriction: productivity must follow a first-order Markov process. This includes the methods of levinsohn2003estimating, ackerberg2015identification, and gandhi2020onthe, as well as dynamic panel approaches arellano1991sometests,blundell1998initial. The Markov assumption is not a regularity condition; it is the identifying restriction that pins down the materials elasticity through the transition equation. When productivity evolves endogenously through R&D, learning, or managerial turnover, omitting the relevant state variables generates a transmission bias deloecker2007doexports,deloecker2013detecting,doraszelski2013rdand. The assumption also presupposes a stationary transition process, ruling out structural breaks from aggregate shocks, regulatory shifts, or technological change. More fundamentally, chen2024identifying show that under the potential outcomes framework, any treatment that alters the transition path of productivity violates the Markov property by construction, even when the treatment variable is included as a control. The bias does not vanish with sample size, nor can it be removed by adding treatment indicators to the Markov transition equation. The Markov-based estimate is therefore inconsistent precisely in the settings where productivity serves as an outcome variable, the dominant use of production function estimation in applied work.
This paper shows that the Markov assumption is unnecessary for identification. I replace it with a static condition: conditional independence of demand shocks across three intermediate inputs (raw materials, electricity, water). Three flexible inputs whose demands respond to the same underlying productivity serve as three noisy measurements of a common latent variable. Because each input is procured from a separate market, the input-specific demand shocks are mutually independent conditional on productivity and observable controls. I recover the productivity distribution from these signals using the spectral decomposition of hu2008instrumental (hereafter HS08), without any restriction on how productivity evolves over time. Identification requires only a single cross-section; the data requirement (firm-level quantities of three separate inputs) is met in manufacturing censuses across several countries.\footnote{These include India's Annual Survey of Industries, Canada's Annual Survey of Manufacturing and Logging, the World Bank Enterprise Survey, and the U.S.\ EIA Form 923. When labor adjusts rapidly to current productivity, two intermediate inputs suffice (footnote (ref)).}
The substitution of assumptions has first-order consequences for economic measurement. In 502 Japanese manufacturing industries, the proposed method yields systematically lower markups than the standard ACF method across the entire distribution: the ACF markup CDF lies strictly to the right at every percentile. At the median, the gap is 0.10 (proposed 0.93 vs.\ ACF 1.03), and the share of industries with markups above unity falls from 54 percent under ACF to 37 percent under the proposed method. The Markov assumption thus inflates the measured degree of market power across the manufacturing sector. Monte Carlo simulations trace the mechanism: under potential outcome dynamics, ACF's bias in the materials elasticity is $+0.19$ (63 percent of the true value), while the proposed estimator is unbiased.
In a difference-in-differences analysis of the 2011 T\={o}hoku earthquake, the standard method overstates the productivity loss by 0.40 percentage points, corresponding to roughly \$3.6 billion (\textyen400 billion) per year. Because identification is static, the estimator can be applied period by period, producing time-varying estimates of production technologies without imposing structural stability on the productivity process. The empirical application documents substantial temporal variation across 2003--2020 and yields divergent conclusions regarding allocative efficiency as assessed through the olley1996thedynamics decomposition. The Markov assumption does not merely introduce statistical noise; it systematically inflates measured market power and distorts policy conclusions.
The substitution involves an honest tradeoff. The Markov assumption, when it holds, provides efficiency gains by exploiting the time-series history of productivity. The conditional independence assumption uses only within-period information, so under correct Markov specification, standard estimators have lower variance. I document this in Monte Carlo simulations under correct Markov specification. The value of the proposed method lies in the broad class of applications where the Markov assumption is questionable or directly contradicted by the research design, including any study in which a treatment alters productivity dynamics chen2024identifying.
The two assumptions differ in the nature of their economic content. The Markov restriction constrains the time-series evolution of unobserved productivity; no economic theory predicts that productivity should follow a first-order autoregression, and the assumption cannot be tested within the proxy variable framework. The conditional independence restriction constrains input market structure: it specifies what threatens identification (common shocks across input markets) and what restores it (conditioning on observable controls $z_{jt}$ that absorb the common component). The threats are enumerable (demand fluctuations, aggregate markup changes, correlated procurement), and the defenses are observable (inventory, aggregate output, fixed effects; Section (ref)). The microfoundations in Appendix (ref) derive the demand shocks from a cost-minimization problem with input-specific markdowns, making the economic content of the assumption precise. No analogous transparency is available for the Markov assumption: within the proxy variable framework, no observable implication distinguishes a correctly specified AR(1) from an AR(2) or a potential outcome process. By contrast, the conditional independence assumption yields a testable necessary condition (Remark (ref)): in 502 industries, the pairwise convergence diagnostic supports the identifying restriction for capital, providing direct evidence on the empirical plausibility of the assumption. When the assumption is violated through a common shock to electricity and water (the most economically salient threat), Monte Carlo analysis (Section (ref), Appendix (ref)) shows that the resulting bias in $\hat{\beta}_m$ is upward, the same direction as the Markov misspecification bias. The empirical finding that the proposed method yields lower $\hat{\beta}_m$ than ACF therefore cannot be explained by conditional independence violation; it is consistent only with Markov misspecification in the standard estimator.
I make three contributions. First, I show that the cross-sectional covariance structure among three flexible intermediate inputs fully substitutes for the Markov restriction, delivering nonparametric identification of the production function and the productivity distribution from a single period. This is a substitution, not a relaxation, of identifying assumptions. The mapping to the HS08 framework provides density identification (Theorem (ref)); the theoretical contribution of this paper lies in what follows. I characterize the residual indeterminacy that arises when Markov is dropped: any two observationally equivalent structures differ only by a location shift $\Delta(k,l)$ applied to productivity, ruling out nonlinear transformations (Theorem (ref)). I provide two routes that close this indeterminacy without dynamic assumptions, an exclusion restriction (Corollary (ref)) and a homothetic regularity condition (Theorem (ref)).
While nonparametric sieve estimation could in principle implement the identification results directly, the high-dimensional numerical integration is computationally prohibitive for census-scale panels. I develop a Cobb--Douglas GMM estimator designed for applied use, and establish its consistency and asymptotic normality (Theorem (ref)); the extension to translog production is developed in Appendix (ref). The conditional independence assumption yields a pairwise convergence diagnostic (Remark (ref)) with no analogue under the Markov assumption: within the proxy variable framework, no restriction distinguishes a correctly specified AR(1) from an AR(2) or a potential outcome process. In 502 industries, this diagnostic converges to zero for capital but not for labor, providing direct evidence on the differential applicability of the exclusion restriction.
Second, I document that the Markov assumption generates a systematic upward bias in measured market power. Monte Carlo simulations show that ACF's bias in $\hat{\beta}_m$ does not vanish as sample size grows: $+0.026$ under AR(2) dynamics and $+0.266$ under potential outcome dynamics. In the empirical application, ACF produces higher materials elasticities and higher markups at every percentile across 502 industries. The gap crosses the competitive threshold and reverses the policy-relevant conclusion about market structure. The recovered productivity measures also show stronger associations with economic fundamentals than those from the standard method, consistent with a higher signal-to-noise ratio from separating input-specific demand shocks (Section (ref)).
Third, I show that productivity measures recovered from the proposed method are valid under the potential outcomes framework (Proposition (ref)), resolving the inconsistency identified by chen2024identifying. Because the estimator uses no transition equation, the recovered productivity is invariant to how a treatment operates on productivity dynamics. The earthquake event study illustrates the practical consequence: the proposed estimate is $-1.28$ percent while ACF yields $-1.68$ percent, a gap that arises because the ACF estimate lacks the theoretical guarantee that the production function parameters are consistently estimated under treatment-induced dynamics. The same static, $\omega$-conditional structure also renders the estimator robust to endogenous exit: under the standard timing convention where exit precedes input choice, conditioning on $\omega$ absorbs survival selection, and no survival probability correction is needed (Remark (ref)).
\paragraph{Related literature.}
Table (ref) positions my identification strategy within the recent literature. The most closely related work is gandhi2020onthe (GNR). GNR's Theorem 1 establishes that proxy variable methods alone cannot identify the gross output production function; additional within-period, cross-sectional information is required. Both approaches supply such information: GNR through the structural link between the production function and the firm's first-order condition, yielding a nonparametric share regression that directly identifies the flexible input elasticity; my approach through the measurement error structure of HS08, using conditional independence across intermediate inputs to recover the distribution of unobserved productivity.
The two approaches rest on different assumptions regarding input markets. GNR's first-order condition requires competitive input markets with common prices and that any unobserved component in the share equation is non-persistent (their Appendix O6, Assumption 7); when input-specific markdowns or procurement frictions are persistent, the FOC-based estimation equation does not hold and the share regression is misspecified. My framework permits persistent, input-specific demand shocks arising from procurement relationships, supply contracts, or input-specific markdowns; identification requires only mutual independence across inputs at each time point, accommodating arbitrary serial dependence within each shock. GNR's second stage recovers capital and labor elasticities using the Markov structure; my approach requires no dynamic assumption at any stage. The scalar unobservability case is a special case of my model, obtained when the input-specific shocks are degenerate (Section (ref)).
Alternative approaches that exploit static first-order conditions grieco2016production,caselli2025productivity avoid dynamic assumptions but generally require parametric restrictions on functional forms and the demand system. Additional related work is summarized in Table (ref). Several recent papers apply the HS08 framework to production functions brandestimating,hu2020estimating,dotyadynamic, but all use lagged variables as instruments and therefore retain the Markov assumption. zeng2023identification avoid the Markov restriction at the estimation stage but presuppose it for the investment policy function. A growing literature on non-Hicks-neutral identification navarrononparametric,ackerberg2022nonparametric,pan2022identification,kasahara2023identification,dotyadynamic, including factor-augmenting approaches doraszelski2018measuring,demirerproduction,raval2019themicro, retains first- or higher-order Markov assumptions; my identification results extend to these models without dynamic restrictions (Appendix (ref)), though the implemented estimator uses the Hicks-neutral Cobb--Douglas specialization (Section (ref)).
The remainder of the paper is organized as follows. Section (ref) presents the model and the nonparametric identification results. Section (ref) develops the GMM estimator. Section (ref) presents Monte Carlo evidence. Section (ref) applies the estimator to 502 Japanese manufacturing industries. Section (ref) concludes.
This section establishes the identification strategy in three steps. First, I show that three conditionally independent input demands identify the joint distribution of productivity and inputs within each capital-labor cell (Theorem (ref)); any two observationally equivalent structures differ only by a location shift $\Delta(k,l)$ (Theorem (ref)). Second, I provide two conditions that eliminate this indeterminacy: an exclusion restriction (Corollary (ref)) and a homothetic regularity condition (Theorem (ref)). The exclusion restriction carries a testable implication (Remark (ref)). The formal statement of density identification (Theorem (ref)) and the technical regularity conditions (Assumptions (ref)--(ref)) are in Appendix (ref). These identification results translate into three groups of moment conditions in the GMM estimator of Section (ref): proxy moments (Block A), covariance moments (Block B), and curvature moments (Block C). When these terms appear below, they refer forward to Section (ref).
I define the general gross output production function for firm $j$ at time $t$ as follows:
Here, $y_{jt}$ is the logarithm of output, $k_{jt}$ and $l_{jt}$ are the logarithms of capital and labor. Following the production function literature olley1996thedynamics,ackerberg2015identification,bond2005adjustment, capital and labor are treated as dynamic or quasi-fixed inputs whose current values are predetermined relative to intermediate input decisions. The model requires at least three distinct intermediate inputs: $m_{jt}$ (raw materials), $e_{jt}$ (electricity), and $w_{jt}$ (industrial water). Three inputs are the minimum required by the hu2008instrumental spectral decomposition: it identifies the latent productivity distribution from three mutually independent measurements of a common latent variable; two measurements do not suffice for nonparametric identification without additional restrictions.\footnote{When labor adjusts rapidly to current productivity, it may serve as a third measurement of $\omega_{jt}$, reducing the required number of flexible intermediate inputs from three to two; see Footnote (ref) for details.} $\omega_{jt}$ is the firm's productivity, unobserved by the econometrician but known to the firm when making input decisions. $\varepsilon_{jt}$ denotes ex-post production shocks (measurement error or unexpected disruptions), unobserved by both the firm and the econometrician at the time of input choice.
The state variable vector $x_{jt}=(k_{jt},l_{jt},z_{jt})$ determines input demand. Here, $k_{jt}$ and $l_{jt}$ are the primary inputs, while $z_{jt}$ represents additional firm-specific state variables such as inventory levels, input prices, or market conditions that do not directly enter the production function but influence input demand. Given $x_{jt}$, the demand for each intermediate input is determined as follows:
The functions $g_{m}(\cdot)$, $g_e(\cdot)$, and $g_w(\cdot)$ are unknown and potentially nonlinear. $\tau_{jt}$, $\nu_{jt}$, and $\eta_{jt}$ are unobserved shock terms specific to each input demand, following \textcites{hu2020estimating}{brandestimating}{dotyadynamic}. These shocks capture optimization errors, supply disruptions, and adjustment frictions not explained by productivity and state variables. Appendix (ref) derives the demand system from a cost-minimization problem under imperfect input markets and shows that these shocks correspond to input-specific markdowns, prices, and wedges; specifically, the components of markdowns and wedges orthogonal to observable state variables.
The presence of input-specific shocks represents a departure from the scalar unobservability assumption maintained in olley1996thedynamics, levinsohn2003estimating, ackerberg2015identification, GNR, and others, which requires productivity to be the sole unobservable affecting input demand. When scalar unobservability fails because firm-level input prices, markdowns, or wedges are unobserved, standard proxy variable estimators are inconsistent jaumandreu2021reexamining,doraszelski2025production. In my framework, all unobserved firm-specific heterogeneity beyond productivity is absorbed into $\tau_{jt}, \nu_{jt}, \eta_{jt}$, and identification requires only that these shocks be mutually independent across inputs, not that they be absent. Scalar unobservability is nested as the special case $\tau = \nu = \eta = 0$ at the model level; the identification strategy requires non-degenerate demand shocks and is therefore complementary to, rather than a generalization of, scalar inversion methods. From the standpoint of the cost-minimization model in Appendix (ref), $\tau = \nu = \eta = 0$ requires that all firms in an industry face identical input prices, identical markdowns in every input market, and make no optimization errors in input choice. In practice, firms negotiate procurement contracts individually, face supplier-specific delivery terms, and adjust input quantities with heterogeneous frictions. The presence of input-specific demand shocks is the empirically relevant case; the proposed framework treats these shocks as a source of identifying information rather than a nuisance to be assumed away.
This formulation also addresses the collinearity problem identified by gandhi2020onthe: under scalar unobservability, flexible inputs determined by static optimization lack sufficient residual variation to identify the gross production function ackerberg2015identification,bond2005adjustment. GNR resolve this problem by exploiting the first-order condition for the flexible input, which identifies its output elasticity from the revenue share. My approach resolves the collinearity through independent input-specific shocks, which supply the cross-sectional variation needed for identification via the measurement error structure of HS08, without relying on the first-order condition or dynamic moment conditions. The practical difference is that the share regression requires the first-order condition to hold with common input prices, whereas my approach permits firm-specific input prices and markdowns (Appendix (ref)).
The identification theory rests on two substantive assumptions stated here, together with three regularity conditions (Assumptions (ref)--(ref)) collected in Appendix (ref).
Role and economic content. This is standard in the production function literature olley1996thedynamics,ackerberg2015identification. The shock $\varepsilon_{jt}$ captures ex-post deviations (measurement error, unexpected disruptions) that are realized after input choices are made and are therefore uncorrelated with all inputs and productivity. It acts as classical measurement error in the dependent variable and inflates standard errors but does not bias the production function estimates (Theorem (ref)).
Role. This is the substantive identifying condition. Together with the regularity conditions in Appendix (ref) (Assumptions (ref)--(ref)), it enables the unique spectral decomposition of the integral equation (ref). Conditional independence is the economically substantive condition; it restricts the data generating process rather than regularity of the operators.
Economic content. The assumption posits that, for a firm with given state variables and productivity level, an unexpected shock to raw material demand (e.g., a supply chain disruption) is independent of a shock to electricity demand (e.g., an unscheduled rate surcharge). This is natural when input markets are segmented: raw materials, electricity, and water are procured through distinct channels, under separate contracts, with different suppliers. The common components of demand variation (product demand fluctuations, aggregate markup changes) are captured by $x_{jt}$; $\tau_{jt}, \nu_{jt}, \eta_{jt}$ represent the residual, input-specific components. The microfoundations in Appendix (ref) make this structure precise.
The general principle is as follows. Common shocks that affect all three input demands (product demand fluctuations, markup variation, aggregate input price movements) can be absorbed by projecting onto observable control variables $z_{jt}$; the shock terms $\tau_{jt}, \nu_{jt}, \eta_{jt}$ are then defined as the orthogonal residuals of this projection (Appendix (ref)). The independence assumption therefore requires only that the residual, input-specific components of demand variation are mutually independent.
Several potential threats illustrate this principle. Unobserved demand shocks generate common variation across all inputs, but can be proxied by inventory fluctuations kumar2019productivity or recovered from revenue data kasahara2020nonparametric, included in $z_{jt}$. Product market power affects all input demands through marginal revenue; following ackerbergproduction,jaumandreu2025robustproduction, low-dimensional sufficient statistics for the markup (e.g., competitors' output, average variable cost) can be included in $z_{jt}$.\footnote{Under Cournot competition, ackerbergproduction show that the total output of competitors serves as a sufficient statistic.} Input market power (markdowns) may generate common bargaining advantages across inputs, but the common component depends on firm attributes (size, liquidity) captured by $(k_{jt}, l_{jt}, z_{jt})$; what remains in the shock terms are idiosyncratic variations from individual supplier relationships. It is economically reasonable that the outcome of negotiations with raw material suppliers is independent of electricity rate negotiations, conditional on firm size and other observables.\footnote{When an intermediate input is traded on competitive commodity markets, the firm is a price-taker and the markdown on that input vanishes. avignon2025markups exploit this property for globally traded dairy commodities to separately identify markups and markdowns on other inputs.} Common input price shocks (e.g., oil price hikes) affect multiple inputs symmetrically and are controlled by time fixed effects or industry-specific deflators in $z_{jt}$. Firm-specific price variations are absorbed as part of the structural shock terms and need only be independent across inputs.
The identification proceeds in two stages: first, I recover the production function and productivity distribution within each capital-labor cell $(k_0, l_0)$; second, I characterize and resolve the residual indeterminacy that arises when linking these cell-specific results across different values of $(k, l)$. \footnote{In the following, firm subscripts $j$ are suppressed as I discuss population-level arguments. The time subscript $t$ is retained only to indicate time-variation in the production function $f_t$.}
The foundational identification result applies the spectral decomposition of HS08, whose conditions I verify under the present assumptions.
The proof, which verifies the conditions of HS08's Theorem 1 for the integral equation (ref), is in Appendix (ref).
As a consequence of Theorem (ref) and equation (ref), for each fixed $(k_0, l_0)$, the following are nonparametrically identified: the conditional densities $f_{m \mid \omega, k_0, l_0}$, $f_{e \mid \omega, k_0, l_0}$, $f_{w \mid \omega, k_0, l_0}$, and $f_{\omega \mid k_0, l_0, m, e, w}$.
Using these identification results, I recover the structure of $f_t$ as a function of $(m, e, w, \omega)$. I focus on the Hicks-neutral specification, widely adopted in the empirical literature, and defer the general case to Appendix (ref). Under this specification $y = g_t(k, l, m, e, w) + \omega + \varepsilon$, Assumption (ref) implies
Here $g_t$ represents the component of the production technology that depends on intermediate inputs, with productivity $\omega$ separated out. The first term on the right-hand side is a conditional expectation identified directly from the data, and the second is computable from the posterior density in equation (ref). Thus $g_t$ is identified as a function of $(m, e, w)$ without additional assumptions. For the general non-Hicks-neutral model, $f_t(k_0, l_0, m, e, w, \omega)$ is identified as a function of $(m, e, w, \omega)$ under additional regularity conditions on the distribution of $\varepsilon$; see Appendix (ref) for details.
For each fixed $(k_0, l_0)$, the conditional distribution $f_{\omega \mid k_0, l_0, m, e, w}$ is fully characterized, and the conditional expectation
provides a firm-level productivity measure for each firm $j$ and period $t$. The empirical applications of this within-$(k,l)$ identification, including markup estimation and policy evaluation, are developed in Section (ref) after the identification theory is completed.
However, to identify $f_t$ as a function of $(k, l)$ as well, additional structure is needed. (When labor adjusts rapidly to current productivity, it serves as an additional measurement, reducing the required intermediate inputs from three to two.\footnote{When labor adjusts within the production period, it serves as a third measurement of $\omega_{jt}$, and the HS08 identification procedure (Theorem (ref)) applies to the triple $(l_{jt}, m_{jt}, e_{jt})$, reducing the required flexible intermediate inputs from three to two. This extension applies when adjustment costs are small enough that $l_{jt}$ responds to within-period productivity innovations; industries with high turnover or temporary staffing (e.g., food processing, garment manufacturing) are natural candidates. When labor is quasi-fixed, $l_{jt}$ reflects past rather than current productivity, and the conditional independence conditions for $(l, m, e)$ do not hold. See Appendix (ref) for details.}) $\omega$ must be defined on a common scale across different values of $(k, l)$. Since Theorem (ref) applies the HS08 procedure independently for each $(k, l)$, there is no automatic correspondence between the $\omega$ values identified at $(k_1, l_1)$ and those identified at $(k_2, l_2)$. I now formalize this problem.
Theorem (ref) identifies the production function within each $(k_0, l_0)$, but a practitioner needs parameters that are comparable across different capital-labor combinations. The next result shows exactly what remains unresolved and rules out the possibility that the indeterminacy takes a nonlinear form.
Under Assumptions (ref)--(ref) and the regularity conditions in Appendix (ref) (Assumptions (ref)--(ref)), the conditional densities $f_{m|\omega}$, $f_{e|\omega}$, $f_{w|\omega}$ and the marginal density $f_\omega$ are nonparametrically identified from the joint density of $(m,e,w)$ conditional on $(k,l,z)$ (Theorem (ref)). This pins down the shape of each conditional distribution but leaves a common location shift $\Delta(k,l)$ applied to the latent variable unresolved. The following theorem characterizes this residual indeterminacy completely.
The proof is given in Appendix (ref); the key steps are as follows. The HS08 eigenvalue-eigenfunction decomposition uniquely determines the functional form of each conditional density within each $(k_0, l_0)$, ruling out nonlinear transformations of $\omega$. Any remaining degree of freedom must therefore be a location shift that varies across $(k, l)$, yielding (ref). The continuity of $\Delta(k,l)$ follows from the continuous dependence of $f_{m \mid \omega, k, l}$ on $(k, l)$ (stated after Assumption (ref)) together with the perturbation theory of compact operators under simple eigenvalues (Assumption (ref); see Appendix (ref) for details).
Nonlinear transformations (including scale transformations) are ruled out because the eigenvalue--eigenfunction decomposition in HS08 uniquely determines the functional form of each conditional density within each $(k_0, l_0)$. Second, the $\Delta(k, l)$ indeterminacy arises inherently from the fact that Theorem (ref) applies the HS08 procedure independently for each $(k, l)$. Within each $(k_0, l_0)$, Assumption (ref) fixes the level of $\omega$, but the reference point of this normalization may depend on $(k_0, l_0)$. The data on conditional distributions of intermediate input demands do not contain information to unify $\omega$ levels across different $(k, l)$.\footnote{ hahn2023identification show that in dynamic approaches such as olley1996thedynamics, the identification of dynamic input elasticities relies on an index restriction that collapses state variables into a one-dimensional scalar. Theorem (ref) does not provide such an index restriction, and hence the indeterminacy with respect to the dynamic elasticities persists.}
Economically, the $\Delta(k, l)$ indeterminacy means that the effect of $(k, l)$ on the production function and $\mathbb{E}[\omega \mid k, l]$ cannot be separated without additional restrictions. As a direct consequence, $f_t$ is identified up to the specification of $\mathbb{E}[\omega \mid k, l]$: fixing $\mathbb{E}[\omega \mid k, l]$ pins down $\Delta = 0$ (Theorem (ref), Appendix (ref)).
The $\Delta(k, l)$ indeterminacy also arises in the existing literature: gandhi2020onthe resolve it in the Hicks-neutral setting by combining first-order conditions with a Markov assumption, which reduces $\Delta(k, l)$ to a constant; for non-Hicks-neutral models, this strategy fails because $\omega$ cannot be separated from the first-order condition.\footnote{In the Hicks-neutral model, $\Delta(k, l)$ shifts the $(k,l)$ component of the production function: $\tilde{g}_t(k, l, \cdot) = g_t(k, l, \cdot) - \Delta(k, l)$. In non-Hicks-neutral models, the FOC $P_t \cdot (\partial f_t / \partial M) = \rho_t$ retains $\omega$ on the left-hand side, precluding a share regression. li2024identification show that heterogeneous output elasticities with respect to flexible inputs remain identifiable under a scalar unobservable assumption on the proxy variable.}
I provide two alternative methods that close the identification gap without dynamic assumptions. Theorem (ref) guarantees that $\Delta(k, l)$ is a continuous function of $(k,l)$ alone, which both methods exploit. Section (ref) imposes exclusion restrictions on intermediate input demands that directly constrain the functional form of $\Delta(k, l)$, achieving nonparametric point identification. Section (ref) parametrically specifies the $(k, l)$ component and introduces a regularity condition on the shape of $\mathbb{E}[\omega \mid k, l]$, achieving parametric identification through the non-constant curvature of the homothetic transformation.
The $\Delta(k,l)$ indeterminacy is the cost of dispensing with the Markov assumption. I now show this cost is payable: two conditions, each operating without dynamic restrictions, eliminate the indeterminacy and deliver point identification.
The $\Delta(k, l)$ indeterminacy arises because the location normalization in Assumption (ref) is applied independently for each $(k, l)$ (Theorem (ref)). If the HS08 location normalization $M[f_{m \mid \omega, k, l}(\cdot \mid \omega)] = \omega$ could be applied uniformly across all $(k, l)$, then $\Delta(k, l) = 0$ would follow immediately. However, for this uniform normalization to hold, $M[f_{m \mid \omega, k, l}(\cdot \mid \omega)]$ must not depend on $(k, l)$; that is, the conditional demand for the intermediate input, given $\omega$, must be independent of $(k, l)$. This observation suggests that exclusion restrictions on intermediate input demands directly constrain $\Delta(k, l)$.
Economically, condition (i) requires that the demand for some intermediate input (e.g., electricity) depends on productivity alone and not on capital or labor intensity; this may hold in energy-intensive industries where electricity consumption is driven by production volume rather than by the composition of capital equipment. Condition (ii) requires that different inputs exclude different primary inputs from their demand: for example, raw material demand does not depend on labor intensity, and fuel demand does not depend on capital intensity. These exclusion restrictions limit the scope of application to industries where institutional knowledge supports them. For settings where such restrictions cannot be justified, I provide a parametric alternative in the next subsection.
The formal statement and proof are given in Proposition (ref) (Appendix (ref)).\footnote{Replacing the linear subtraction of $\hat{\omega}^h$ in Proposition (ref) with a polynomial regression is not consistent in general; see Appendix (ref) for details.}
As an alternative when exclusion restrictions cannot be justified, I parametrically specify the $(k, l)$ component and introduce a regularity condition on $\mathbb{E}[\omega \mid k, l]$. Consider the additively separable model
where $g$ is parametric with known functional form and $q$ is nonparametric. From Section (ref), $q$ is nonparametrically recoverable for each fixed $(k_0, l_0, \omega_0)$.
Specializing to $g(k, l;\, \theta) = \beta_k k + \beta_l l$, Theorem (ref) reduces the identification indeterminacy to
To eliminate this two-dimensional indeterminacy, I introduce the following regularity condition.
All three conditions are necessary for Theorem (ref): (A) prevents observational equivalence with linear functions; (B) ensures the counterfactual index is also translation homogeneous, so that the MRS of $\tilde{v}$ is translation invariant; (C) excludes Cobb--Douglas, where $v_k/v_l$ is constant and a one-dimensional indeterminacy persists. Economically, (A) requires nonlinear returns to the input bundle, (B) corresponds to constant returns to scale in the level variables (since translation homogeneity on the log scale is equivalent to degree-one homogeneity in levels), and (C) requires a finite and non-unit elasticity of substitution, satisfied by CES, translog, and normalized quadratic forms. Assumption (ref) can be checked from Blocks A and B alone (Section (ref)); detailed necessity arguments and testability procedures are in Appendix (ref).
To illustrate, consider the CES specification where $v(k, l) = \frac{1}{\rho_v}\log\bigl(\alpha e^{\rho_v k} + (1-\alpha) e^{\rho_v l}\bigr)$ is translation homogeneous on the log scale: $v(k+c, l+c) = v(k,l) + c$. With $h(v) = \gamma v$ (for $\gamma \neq 0$ and higher-order terms $\rho_2 v^2 + \rho_3 v^3$ with $\rho_2 \neq 0$ or $\rho_3 \neq 0$), $h'$ is non-constant (satisfying (A)), $v$ is translation homogeneous (satisfying (B)), and the MRS $v_k/v_l = [\alpha/(1-\alpha)] e^{\rho_v(k-l)}$ is non-constant for $\rho_v \neq 0$ (satisfying (C)). The Cobb--Douglas case ($\rho_v \to 0$, so $v \to \alpha k + (1-\alpha) l$) yields a linear $v$ and a constant MRS, violating conditions (A) and (C) simultaneously; the rank condition in Theorem (ref) fails, and $(\beta_k, \beta_l)$ cannot be separately identified. More generally, when $\rho_v$ is close to zero, identification of $(\beta_k, \beta_l)$ through Block C becomes weak: the marginal rate of substitution $v_k/v_l$ approaches a constant as $\rho_v \to 0$, so the cross-sectional variation in $(k_{jt}, l_{jt})$ provides little leverage on the curvature parameters. In the empirical analysis, the $t$-statistics for $\hat{\rho}_2$ and $\hat{\rho}_3$ (Section (ref)) provide a direct diagnostic for this failure; industries where both are statistically insignificant should not be relied upon for separate identification of $\beta_k$ and $\beta_l$ through Block C alone. When $\rho_v = 0$, the exclusion restriction of Corollary (ref) provides an alternative identification route.
Theorem (ref) is stated and proved for the CES specification of $v(k,l)$; the argument extends to other parametric forms (e.g., translog) subject to verifying the rank condition specific to each functional form.\footnote{For instance, with a translog specification $g = \beta_k k + \beta_l l + \beta_{kk} k^2 + \beta_{ll} l^2 + \beta_{kl} kl$, $\Delta$ is restricted to the corresponding polynomial class and the homothetic regularity condition eliminates the indeterminacy by a similar argument, but the conditions on the MRS differ from the CES case.}
The within-$(k_0, l_0)$ identification results have direct empirical applications that differ in what they require. Markup estimation requires only $\beta_m$, which is identified by Blocks A and B alone. Event studies and difference-in-differences designs similarly require only Block A+B: because the estimator uses no transition equation for $\omega$, the recovered $\hat{\omega}_{jt}$ is valid under any productivity dynamics, including treatment-induced non-Markov paths (Proposition (ref), Appendix (ref)). Full productivity-level analysis (including the identification of $\beta_k$ and $\beta_l$) requires Block C in addition.
\paragraph{Applications.} Because estimation does not employ a transition process for $\omega$, the estimates are invariant to how a policy $D_{jt}$ affects productivity dynamics (Proposition (ref), Appendix (ref)). For markup estimation, the within-$(k_0, l_0)$ results suffice: output elasticities $\partial f_t / \partial m$ are identified for each fixed $(k_0, l_0)$, which recovers markups as the ratio of the output elasticity to the revenue share deloecker2012markups.
The nonparametric identification results of Section (ref) establish that the production function and productivity distribution are identified from the joint density of intermediate inputs; nonparametric sieve estimation could in principle implement this directly, but the high-dimensional numerical integration required is computationally prohibitive for census-scale panels spanning hundreds of industries. I therefore develop a GMM estimator that specializes to a linear production function and linear demand functions. Under this parametric restriction, the observational equivalence class of Theorem (ref) reduces to a two-dimensional indeterminacy $(c_k, c_l)$ (equation (ref)), and the identification results of Corollary (ref) and Theorem (ref) carry through directly.
As noted in Remark (ref), the identification results of Section (ref) apply to general production functions. The parametric implementation below specializes to the Cobb--Douglas case, where input demand functions are linear in productivity (Appendix (ref)). This linearity yields the tractable linear GMM system of Blocks A--B. Extension to flexible functional forms such as translog is developed in Appendix (ref); the identification source remains the conditional independence of demand shocks.
The GMM estimator jointly recovers the production function and demand parameters from three blocks of moment conditions:
Blocks A and B identify the intermediate input elasticities $(\beta_m, \beta_e, \beta_w)$, the demand function parameters $(\theta_g, \psi_\omega)$, and certain composite functions of $(\beta_k, \beta_l)$ and the demand slopes. However, as shown in Section (ref), these blocks alone cannot separate $\beta_k$ and $\beta_l$ from the demand function slopes on $(k, l)$ due to the $\Delta(k,l)$ observational equivalence (Theorem (ref)). Block C resolves this indeterminacy through the nonlinear curvature of $\mathbb{E}[\omega \mid k, l]$ imposed by Assumption (ref), thereby achieving point identification of all structural parameters (Theorem (ref)). When its identifying conditions are weak, the exclusion restriction of Corollary (ref) provides an alternative route. Figure (ref) (Appendix (ref)) provides a visual overview of the full estimation and inference pipeline, including the diagnostic branches that determine which identification route applies.
The parametric specialization below implements the identification results of Section (ref) under additive separability; this restriction reduces the nonparametric problem to a finite-dimensional GMM system while preserving all theoretical properties of Theorems (ref)--(ref). To apply GMM, I impose additive separability on both the production and demand functions.
\paragraph{Production function.} Following the parametric model of Section (ref), the production function is specified as:
Here $g(k,l;\theta) = \beta_k k + \beta_l l$ is the parametric $(k,l)$ component and $q(m,e,w) = \beta_m m + \beta_e e + \beta_w w$ is the (linear) intermediate input component, corresponding to the additively separable model (ref).
\paragraph{Demand functions.} The intermediate input demands take the additively separable form:
where the functions $h_m, h_e, h_w$ are left unrestricted and $\psi_\omega = (\gamma_\omega, \delta_\omega, \zeta_\omega)$ are the productivity loading coefficients. \footnote{The Cobb--Douglas first-order condition (Appendix (ref)) structurally constrains the demand function to be linear in $(k, l, \omega)$, but imposes no restriction on the functional form of the dependence on $z$. The state variables $z_{jt}$ enter through input prices $\ln P_{h,jt}$, the common market factor $\ln(P_{jt}/\mu_{jt})$, markdowns $\ln \psi_{h,jt}$, and wedges $\ln \Upsilon_{h,jt}$ (equation (ref)), each of which may depend nonlinearly on $z$.} The demand slope parameters $\theta_g = (\gamma_k, \gamma_l, \delta_k, \delta_l, \zeta_k, \zeta_l)$ and the productivity loadings $\psi_\omega$ are estimated jointly by GMM together with the $3\,d_z$ coefficients of $h_m, h_e, h_w$ on the polynomial basis in $z$.
\paragraph{Homothetic structure of $\mathbb{E}[\omega \mid k, l]$.} Under Assumption (ref) (Homothetic Weak Separability), the conditional expectation of productivity admits the representation $\mathbb{E}[\omega_{jt} \mid k_{jt}, l_{jt}] = h(v(k_{jt}, l_{jt}))$. The economic motivation is discussed in Section (ref). I parametrize the index function using a CES aggregator:
which, in levels, corresponds to the CES aggregator $V = \bigl(\alpha\, K^{\rho_v} + (1 - \alpha)\, L^{\rho_v}\bigr)^{1/\rho_v}$. This nests the Cobb--Douglas case ($\rho_v \to 0$, where $v \to \alpha\,k + (1-\alpha)\,l$) as a special case and satisfies the degree-one homogeneity requirement (Assumption (ref)(B)) and the strict convexity of isoquants (Assumption (ref)(C)) for $\alpha \in (0,1)$ and any $\rho_v \neq 0$. The transformation function $h$ is approximated by a cubic polynomial:
where the constant $\rho_0$ is absorbed by de-meaning prior to estimation. Under the normalization $\mathbb{E}[\omega] = 0$, the constant satisfies $\rho_0 = -\mathbb{E}[\rho_1 v + \rho_2 v^2 + \rho_3 v^3]$; this constant is not separately identified from the production function intercept and is recovered post-estimation. Condition (A) of Assumption (ref) ($h'$ non-constant) requires $\rho_2 \neq 0$ or $\rho_3 \neq 0$; this is a necessary condition for the identification of $\beta_k$ and $\beta_l$ (Theorem (ref)). I report results for polynomial orders 3 through 5 as a robustness check; computational details including the parametrization of $h$ are in Appendix (ref).
\paragraph{Parameter classification.} The full parameter vector is $\Theta = (\theta_1', \theta_2')'$, where:
\paragraph{Residuals.} Define the observable residuals, where the nuisance functions $h_m(z), h_e(z), h_w(z)$ are estimated jointly as described below:
The equalities following the definition signs hold at the true parameter values. The nuisance functions $h_m(z), h_e(z), h_w(z)$ are approximated by second-degree polynomials in $z$ and estimated jointly with the structural parameters; details are in Appendix (ref).
Under Block A+B estimation, the normalization $\beta_k = \beta_l = 0$ is adopted; this is without loss of generality because the $\Delta(k,l)$ observational equivalence (Theorem (ref)) implies that $\beta_k$ and $\beta_l$ are not separately identified from the demand function slopes on $(k, l)$ without Block C. Under this normalization, $\tilde{y}_{jt} = \omega_{jt} + \varepsilon_{jt}$.
\paragraph{Block A: Proxy Moments.} \footnote{The moment conditions require Assumption (ref) (Appendix (ref)), which is implied by the zero conditional mean condition together with Assumption (ref).} By eliminating $\omega_{jt}$ across pairs of residuals, I construct three error terms that depend only on the structural shocks:
An asymmetric instrument strategy assigns different instruments to each error based on the shock composition. Since $u_{i,jt}$ excludes certain shocks, the corresponding intermediate inputs serve as valid additional instruments (Appendix (ref)):
where $Z_{\mathrm{base},jt} = (k_{jt},\, l_{jt},\, z_{jt})$. Block A is invariant to the $\Delta(k,l)$ transformation of Theorem (ref) and therefore cannot separately identify $\beta_k$ from the demand slopes on $(k, l)$ (Appendix (ref)).
\paragraph{Block B: Covariance Moments.} Let $\phi_h$ denote the productivity loading of residual $\tilde{h}$: $\phi_m \equiv \gamma_\omega$, $\phi_e \equiv \delta_\omega$, $\phi_w \equiv \zeta_\omega$, and $\phi_y \equiv 1$. The mutual exogeneity of shocks (Assumption (ref)(3)) implies that $\mathrm{Cov}(\tilde{h}_1, \tilde{h}_2) = \phi_{h_1}\,\phi_{h_2}\,\mathrm{Var}(\omega)$ for each pair $(h_1, h_2) \in \{y, m, e, w\}$, $h_1 \neq h_2$. Eliminating $\mathrm{Var}(\omega)$ across the six distinct pairs yields six covariance relations of the form
for each pair (Appendix (ref) lists the individual conditions). Of these six relations, four are algebraically implied by the Block A instrumental variable moments: the conditions involving cross-products of the demand residuals $\tilde{e}$ and $\tilde{w}$ with the proxy equation errors are already encoded in the Block A moment conditions through the instruments $Z_3 = (k, l, \tilde{e}, \tilde{w})$. Consequently, Block B contributes only two independent moment conditions beyond Block A, and the combined Block A+B system is just-identified. The concentrated covariance-ratio formulas derived in Appendix (ref) remain useful for obtaining closed-form scale parameter estimates, improving computational efficiency. As with Block A, Block B is invariant to the $\Delta(k,l)$ transformation.
\paragraph{Block C: Curvature Moments.} Block C resolves the $\Delta(k,l)$ indeterminacy by implementing the homothetic regularity condition (Assumption (ref), Theorem (ref)).
Define the net output residual $\tilde{y}_{jt}(\theta_1) = y_{jt} - \beta_m m_{jt} - \beta_e e_{jt} - \beta_w w_{jt}$. Evaluating at the true parameter vector $\Theta_0$, the production function (ref) gives $\tilde{y}_{jt} = \beta_k k_{jt} + \beta_l l_{jt} + \omega_{jt} + \varepsilon_{jt}$. Taking the conditional expectation with respect to $(k_{jt}, l_{jt})$:
The first step uses $\mathbb{E}[\varepsilon_{jt} \mid k, l] = 0$, which follows from Assumption (ref) by the law of iterated expectations. The second step uses Assumption (ref). No structural decomposition of $\omega_{jt}$ is postulated; equation (ref) follows entirely from the definition of conditional expectation and the regularity condition on its functional form.
Define the structural error:
Equation (ref) implies $\mathbb{E}[u_{jt} \mid k_{jt}, l_{jt}] = 0$ at $\Theta_0$, which yields valid moment conditions with any function of $(k, l)$ as instruments. I use the polynomial instrument vector:
giving the moment conditions:
As with Block A, the constant term is excluded from $Z_{2,jt}$ and $\rho_0$ (the intercept of $h$) is recovered post-estimation from the de-meaned residuals.
\paragraph{Identification mechanism.} The structural error $u_{jt}$ depends on $\theta_2 = (\beta_k, \beta_l, \alpha, \rho_1, \rho_2, \rho_3)$. Theorem (ref) establishes that under Assumption (ref), the $\Delta(k,l) = c_k k + c_l l$ transformation is incompatible with the homothetic structure unless $(c_k, c_l) = (0,0)$. Operationally, this identification works through the higher-order instruments in $Z_{2,jt}$: the nonlinear terms $v^2$ and $v^3$ in $h$ interact with the homogeneity of $v$ in a manner that uniquely pins down $\beta_k$ and $\beta_l$.
If $\rho_2 = \rho_3 = 0$ (i.e., $h$ is linear), then $\beta_k$ and $\rho_1 \alpha$ are linearly confounded and identification fails. The significance of $\hat{\rho}_2$ and/or $\hat{\rho}_3$ therefore serves as a diagnostic for the strength of identification. I report estimates and standard errors of these parameters in both the simulation and the empirical analysis.\footnote{ In practice, even when $\rho_2$ and $\rho_3$ are nonzero, the near-collinearity between $\rho_1 v(k,l)$ and $(\beta_k k, \beta_l l)$ can impede numerical optimization. I orthogonalize the polynomial basis $(v, v^2, v^3)$ against the linear span of $(1, k, l)$ before constructing $h$, so that only the nonlinear component of $h(v)$ (the source of identification, Theorem (ref)) enters the Block C moment conditions. This is a reparametrization: the structural parameters $(\beta_k, \beta_l, \alpha)$ are invariant, while the polynomial coefficients $(\rho_1, \rho_2, \rho_3)$ are redefined as loadings on the orthogonalized basis.}
\paragraph{De-meaning and estimation procedure.} All variables are de-meaned prior to estimation and the constant is excluded from all instrument vectors. All parameters $\Theta = (\theta_1, \theta_2)$ are estimated simultaneously by two-step GMM:
where $\bar{g}_j(\Theta) = T^{-1}\sum_{t=1}^T g_{jt}(\Theta)$ stacks all moment conditions, and $\hat{W}$ is the optimal weighting matrix estimated from a first-step identity-weighted GMM. Post-estimation intercepts and further implementation details are in Appendix (ref).
Given the estimated parameters $\hat{\Theta}$, the firm-level productivity measure is computed as:
If $\varepsilon_{jt} = 0$, this equals $\omega_{jt}$. Otherwise, $\hat{\omega}_{jt} = \omega_{jt} + \varepsilon_{jt}$; the ex-post shock acts as classical measurement error when $\hat{\omega}_{jt}$ is used in subsequent regressions (Theorem (ref)).
\paragraph{Practical treatment of the $\Delta(k, l)$ indeterminacy.} When Block C is not imposed, $\hat{\omega}_{jt}$ includes a location shift $c(k_{jt}, l_{jt})$ (Theorem (ref)). Since $c$ depends only on $(k, l)$, flexible controls in $(k, l)$ absorb this shift in regression analysis; in difference-in-differences designs with parallel $(k, l)$ trends, $c$ is automatically differenced out. When the identifying restrictions of Section (ref) are imposed, $\Delta(k, l)$ reduces to a constant absorbed by fixed effects. The proposed estimator therefore supports event studies and productivity regressions without requiring Block C: the $\Delta(k,l)$ component is controlled via polynomial $(k,l)$ regressors in all subsequent regressions (Section (ref)). The formal justification is provided by Proposition (ref) and the ATT identification result in Appendix (ref).
Under standard regularity conditions (Appendix (ref)), the GMM estimator satisfies:
Standard errors are clustered at the firm level to accommodate arbitrary within-firm serial dependence. The proof and regularity conditions are in Appendix (ref).
Computational details are in Appendix (ref).
\paragraph{Identification count.} The combined Block A+B system is just-identified: Block A contributes 10 moment conditions, and Block B contributes exactly two independent moment conditions beyond Block A (four of the six Block B covariance relations are algebraically redundant with Block A; Section (ref)), giving 12 moment conditions matching the 12 free parameters in $\theta_1$. The scale parameters $(\gamma_\omega, \delta_\omega, \zeta_\omega)$ are estimated via closed-form covariance ratios for computational efficiency.
\paragraph{Strength of identification for $\beta_k, \beta_l$.} As discussed in Section (ref), the identification of $\beta_k$ and $\beta_l$ relies on the nonlinearity of $h$ ($\rho_2 \neq 0$ or $\rho_3 \neq 0$). I report the estimates and $t$-statistics of $\hat{\rho}_2$ and $\hat{\rho}_3$ as diagnostics. If both are insignificant, the identification of primary input elasticities may be weak, and the researcher should interpret $\beta_k$ and $\beta_l$ with caution or consider exclusion restrictions (Corollary (ref)) as an alternative identification strategy.
\paragraph{Reduced-form check of Assumption (ref).} As a pre-estimation diagnostic, one may estimate $\theta_1$ from Blocks A and B alone (which does not require Assumption (ref)), construct $\tilde{y}_{jt}(\hat{\theta}_1)$, and examine whether $\mathbb{E}[\tilde{y} \mid k, l]$ exhibits a homothetic structure via nonparametric regression. A visual departure from homotheticity would indicate a violation of the identifying assumption.
\paragraph{Polynomial degree selection.} The cubic specification of $h$ can be extended to higher-order polynomials. I recommend reporting results for polynomial orders 3 through 5 and selecting via information criteria.
Appendix (ref) reports the full-sample Block C recovery results for $(\beta_k, \beta_l)$ across all 502 industries, comparing the homothetic approach with the exclusion restriction and ACF estimators.
The identification results in Section (ref) show that the Markov assumption is unnecessary; this section asks whether removing it matters quantitatively. I use Monte Carlo simulations to measure the bias that the Markov assumption introduces in the materials elasticity and to trace its propagation into downstream objects. The primary comparison is between the proposed estimator, which imposes no restriction on productivity dynamics, and the standard ACF estimator, which requires first-order Markov.
All DGPs share a common structure for the production function, demand functions, and dynamic input decisions, differing only in the productivity process. Detailed parameter settings are in Appendix (ref).
The firm's production function is Cobb--Douglas in all inputs:\footnote{The Cobb--Douglas specification is standard in Monte Carlo studies of production function estimators ackerberg2015identification,gandhi2020onthe. Evaluating the proposed method under more flexible production functions (e.g., translog) is left for future work; the identification results (Theorems (ref)--(ref)) do not require Cobb--Douglas.}
with true parameter values $(\beta_0, \beta_k, \beta_l, \beta_m, \beta_e, \beta_w) = (0.1, 0.2, 0.3, 0.3, 0.15, 0.1)$ and $\varepsilon_{jt} \sim \text{i.i.d.}\ N(0, 0.05^2)$.
Intermediate input demands are log-linear in $(k, l, \omega)$ with input-specific demand shocks following independent AR(1) processes ($\rho = 0.5$, $\sigma = 0.15$). The demand function coefficients are calibrated from the first-order conditions of cost minimization under input-specific markdowns (Appendix (ref)); the productivity loading coefficients $(\gamma_\omega, \delta_\omega, \zeta_\omega) = (2.2, 2.0, 1.8)$ differ across inputs, reflecting heterogeneous markdowns.\footnote{Under perfect competition with a Cobb--Douglas production function, the first-order condition implies $\gamma_\omega = 1/(1 - \beta_m) \approx 1.43$; the larger values incorporate input-specific markdowns and procurement frictions (see Appendix (ref)).} Conditional independence (Assumption (ref)) is a cross-sectional condition requiring mutual independence across inputs at each point in time; it is unaffected by the serial correlation of individual shocks, since each AR(1) has mutually independent innovations.
Primary inputs are endogenously determined. Capital accumulates through dynamic investment, and labor is chosen based on forecasted productivity from an AR(1) model. The labor demand function is structured so that Assumption (ref) holds: the conditional expectation $\mathbb{E}[\omega_{jt} \mid k_{jt}, l_{jt}]$ is a function of a CES aggregator with $(\alpha, \rho_v) = (0.4, 0.3)$. Full parameter details are provided in Appendix (ref).
To test the robustness of the proposed method, I generate productivity under three scenarios:
Using the generated data, I organize the estimation into two parts to isolate the contributions of each block of moment conditions.
\paragraph{Part 1: Flexible input parameters.} I estimate the intermediate input elasticities $(\beta_m, \beta_e, \beta_w)$ and compare four estimators.\footnote{The main text figures report two of the four estimators (ACF and Proposed). ACF-Mod results are in Appendix (ref); GNR results are also reported there.} The four estimators are:
\paragraph{Part 2: Fixed input parameters.} I additionally estimate $(\beta_k, \beta_l)$ by adding the homothetic regularity condition (Block C) to the proposed estimator:
I report bias, $\text{Bias}(\hat{\beta})=\mathbb{E}_{R}[\hat{\beta}^{(r)}]-\beta_{\text{true}}$, and RMSE, $\text{RMSE}(\hat{\beta})=\sqrt{\mathbb{E}_{R}[(\hat{\beta}^{(r)}-\beta_{\text{true}})^{2}]}$, averaged over $R$ Monte Carlo repetitions.
For Part 1, I run $R = 100$ replications for each combination of DGP and estimation method, varying the number of firms $N \in \{50, 200, 500\}$ and the observation period $T \in \{10, 20, 50\}$ to examine the impact of sample size. For Part 2, I run $R = 100$ replications at $(N, T) = (200, 50)$. The parameter estimates obtained in each repetition are collected, and mean bias and RMSE are calculated for comparison. With $R = 100$, the simulation standard error of the estimated bias is approximately $\text{SD}/\sqrt{R}$; for the typical standard deviation of $\hat{\beta}_m$ ($\approx 0.005$), this yields a simulation uncertainty of $\approx 0.0005$, which is small relative to the reported biases.
I report the Part 1 results in Figures (ref) and (ref) and the Part 2 results in Figure (ref).\footnote{Additional summary tables, including GNR results, are provided in Appendix (ref). RMSE convergence plots are in Figure (ref).} These results confirm that the proposed estimator performs well under all DGPs considered and illustrate the sensitivity of the ACF framework to violations of the Markov assumption.
\paragraph{Remark on GNR.} GNR shares the static identification strategy of the proposed method (both recover $\beta_m$ from within-period variation without a Markov assumption) but requires competitive input markets with non-persistent demand shocks (their Appendix O6, Assumption 7). The present DGP, which features persistent input-specific shocks ($\rho_\tau = 0.5$), is therefore outside GNR's maintained assumptions by design: the DGP is calibrated to the proposed method's setting, not GNR's. Under GNR's own assumptions ($\tau = \nu = \eta = 0$), the share regression recovers $\beta_m$ consistently regardless of the productivity process. The simulation results for GNR (Appendix (ref)) should accordingly be read as illustrating the sensitivity of the FOC-based approach to input market imperfections, not as a general performance comparison.
\paragraph{DGP1 (AR(1) Baseline):}
Under DGP1, where the Markov assumption holds, all three estimators (ACF, ACF-Mod, and Proposed) are consistent. The bias for each method decays toward zero as $T$ increases (Figure (ref)). The boxplots in Figure (ref) corroborate this finding; the proposed method remains centered on the true values. The ACF and ACF-Mod estimators show small positive finite-sample bias that diminishes with sample size (see Appendix (ref) for detailed tables). However, the proposed estimator exhibits larger variance than the ACF estimator under DGP1, resulting in higher RMSE when the Markov assumption is correctly specified (Appendix Table (ref)). This is the efficiency cost of the static approach: the proposed method trades time-series information for robustness to dynamic misspecification. Under DGP2 and DGP3, this ranking reverses: ACF's bias dominates its variance advantage, yielding larger mean squared error. The static identification strategy is also the only approach in this literature that permits event study and difference-in-differences designs, where the treatment itself violates the Markov assumption (Section (ref)).
\paragraph{DGP2 (AR(2)) and DGP3 (Potential Outcome):}
Under DGP2 and DGP3, where the first-order Markov assumption does not hold, Figure (ref) reveals a clear divergence. The ACF estimator exhibits positive bias in $\hat{\beta}_m$ that does not vanish with increasing $T$. Under DGP2, this reflects standard omitted-variable inconsistency: the AR(2) component of productivity persistence is not captured by the first-order transition equation. Under DGP3, the issue is more fundamental: the Markov transition equation is structurally incompatible with the potential outcome process chen2024identifying, so the ACF moment condition lacks a structural interpretation and the resulting estimate does not converge to the true $\beta_m$. An infeasible oracle benchmark (ACF-Mod) that removes scalar unobservability by treating demand shocks as observed shows comparable bias under both DGPs, confirming that the source is Markov misspecification rather than demand shock heterogeneity (Appendix (ref)).
The proposed method, by contrast, exhibits negligible bias across these specifications. The bias remains close to zero for all values of $T$ under both DGP2 and DGP3. Because the estimator relies solely on static conditional independence, it remains invariant to the underlying productivity dynamics. The main text figures report results for $N = 500$; increasing $N$ reduces variance for all estimators but does not mitigate ACF's asymptotic bias under DGP2 or DGP3 (Appendix Figure (ref)), confirming that the bias is asymptotic rather than finite-sample.
\paragraph{Block A+B vs.\ Block A+B+C (Part 2):}
Part 2 supplements Part 1 by adding Block C to recover $(\beta_k, \beta_l)$. I use the design point $(N, T) = (200, 50)$, which matches the Part 1 baseline, to examine whether Block C disturbs the Block A+B parameters. Figure (ref) presents the results, where Block C is added to identify $(\beta_k, \beta_l)$. In the baseline DGP1, the proposed method recovers both parameters with negligible bias. Under DGP3, where ACF estimates of $\beta_k$ and $\beta_l$ collapse toward zero (RMSE $\approx 0.20$--$0.30$), the proposed method achieves substantially lower error (RMSE $\approx 0.02$). The intermediate input elasticities $(\beta_m, \beta_e, \beta_w)$ remain stable between Part 1 and Part 2, confirming that the addition of Block C moments does not contaminate the well-identified flexible input parameters. This stability shows in finite samples that the joint GMM system does not transmit Block C misspecification into the flexible input estimates: the intermediate input elasticities are identified by Blocks A and B alone (Theorem (ref), specialized to the Cobb--Douglas parametric model of Section (ref)), so any misspecification in Block C affects only $(\beta_k, \beta_l)$. Because markups depend solely on $\beta_m$ (equation (ref)), the primary empirical application is insulated from Block C specification.
\paragraph{DGP4 (Conditional Independence Violation):}
DGP4 examines the cost of violating Assumption (ref) by introducing correlation between the electricity demand shock $\nu_{jt}$ and the water demand shock $\eta_{jt}$, arguably the most economically salient threat to conditional independence, since both are utility services subject to common energy prices and infrastructure constraints. The materials shock $\tau_{jt}$ remains independent. The correlation $\rho_{ew} \equiv \mathrm{Corr}(\nu_{jt}, \eta_{jt})$ varies from 0 to 0.30.
The bias mechanism operates through the scale parameter $\zeta_\omega$. Positive $\mathrm{Cov}(\nu, \eta)$ inflates $\mathrm{Cov}(\tilde{e}, \tilde{w})$, causing the concentrated scale estimator $\hat{\zeta}_\omega = \mathrm{Cov}(\tilde{e}, \tilde{w}) / \mathrm{Cov}(\tilde{y}, \tilde{e})$ to overestimate $\zeta_\omega$ (Appendix (ref)). The overestimated $\hat{\zeta}_\omega$ introduces a positive productivity component into the Block A residual $u_2 = \hat{\zeta}_\omega \tilde{m} - \hat{\gamma}_\omega \tilde{w}$, which the GMM compensates by increasing $\hat{\beta}_m$, yielding an upward bias.
Table (ref) (Appendix Figure (ref)) reports the results. When $\rho_{ew} = 0$, the proposed method is approximately unbiased. As $\rho_{ew}$ increases, $\hat{\beta}_m$ exhibits increasing upward bias. The magnitudes suggest that the estimator is robust to moderate violations. The bias direction is the same as the Markov misspecification bias documented in DGPs 2 and 3 for ACF: both push $\hat{\beta}_m$ upward. Therefore, the empirical finding that the proposed estimator yields lower $\hat{\beta}_m$ than ACF (Section (ref)) cannot be attributed to CI violation; it must reflect Markov misspecification bias in ACF.\footnote{ACF uses only the materials demand proxy and does not exploit cross-shock variation, so it is unaffected by $\mathrm{Corr}(\nu, \eta)$.}
Table (ref) summarizes the bias properties across DGPs 1--3. The proposed method is unbiased across all three specifications, while ACF exhibits positive bias under Markov misspecification.
Taken together, the Monte Carlo simulations confirm that the proposed method recovers production function parameters without imposing restrictions on the productivity process. In contrast, standard methods exhibit substantial positive bias in $\hat{\beta}_m$ when the assumed law of motion for productivity does not match the true data generating process. The simulations establish that Markov misspecification generates a detectable and economically meaningful bias. The empirical application then examines whether these patterns hold in Japanese manufacturing data.
The empirical analysis has two objectives: to test whether the conditional independence framework produces economically plausible estimates across the manufacturing sector, and to assess the relative plausibility of the static and dynamic identifying assumptions through the convergence diagnostic of Remark (ref). I estimate the production function for all 502 manufacturing industries using Block A+B, reporting analytical standard errors. A practical consequence of this block structure: the markup estimates, productivity determinants, and convergence diagnostics reported below require only Blocks A and B. These results do not depend on the resolution of the $\Delta(k,l)$ indeterminacy and are available for all 502 industries. Block A+B+C is used for a subset of industries where $(\beta_k, \beta_l)$ recovery is needed for productivity level analysis.
The section is organized as follows. Sections (ref) and (ref) describe the data and estimation specifications. Section (ref) presents the exclusion restriction diagnostic. Section (ref) reports production function parameters and markup estimates. Section (ref) examines productivity determinants. Section (ref) presents an event study application exploiting the non-Markov validity of the estimator. Section (ref) reports estimates of $(\beta_k, \beta_l)$ from two independent identification routes.
I apply the proposed method to the Japanese Census of Manufactures and the Economic Census for Business Activity. I estimate the production function for all manufacturing industries with at least 50 firm-year observations in the extended panel (2003--2020), yielding Block A+B estimates for 502 industries covering 559{,}381 firm-year observations. \footnote{The identification results of Section (ref) require only the joint distribution of $(m_{jt}, e_{jt}, w_{jt}, k_{jt}, l_{jt})$ at a single point in time; no assumption on the time-series dynamics of $\omega_{jt}$ is needed (Appendix (ref)). The panel dimension is exploited solely to improve estimation efficiency by time-averaging the sample moment conditions, $\bar{g}_j(\Theta) = T^{-1}\sum_{t} g_{jt}(\Theta)$, which reduces finite-sample variance without affecting consistency.} These estimates provide markup distributions and productivity determinants at the level of the entire manufacturing sector. Four industries (food processing [Bread, industry code 971], paper products (Corrugated board boxes, code 1453), chemicals (Plastic film, code 1821), and machinery [Industrial robots, code 2694]) serve as representative cases for the time-varying parameter analysis in Appendix (ref), covering major manufacturing sectors (food, paper, chemicals, machinery). Analytical standard errors from the GMM sandwich formula are reported for both the proposed method and the ACF benchmark.
The core variables include the logarithm of real output, $y_{jt}$, the logarithm of real capital stock, $k_{jt}$, and the logarithm of labor input, $l_{jt}$.
I map the theoretical input triplet to observable data as follows. I designate the real value of primary raw materials as $m_{jt}$, the quantity of electricity as $e_{jt}$, and the quantity of industrial water as $w_{jt}$. This selection exploits the fact that industrial water and electricity prices are typically regulated, limiting firm-specific bargaining. This institutional feature reduces the risk of unobserved common price shocks inducing correlation between $\nu_{jt}$ and $\eta_{jt}$, thereby supporting the validity of the conditional independence assumption ($\tau_{jt} \perp \nu_{jt} \perp \eta_{jt} \mid (\omega_{jt}, x_{jt})$). The principal remaining threat is commodity price shocks that jointly affect raw materials costs and electricity generation costs. Two features mitigate this concern: (i) industrial electricity prices exhibit less high-frequency variation than raw materials procurement costs, as the fuel cost adjustment mechanism smooths commodity price pass-through on a quarterly basis; and (ii) even if a residual common utility shock induces positive $\mathrm{Corr}(\nu_{jt}, \eta_{jt})$, the resulting bias in $\hat{\beta}_m$ is upward (the same direction as ACF's Markov bias), so the empirical gap between methods cannot be attributed to CI violation (Section (ref), Appendix (ref)).
I augment $x_{jt}$ with control variables $z_{jt}$ consisting of beginning-of-period total inventory ($z_{1,jt}$), its square ($z_{2,jt} \equiv z_{1,jt}^2$), plant fixed effects, and year fixed effects. These controls directly implement the conditioning strategy of Section (ref), where common shocks are absorbed by $z_{jt}$ so that the residual shock terms $\tau_{jt}, \nu_{jt}, \eta_{jt}$ satisfy the conditional independence assumption. Inventory proxies for unobserved product demand fluctuations kumar2019productivity: a firm anticipating high demand accumulates more stock in advance, so inventory captures the common demand component that would otherwise enter all three input demands simultaneously (Section (ref)). The quadratic term $z_{2,jt}$ accommodates a nonlinear relationship between inventory and unobserved demand, consistent with the structural decomposition in equation (ref) where demand-related terms enter input prices nonlinearly. Year fixed effects absorb common input price shocks (e.g., energy price movements) that affect all inputs simultaneously, as discussed in Section (ref). Plant fixed effects absorb time-invariant plant-level heterogeneity in input prices and buyer-supplier relationships, capturing the firm-attribute component of input market power noted in Section (ref). The use of fixed effects exploits the panel dimension for efficiency and enriches the conditioning set for the conditional independence assumption, but does not impose any restriction on the time-series dynamics of $\omega_{jt}$. Standard proxy variable estimators use the Markov transition equation to address exit-driven selection olley1996thedynamics. The proposed method does not require this correction: because identification conditions on $\omega_{jt}$, endogenous exit based on $(\omega_{jt}, k_{jt})$ is absorbed by the conditioning and the moment conditions hold on the surviving population without a survival probability correction (Remark (ref)). Plant fixed effects further reduce the influence of systematic level differences across plants.
I contrast the results of my approach with those obtained from the standard ACF framework.
First, I implement the proposed method using the GMM estimator derived in Section (ref). The estimator jointly recovers the production function and demand parameters as described in Section (ref). The CES aggregator parameters $(\rho_v, \alpha)$ are selected via profile GMM: for a grid of $(\rho_v, \alpha)$ values, the remaining parameters are estimated by minimizing the GMM objective, and the pair yielding the smallest $J$-statistic is selected.\footnote{Under strong identification of $(\rho_v, \alpha)$, the profile GMM procedure yields a $J$-statistic with the standard $\chi^2$ distribution asymptotically newey1994chapter. When identification of these parameters is weak, the minimum-$J$ selection may bias the test toward under-rejection, making the test conservative. The block bootstrap standard errors reported below account for the uncertainty in $(\rho_v, \alpha)$ selection by re-running the profile grid search within each bootstrap replication.} The nuisance functions $h_m, h_e, h_w$ are approximated by second-degree polynomials in $(z_{1,jt}, z_{2,jt})$, giving a polynomial basis of dimension $d_z = 2$ and thus $\dim\Theta = 24$ where Block A+B is just-identified (Section (ref)). I report the estimates of $\hat{\rho}_2$ and $\hat{\rho}_3$ as diagnostics for the strength of identification of $\beta_k$ and $\beta_l$ (Section (ref)).
As a benchmark, I estimate the ACF two-step GMM with the same control variables $z_{jt}$ to ensure comparability. Analytical standard errors from the GMM sandwich formula are reported for both methods.
\paragraph{Identifying assumptions in practice.} The ACF framework requires scalar unobservability (productivity as the sole unobservable in input demand) and a first-order Markov process for productivity. GNR requires scalar unobservability and competitive input markets. The proposed method requires conditional independence of input-specific demand shocks. Scalar unobservability rules out procurement relationships, supply contracts, and input-specific markdowns; the proposed method permits these. The GNR competitive input market assumption precludes markup estimation, since the identifying condition coincides with the object of interest. The conditional part of the independence assumption depends on the adequacy of the control variables $z_{jt}$, but this dependence is shared by the ACF proxy equation.
Two diagnostics probe different layers of the identification strategy before any structural results are interpreted:
Table (ref) summarizes these two diagnostics and their empirical outcomes.
\paragraph{Exclusion restriction diagnostic.} Figure (ref) applies the exclusion-based OLS recovery of Proposition (ref) to all 502 manufacturing industries, plotting $\hat{\beta}_k^{(m)}$ against $\hat{\beta}_k^{(e)}$ (panel a) and $\hat{\beta}_l^{(m)}$ against $\hat{\beta}_l^{(e)}$ (panel b). Under the exclusion restriction, both panels should cluster along the 45-degree line. Panel (a) confirms this for capital: points concentrate tightly around the diagonal, consistent with $a_{k}^{h} = 0$ across industries. Panel (b) reveals the opposite for labor: points scatter widely, indicating that different proxy equations yield systematically different $\hat{\beta}_l$ values.
The asymmetry between capital and labor is the central diagnostic finding. Capital is quasi-fixed within the production period and does not directly influence short-run intermediate input procurement, so $a_{k}^{h} = 0$ is economically plausible. The systematic failure for labor is consistent with labor affecting production scheduling, shift patterns, and input utilization through input-specific channels (Appendix (ref)). The formal Wald test of $d_k = d_l = 0$ is rejected for 37% of industries at the 5% level, while the labor-only Wald test ($d_l = 0$) is rejected for 28% of industries, confirming that the labor exclusion restriction is violated for a substantial share of the sample while capital passes in most cases.
The intermediate input elasticities $(\beta_m, \beta_e, \beta_w)$ are identified by Blocks A and B alone (Theorem (ref), specialized to the Cobb--Douglas parametric model of Section (ref)), without requiring Block C or Assumption (ref). The ACF estimates of $\hat{\beta}_m$ are systematically higher than those from the proposed method (Table (ref) in Appendix (ref)), consistent with the Markov misspecification bias documented in the Monte Carlo simulations (Section (ref)). The cross-industry distribution of all Block A+B and Block C parameter estimates is reported in Table (ref) in Appendix (ref).
Electricity and water elasticities are small across industries (median $\hat{\beta}_e = 0.001$ and $\hat{\beta}_w = 0.006$, respectively), consistent with these inputs serving auxiliary rather than central production roles in Japanese manufacturing. Their demand shocks nevertheless remain valid exclusion restrictions for identifying $\beta_m$ in the proposed GMM.
Under perfect competition, $\beta_m$ equals the revenue share, which is the basis of GNR's share regression. My estimator identifies $\beta_m$ independently of the first-order condition, permitting imperfect competition in both product and input markets.
\paragraph{Markups.} Markups are computed following deloecker2012markups. Under the Cobb-Douglas specification maintained throughout, the markup formula simplifies to
where $s_{m,jt}$ denotes the expenditure share of raw materials in total revenue. Unlike the standard production approach, in which Hicks-neutral productivity and scalar unobservability jointly imply $\hat{\beta}_h/s_{h,jt} = \mu_{jt}$ for every variable input $h$ (so that materials, labor, and energy serve as interchangeable markup proxies), this paper allows input-specific markdowns $\psi_{h,jt}$ for each static input $h \in \{m,e,w\}$, captured by the demand shocks $(\tau_{jt},\nu_{jt},\eta_{jt})$ (Appendix (ref)). Consequently, $\hat{\beta}_h/s_{h,jt}$ will generally differ across inputs by design; this divergence reflects the richer structure of the framework, not an overidentification failure. Raw materials are selected for markup computation because competitive commodity markets support the absence of buyer-side market power ($\psi_{m,jt}\approx 1$; avignon2025markups), giving $\hat{\beta}_m/s_{m,jt}\approx\mu_{jt}$; this is a maintained assumption. \footnote{The empirical specification imposes Hicks-neutral Cobb-Douglas production; if factor-augmenting productivities differ across inputs, $\hat{\beta}_m$ may absorb non-neutral components and bias the markup estimate raval2023testing. The identification theory accommodates non-Hicks-neutral production (Appendix (ref)), but the implemented GMM does not exploit this generality.} These estimates require only Blocks A and B and are invariant to the $\Delta(k,l)$ indeterminacy (Theorem (ref)), since equation (ref) depends only on $\hat{\beta}_m$ and the observable cost share. I restrict the comparison to industries with at least 50 firms ($N_{\text{firms}} \geq 50$), which removes industries where the lower bound $\hat{\beta}_m \approx 0$ reflects identification failure rather than true low input elasticities. \footnote{The value-added markup $\mu^{VA} = \beta_l^{VA}/s_l^{VA}$ can differ substantially from the gross output markup when the materials share is large. gandhihowheterogeneous document that gross output and value-added specifications yield fundamentally different productivity estimates. I report gross output markups throughout.}
\paragraph{Comparison with ACF.} Figure (ref) plots the empirical CDF of industry-level median markups under the proposed method and the ACF benchmark for the $N_{\text{firms}} \geq 50$ subsample. The two distributions are stochastically ordered: the ACF CDF lies strictly to the right of the proposed CDF at every percentile (Table (ref)). The proposed method yields a median markup of 0.926, while ACF yields 1.027, a gap of 0.101 at the median. At the 90th percentile the gap widens to approximately 0.15. Under the proposed method, 37% of industries show markups above unity, compared with 54% under ACF.
The Monte Carlo evidence in Section (ref) provides a structural interpretation. Under DGP 3 (potential-outcome dynamics, Table (ref)), ACF incurs a bias of $+0.19$ in $\hat{\beta}_m$ (true value $0.30$), a 63% relative overestimate, while the proposed estimator is essentially unbiased ($\text{bias} = 0.001$). The empirical gap of $+0.10$ at the median corresponds to a relative overestimate of roughly 11% in $\hat{\beta}_m$, well within the range predicted by the DGP 3 calibration. The evidence is therefore consistent with the theoretical prediction that ACF overestimates $\hat{\beta}_m$ when productivity dynamics deviate from the Markov assumption. Because markups recovered from production functions are widely used to assess the evolution of market power deloecker2020rise, the systematic gap documented here raises the question of whether existing markup estimates are sensitive to the choice of identifying assumption.
The following analysis requires only Blocks A and B. Because the $\Delta(k,l)$ indeterminacy (Theorem (ref)) varies only through $(k_{jt}, l_{jt})$, it is absorbed by the cubic polynomial controls in $(k_{jt}, l_{jt})$ included in the regression. The same argument applies to proportional common shocks (Section (ref)): if an unobserved common component $\xi_{jt}$ loads proportionally on all intermediate input demands, it is absorbed into the recovered productivity $\hat{\omega}_{jt}$, and its $(k_{jt}, l_{jt})$-dependent component is absorbed by firm fixed effects. The determinants regression therefore identifies the association between covariates and the total latent efficiency measure that drives input allocation decisions, regardless of whether this measure coincides with physical productivity.
As a validation of the recovered productivity measures, I examine their association with observable economic fundamentals. I regress the productivity residual jointly on three firm-level covariates (log investment, exporter status, and log wages) with firm and year fixed effects: \[ \hat{\omega}_{jt} = \phi_j + \gamma_t + \mathbf{x}_{jt}'\boldsymbol{\beta} + u_{jt}, \] clustering standard errors at the firm level. For the proposed method, I additionally include a cubic polynomial in $(k_{jt}, l_{jt})$ as nonparametric controls, since the $\Delta(k,l)$ indeterminacy enters through capital and labor. The ACF regression omits these controls, as the ACF residual already subtracts $\hat{\beta}_k k + \hat{\beta}_l l$. Log wages is included as a correlate of productivity; a maintained caveat is that wages may be endogenous, as high-productivity firms can share rents with workers, so the coefficient captures association rather than a causal effect.
Table (ref) reports the results. Under the proposed method, log wages are strongly positively associated with estimated productivity ($\hat{\beta} \approx 0.139$, $p < 0.01$), and log investment is positive but small ($\hat{\beta} \approx 0.001$, $p < 0.01$), while exporter status is negligible and imprecisely estimated. The ACF regression yields a smaller wage coefficient ($\hat{\beta} \approx 0.109$). The mechanism is as follows: under ACF, the upward bias in $\hat{\beta}_m$ propagates into $\hat{\omega}^{\text{ACF}} = y - \hat{\beta}_m^{\text{ACF}} m - \hat{\beta}_k k - \hat{\beta}_l l$, subtracting too large a materials component and systematically depressing the recovered productivity level for materials-intensive firms. This distortion attenuates the association between $\hat{\omega}$ and economic fundamentals that covary with input intensity. The magnitude of the improvement depends on the relative variance of demand shocks and productivity; whether the pattern generalizes beyond this application requires further investigation.
Because the proposed estimator recovers productivity from static covariances alone, its estimates are valid under any productivity dynamics, Markov or otherwise (Proposition (ref)). Standard proxy variable estimators embed a Markov transition equation that is structurally incompatible with the potential outcomes framework when a treatment alters the transition path of productivity chen2024identifying; the moment condition that identifies the production function parameters has no structural interpretation under treatment, so the resulting estimates lack economic meaning for policy evaluation. Neither problem arises here, since estimation does not employ a transition equation.
As an illustration, I examine the 2011 T\={o}hoku earthquake using a difference-in-differences design. The treatment group consists of plants in the three core prefectures directly struck by the earthquake and tsunami (Iwate, Miyagi, and Fukushima; seismic intensity $\geq$ 6-strong), where physical destruction and the nuclear disaster caused severe and sustained disruption to production. The control group consists of plants in Kinki and western prefectures (prefectures 25--47). Supply chain contamination of the control group is mitigated by the industry$\times$year fixed effects, which absorb any industry-level aggregate shocks that propagate nationally. Pre-treatment coefficients are flat (max$|\hat{\delta}_t| < 0.013$ for the proposed method, $< 0.020$ for ACF); the full event-study figure is in Appendix (ref).
Table (ref) reports the difference-in-differences estimates under both methods. For the proposed method, cubic polynomial controls in $(k,l)$ are included to absorb the $\Delta(k,l)$ indeterminacy in the residual $\hat{\omega}$; the ACF method requires no such controls, as $\hat{\omega}^{\mathrm{ACF}}$ already subtracts $\hat{\beta}_k k + \hat{\beta}_l l$. Both methods detect a negative and statistically significant post-treatment effect on productivity. Under the proposed method, the DiD estimate is $-1.28$ percent (s.e.\ $0.52$, $p < 0.05$); under ACF, the corresponding estimate is $-1.68$ percent (s.e.\ $0.41$, $p < 0.01$). The gap between the two estimates is approximately $0.40$ percentage points. To illustrate the potential economic magnitude: if a bias of this order applied to the aggregate manufacturing sector, it would correspond to roughly \$3.6 billion (\textyen400 billion) per year, given Japan's manufacturing value added of approximately \$0.9 trillion (\textyen100 trillion) at the 2003--2020 average exchange rate (National Accounts, Cabinet Office of Japan).\footnote{This back-of-envelope calculation extrapolates the local DiD gap to the national level under the assumption that the Markov misspecification bias is of comparable magnitude across industries. The cross-industry markup comparison (Table (ref)) shows that ACF yields systematically higher $\hat{\beta}_m$ at every percentile, consistent with the assumption, but the magnitude varies by industry. The figure should be interpreted as indicative of the scale at stake, not as a structural estimate of aggregate mismeasurement.} Proposition (ref) guarantees that the proposed estimates recover $\mathbb{E}[\omega_{jt}\mid D_{jt}]$ under Conditions (i)--(ii) of that proposition. Condition (ii) is satisfied by construction: the earthquake is a natural disaster whose occurrence is orthogonal to firm-level input demand shocks $(\tau,\nu,\eta)$. Condition (i) requires that the earthquake does not alter the functional form of the demand functions $g_m, g_e, g_w$. Difference-in-differences estimates of the post-treatment change in intermediate input shares show no significant shift in the materials share ($t = 0.87$) or water share ($t = 1.14$). The electricity share shows a small post-treatment increase ($t = 8.91$, $\Delta s_e \approx 0.002$); this is a mechanical compositional effect of the simultaneous contraction in materials usage ($t = -4.53$), which raises the electricity expenditure share $s_e$ without altering the structural demand function $g_e$.\footnote{The level of electricity consumption does not show a significant post-treatment increase when measured in physical units (kWh) rather than expenditure shares, supporting the compositional interpretation.} The ACF estimator does not carry this guarantee. Its residual subtracts $\hat{\beta}_k k + \hat{\beta}_l l$, so
The bias terms vanish only if $\hat{\beta}^{\mathrm{ACF}} = \beta^{\mathrm{true}}$ (exact identification) or if treatment is orthogonal to $(k,l)$. Monte Carlo evidence (Section (ref)) shows that ACF incurs positive bias in $\hat{\beta}_m$ under Markov misspecification; the condition of orthogonality also fails here (DiD$(l) = -0.029$, $t = -7.86$). The proposed estimates, resting on the theoretical guarantee of Proposition (ref), provide a theoretically justified point of comparison.
Identifying $(\beta_k, \beta_l)$ requires closing the $\Delta(k,l)$ indeterminacy documented in Theorem (ref). The paper provides two independent routes: the exclusion restriction (Corollary (ref)) and the homothetic regularity condition (Theorem (ref)).
The exclusion-based OLS recovery (Proposition (ref)) is applied to all 502 manufacturing industries and produces mutually consistent estimates of $\beta_k$ across the three proxy equations (materials, electricity, water), while $\beta_l$ estimates diverge systematically, confirming the diagnostic pattern in Figure (ref) that the exclusion restriction holds for capital but not labor.
To identify $(\beta_k, \beta_l)$ jointly, I apply the Block C homothetic CES approach (Section (ref)). Block C is the primary identification route for both $\beta_k$ and $\beta_l$; the exclusion restriction provides an independent check on $\beta_k$ only, since the restriction fails for labor (Figure (ref), Panel b). The two strategies yield mutually consistent estimates of $\beta_k$ for the 302 industries where the exclusion restriction is validated for capital (Table (ref), Appendix (ref)).
Table (ref) summarizes the cross-industry distributions of $\hat{\beta}_k$ and $\hat{\beta}_l$ across three approaches: exclusion restriction (broken out by proxy input), Block C (homothetic CES), and ACF.
Estimates of $\hat{\beta}_k$ are broadly consistent across all three approaches (median $\approx 0.01$--$0.04$), corroborating the identification cross-check in Table (ref) (Appendix (ref)). Labor elasticity estimates diverge more substantially: Block C yields a median $\hat{\beta}_l = 0.33$, while ACF produces a higher median of $0.50$. The Monte Carlo simulations (Table (ref)) show that under DGP 3, ACF $\hat{\beta}_l$ collapses to near zero (bias $\approx -0.30$), while the proposed method recovers the true value accurately (bias $\approx +0.01$). The empirical ACF estimate lies above the proposed estimate, which is the opposite direction from the MC collapse. Both patterns reflect the same fragility: ACF labor elasticity identification breaks down when the Markov assumption is violated, with the direction of the deviation depending on the specific dynamics of the data-generating process. Figure (ref) shows the full cross-industry density distributions for all three methods. A four-group comparison across identification strategies is reported in Table (ref) (Appendix (ref)).
Can the production function be identified without restricting how productivity evolves over time? This paper answers in the affirmative. Replacing the Markov assumption with conditional independence across three intermediate inputs, the paper shows that the production function and the distribution of productivity are nonparametrically identified from a single cross-section. No assumption on the law of motion for $\omega_{jt}$ is required at any stage of estimation. The empirical analysis, covering 502 Japanese manufacturing industries, confirms that the choice between the two identification strategies has quantitative consequences for every downstream object: input elasticities, markups, allocative efficiency, and the measured response of productivity to economic shocks.
The consequences are economically large. The proposed method yields systematically lower markups than the standard proxy variable estimator across the entire distribution (median 0.93 vs.\ 1.03; the share of industries above unity falls from 54 to 37 percent), shifting the measured degree of market power in the manufacturing sector. In the earthquake event study (Section (ref)), the difference-in-differences estimate of the productivity effect on plants in the three most severely affected prefectures is $-1.28\%$ under the proposed method and $-1.68\%$ under the standard method; the 0.40 percentage point gap corresponds to roughly \$3.6 billion (\textyen400 billion) per year in mismeasured productivity when scaled to aggregate manufacturing output. The olley1996thedynamics decomposition and the productivity determinant regressions reinforce the same pattern: the log-wage coefficient is roughly 25% larger under the proposed method (Table (ref)), consistent with a higher signal-to-noise ratio in the recovered productivity measure once input-specific demand shocks are separated out. The Monte Carlo simulations, the convergence diagnostic, and the determinant regressions all point in the same direction, and the underlying mechanism is general: in any setting where a policy, shock, or institutional change alters the transition path of productivity, the Markov transition equation is structurally misspecified and the resulting production function parameters lack a consistent interpretation. Trade liberalization, R&D subsidies, natural disasters, and mergers all generate such dynamics. The proposed method accommodates these settings because it imposes no restriction on how productivity evolves. The Cobb--Douglas functional form is shared by both the proposed method and the ACF benchmark, so the gap between estimates reflects the difference in identifying assumptions, not in functional form.
These findings connect to two broader debates. First, the recent literature on rising global markups deloecker2020rise relies on production function estimates that impose the Markov assumption. The present results suggest that markup levels, and potentially trends, are sensitive to this assumption; replication of the global markup finding under conditional independence identification is a natural next step. Second, chen2024identifying show that standard proxy variable estimators are structurally incompatible with a potential outcomes framework: the Markov transition equation has no structural interpretation when a treatment alters the productivity process, so the resulting estimates lack economic meaning under policy evaluation. The proposed method avoids this problem because it uses no transition equation; Proposition (ref) establishes that the recovered productivity measure retains a causal interpretation under treatment assignment mechanisms satisfying conditions (i)--(ii) of that proposition (the treatment does not alter demand function structure and is orthogonal to input-specific demand shocks). The same static structure accommodates time-varying parameters without additional assumptions, since no intertemporal link is imposed.
A broader implication concerns the nature of identifying assumptions in production function estimation. The Markov restriction is a constraint on the time-series behavior of an unobservable; the conditional independence restriction is a constraint on the structure of input markets. The latter is grounded in economic primitives (separate suppliers, distinct procurement channels, independent regulatory regimes), and the researcher can specify which observable controls restore the assumption when a particular threat is identified. This transparency provides a basis for evaluating the credibility of the estimates that has no analogue under the Markov framework.
On the theoretical side, Theorem (ref) characterizes the residual indeterminacy that arises once the Markov assumption is dropped. Two routes close this indeterminacy: an exclusion restriction (Corollary (ref)) with a testable necessary condition (Remark (ref)), and a homothetic regularity condition (Theorem (ref)). The two routes yield mutually consistent estimates in industries where both apply.
Three limitations and corresponding directions for future work deserve mention. First, the identification strategy requires at least three intermediate inputs with separately observable quantity data, though this requirement is met in several settings beyond the Japanese Census of Manufactures, including the U.S.\ EIA Form 923 fabrizio2007dothey,cicala2015when and emissions data in environmental economics.\footnote{ Additional datasets satisfying this requirement include India's Annual Survey of Industries (ASI), which reports firm-level electricity and fuel consumption alongside materials; Canada's Annual Survey of Manufacturing and Logging (ASML), which covers electricity and water use at the establishment level; and the World Bank Enterprise Survey (WBES), which collects firm-level electricity expenditure and water source data across over 100 countries. These datasets enable direct application of the proposed estimator in diverse institutional settings.} When labor adjustment is rapid, labor itself serves as an additional productivity signal, reducing the required number of intermediate inputs from three to two (footnote (ref)); extending the framework to such settings is a natural direction. Second, no targeted test of the conditional independence assumption alone exists; the convergence diagnostic of Remark (ref) provides a necessary condition for the exclusion restriction. Extending the moment system to achieve overidentification (for instance via Block C structural constraints or cross-equation demand restrictions under Cobb--Douglas) would enable formal specification testing. Third, Block C identification of $(\beta_k, \beta_l)$ requires non-negligible curvature in $h(v)$; when the capital-labor ratio varies little, the exclusion restriction route becomes preferable, and combining the static identification of flexible input elasticities with semiparametric methods for the capital-labor component is left for future research.
\printbibliography
\thispagestyle{empty}