The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
154,546 characters
Nonparametric Identification and Estimation of Production Functions Invariant to Productivity Dynamics
\IfFileExists{Table/Empirical/cross_val_macros.tex}{
}{
}
\title{Nonparametric Identification and Estimation of Production Functions
Invariant to Productivity Dynamics\thanks{I am grateful to Yasutora Watanabe, Yuta Toyama, Shosei Sakaguchi,
Hidehiko Ichimura, and Satoshi Imahie for their insightful comments and detailed discussions.
I also thank Takanori Adachi, Daiya Isogawa, Yuta Kikuchi, Toshifumi Kuroda,
Yusuke Matsuki, Kentaro Nakamura, Masato Nishiwaki, Tatsushi Oka, Ryo Okui,
Naoya Sueishi, Hidenori Takahashi, and Naoki Wakamori for helpful comments,
as well as participants at the Japan Empirical Industrial Organization Workshop
and the Kansai Econometric Society Meeting. This research was financially supported
by the Project Research Program of the Joint Usage/Research Center Programs
at the Institute of Economic Research, Hitotsubashi University
(Grant Number: IERPK2437); the JST SPRING fellowship; and a
Grant-in-Aid for JSPS Fellows (Grant Number: 25KJ0910).
This research was conducted under approval number
20240708-stat-No1 dated July 8, 2024, by the Statistics Bureau,
Ministry of Internal Affairs and Communications. I utilized
microdata from the \textit{Census of
Manufactures} (Ministry of Economy, Trade and Industry) and the
\textit{Economic Census for Business Activity} (Ministry of
Internal Affairs and Communications; Ministry of Economy, Trade
and Industry). The views expressed in this paper are those of the
author and do not necessarily reflect the views of the Japanese
government or the ministries.
All remaining errors are my own.}}
\author{Rentaro Utamaru\thanks{Institute for Research in Contemporary Political and Economic, Waseda University,
1-104 Totsukamachi, Shinjuku-ku, Tokyo 169-8050, Japan. Email: [email removed]}}
\date{April 2026\\[0.5em]
\normalsize\href{https://r-utamaru.github.io/files/utamaru_nonmarkov_production.pdf}{Click here for the latest version}}
\maketitle
\vspace{-1.5em}
\begin{center}
\fbox{\parbox{0.85\textwidth}{\centering\small
\textbf{Preliminary Draft. Comments Welcome.}\\[0.2em]
The core identification theory and GMM estimator are complete.
Empirical results and Monte Carlo simulations are subject to revision.
Extensions to the GMM implementation of exclusion restrictions and
formal specification testing are in progress.}}
\end{center}
\vspace{-0.5em}
\begin{abstract}
Production function estimates underpin the measurement of firm-level
markups, allocative efficiency, and the productivity effects of
policy interventions. Since \textcite{olley1996thedynamics}, every major
proxy variable estimator has identified the production function
through a first-order Markov assumption on unobserved productivity;
I show that misspecification of this assumption generates persistent
upward bias in the materials elasticity that propagates into
overestimated markups and inflated treatment effects. I replace
the Markov restriction with conditional independence across three
intermediate input demands, a static condition grounded in input
market segmentation, and establish nonparametric identification
from a single cross-section. I develop a GMM estimator and
establish consistency and asymptotic normality. Monte Carlo
simulations confirm that the proposed estimator is unbiased across
Markov and non-Markov environments, while the standard estimator
exhibits persistent bias of up to 63 percent of the true materials
elasticity. In 502 Japanese manufacturing industries, the proposed
method yields systematically lower markups than the standard method
across the entire distribution (median 0.93 vs.\ 1.03), reducing
the share of industries with markups above unity from 54 to 37
percent. In a difference-in-differences analysis of the 2011
T\={o}hoku earthquake, the standard method overstates the
productivity loss by 0.40 percentage points, roughly \$3.6
billion (\textyen{}400 billion) per year.\\
\noindent\textsl{Keywords}: Production Function, Productivity,
Nonparametric Identification, Markups, Market Power\\
\textsl{JEL Classification Codes}: C13, C14, D24, L11, L40
\end{abstract}
\clearpage{}
\section{Introduction}
Estimated production functions underpin the measurement of market
power, allocative efficiency, and the effects of policy on firm
performance. The ratio of the materials elasticity to the materials
revenue share gives the firm-level markup
\parencite{deloecker2012markups}; the dispersion of the productivity
residual measures resource misallocation
\parencite{hsieh2009misallocation}; and the productivity level itself
serves as the outcome variable in studies of trade liberalization
\parencite{deloecker2013detecting}, R\&D investment
\parencite{doraszelski2013rdand}, and disaster recovery. These
downstream analyses inherit the production function estimate: if the
materials elasticity is biased, so is the markup, the misallocation
measure, and the treatment effect. The recent finding that markups
have risen across the global economy \parencite{deloecker2020rise}
relies on such estimates, making the consistency of the underlying
production function a first-order concern.
This paper asks whether the production function can be identified
without restricting how productivity evolves over time, and
documents the consequences when this restriction is removed.
Since \textcite{olley1996thedynamics}, every major production function
estimator has relied on the same structural restriction: productivity
must follow a first-order Markov process. This includes the methods of
\textcite{levinsohn2003estimating},
\textcite{ackerberg2015identification}, and
\textcite{gandhi2020onthe}, as well as dynamic panel approaches
\parencite{arellano1991sometests,blundell1998initial}. The Markov
assumption is not a regularity condition; it is the identifying
restriction that pins down the materials elasticity through the
transition equation. When productivity evolves endogenously through
R\&D, learning, or managerial turnover, omitting the relevant state
variables generates a transmission bias
\parencite{deloecker2007doexports,deloecker2013detecting,doraszelski2013rdand}.
The assumption also presupposes a stationary transition process,
ruling out structural breaks from aggregate shocks, regulatory
shifts, or technological change. More fundamentally,
\textcite{chen2024identifying} show that under the
potential outcomes framework, any treatment that alters the transition
path of productivity violates the Markov property by construction,
even when the treatment variable is included as a control.
The bias does not vanish with sample size, nor can it be removed by
adding treatment indicators to the Markov transition equation. The
Markov-based estimate is therefore inconsistent precisely in the
settings where productivity serves as an outcome variable, the
dominant use of production function estimation in applied work.
This paper shows that the Markov assumption is unnecessary for
identification. I replace it with a static condition: conditional
independence of demand shocks across three intermediate inputs (raw
materials, electricity, water). Three flexible inputs whose demands
respond to the same underlying productivity serve as three noisy
measurements of a common latent variable. Because each input is
procured from a separate market, the input-specific demand
shocks are mutually independent conditional on productivity and
observable controls. I recover the productivity distribution from
these signals using the spectral decomposition of
\textcite{hu2008instrumental} (hereafter HS08), without any
restriction on how productivity evolves over time. Identification
requires only a single cross-section; the data requirement
(firm-level quantities of three separate inputs) is met in
manufacturing censuses across several countries.\footnote{These
include India's Annual Survey of Industries, Canada's Annual Survey
of Manufacturing and Logging, the World Bank Enterprise Survey, and
the U.S.\ EIA Form~923. When labor adjusts rapidly to current
productivity, two intermediate inputs suffice
(footnote~\ref{rem:static_labor}).}
The substitution of assumptions has first-order consequences for
economic measurement. In 502 Japanese manufacturing industries, the
proposed method yields systematically lower markups than the
standard ACF method across the entire distribution: the ACF markup
CDF lies strictly to the right at every percentile. At the median,
the gap is 0.10 (proposed 0.93 vs.\ ACF 1.03), and the share of
industries with markups above unity falls from 54 percent under ACF
to 37 percent under the proposed method. The Markov assumption thus
inflates the measured degree of market power across the
manufacturing sector. Monte Carlo simulations
trace the mechanism: under potential outcome dynamics, ACF's bias in
the materials elasticity is $+0.19$ (63 percent of the true value),
while the proposed estimator is unbiased.
In a difference-in-differences analysis of the 2011 T\={o}hoku
earthquake, the standard method overstates the productivity loss by
0.40 percentage points, corresponding to roughly \$3.6 billion
(\textyen{}400 billion) per year. Because identification is static, the estimator
can be applied period by period, producing time-varying estimates
of production technologies without imposing structural stability on
the productivity process. The empirical application documents
substantial temporal variation across 2003--2020 and yields
divergent conclusions regarding allocative efficiency as assessed
through the \textcite{olley1996thedynamics} decomposition. The
Markov assumption does not merely introduce statistical noise; it
systematically inflates measured market power and distorts policy
conclusions.
The substitution involves an honest tradeoff. The Markov assumption,
when it holds, provides efficiency gains by exploiting the
time-series history of productivity. The conditional independence
assumption uses only within-period information, so under correct
Markov specification, standard estimators have lower variance. I
document this in Monte Carlo simulations under correct Markov
specification. The value of
the proposed method lies in the broad class of applications where the
Markov assumption is questionable or directly contradicted by the
research design, including any study in which a treatment alters
productivity dynamics \parencite{chen2024identifying}.
The two assumptions differ in the nature of their economic content.
The Markov restriction constrains the time-series evolution of
unobserved productivity; no economic theory predicts that
productivity should follow a first-order autoregression, and the
assumption cannot be tested within the proxy variable framework.
The conditional independence restriction constrains input market
structure: it specifies what threatens identification (common
shocks across input markets) and what restores it (conditioning on
observable controls $z_{jt}$ that absorb the common component).
The threats are enumerable (demand fluctuations, aggregate markup
changes, correlated procurement), and the defenses are
observable (inventory, aggregate output, fixed effects;
Section~\ref{subsec:relax_indep}). The microfoundations in
Appendix~\ref{sec:micro_fnd} derive the demand shocks from a
cost-minimization problem with input-specific markdowns, making
the economic content of the assumption precise. No analogous transparency is available for the Markov assumption:
within the proxy variable framework, no observable implication
distinguishes a correctly specified AR(1) from an AR(2) or a
potential outcome process. By contrast, the conditional
independence assumption yields a testable necessary condition
(Remark~\ref{rem:testable_exclusion}): in 502 industries, the
pairwise convergence diagnostic supports the identifying
restriction for capital, providing direct evidence on the
empirical plausibility of the assumption. When the assumption is
violated through a common shock to electricity and water (the
most economically salient threat), Monte Carlo analysis
(Section~\ref{sec:simulation}, Appendix~\ref{app:ci_violation})
shows that the resulting bias in $\hat{\beta}_m$ is
\emph{upward}, the same direction as the Markov misspecification
bias. The empirical finding that the proposed method yields
\emph{lower} $\hat{\beta}_m$ than ACF therefore cannot be
explained by conditional independence violation; it is consistent
only with Markov misspecification in the standard estimator.
\begin{table}[tbph]
\centering
{\footnotesize\caption{{\footnotesize{} Comparison with other studies}}\label{tab:comparison}
}{\footnotesize\par}
\begin{adjustbox}{max width=\linewidth}
\begin{tabular}{lV{\linewidth}ccccc}
\toprule
& {\small Identification}{\small\par}
{\small Method} & {\small Req.\ Markov} & \begin{cellvarwidth}[t]
\centering
{\small Req.\ Scalar}{\small\par}
{\small Unobs.}
\end{cellvarwidth} & \begin{cellvarwidth}[t]
\centering
{\small Nonpara}{\small\par}
{\small Non-Hicks}
\end{cellvarwidth} & \begin{cellvarwidth}[t]
\centering
{\small Function}{\small\par}
{\small Type}
\end{cellvarwidth} & \begin{cellvarwidth}[t]
\centering
{\small Proxy or}{\small\par}
{\small Control}
\end{cellvarwidth}\tabularnewline
\midrule
{\small Proposed method} & \citeauthor*{hu2008instrumental} & & & {\small$\checkmark$} & {\small Gross} & {\small$e_{jt},w_{jt}$}\tabularnewline
{\small\textcite{gandhi2020onthe}} & {\small FOC + Markov} & {\small$\checkmark$} & {\small$\checkmark$} & & {\small Gross} & {\small$s_{jt}$ (share)}\tabularnewline
{\small\textcite{dotyadynamic}} & \citeauthor*{hu2008instrumental} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small Gross} & {\small$I_{jt},y_{jt+1}$}\tabularnewline
{\small\textcite{hu2020estimating}} & \citeauthor*{hu2008instrumental} & {\small$\checkmark$} & {\small$\checkmark$} & & {\small Gross} & {\small$I_{jt},m_{jt+1}$}\tabularnewline
{\small\textcite{brandestimating}} & \citeauthor*{hu2008instrumental} & {\small$\checkmark$} & {\small$\checkmark$} & & {\small Gross} & {\small$y_{jt-1},y_{jt+1}$}\tabularnewline
{\small\textcite{zeng2023identification}} & \begin{cellvarwidth}[t]
{\small\citeauthor*{matzkin2003nonparametric}}\\
{\small\citeauthor*{imbens2009identification}}
\end{cellvarwidth} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small Value} & {\small$K_{jt-1},I_{jt-1}$}\tabularnewline
{\small\textcite{ackerberg2022nonparametric}} & \begin{cellvarwidth}[t]
{\small\citeauthor*{matzkin2003nonparametric}}\\
{\small\citeauthor*{imbens2009identification}}
\end{cellvarwidth} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small Gross} & {\small$\left\{ y_{j\tau},x_{j\tau}\right\} ^{t-1}_{\tau=t-M}$}\tabularnewline
{\small\textcite{navarrononparametric}} & \begin{cellvarwidth}[t]
{\small\citeauthor*{matzkin2003nonparametric}}\\
{\small\citeauthor*{imbens2009identification}}
\end{cellvarwidth} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small Gross} & {\small$x_{jt-1},\mathcal{Y}_{jt-1}$}\tabularnewline
{\small\textcite{pan2022identification}} & \begin{cellvarwidth}[t]
{\small\citeauthor*{matzkin2003nonparametric}}\\
{\small\citeauthor*{imbens2009identification}}
\end{cellvarwidth} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small Gross} & {\small$\left\{ y_{j\tau},x_{j\tau}\right\} ^{t-1}_{\tau=t-M}$}\tabularnewline
\bottomrule
\end{tabular}
\end{adjustbox}
\medskip
\small\raggedright \textit{Notes: ``Req.\ Markov'' indicates whether the
method requires a Markov assumption on productivity; a blank cell
indicates the method does not.
``Req.\ Scalar Unobs.'' indicates whether the method requires
scalar unobservability (productivity as the sole unobservable
in input demand); a blank cell indicates that the method permits
input-specific demand shocks.
``Nonpara Non-Hicks'' indicates nonparametric identification
under non-Hicks-neutral specifications.
``Function Type'' distinguishes gross output from value-added
production functions. ``Proxy or Control'' lists the proxy
variables or control variables used for identification.
For the proposed method, the ``Nonpara Non-Hicks'' checkmark
refers to the identification result in
Appendix~\ref{app:prod_recovery}; the implemented estimator
is Hicks-neutral (equation~\eqref{eq:gmm_prod}).
The proposed method requires conditional independence of
input-specific demand shocks
(Assumption~\ref{ass:cond_indep}) in place of the Markov and
scalar unobservability conditions; both blank cells in its row
reflect this substitution, not an absence of identifying
assumptions.}
\end{table}
I make three contributions. First, I show that the cross-sectional
covariance structure among three flexible intermediate inputs fully
substitutes for the Markov restriction, delivering nonparametric
identification of the production function and the productivity
distribution from a single period. This is a substitution, not a
relaxation, of identifying assumptions. The mapping to the HS08
framework provides density identification
(Theorem~\ref{thm:density_id}); the theoretical contribution of
this paper lies in what follows. I characterize the residual
indeterminacy that arises when Markov is dropped: any two
observationally equivalent structures differ only by a location
shift $\Delta(k,l)$ applied to productivity, ruling out nonlinear
transformations (Theorem~\ref{thm:obs_equiv}). I provide two routes
that close this indeterminacy without dynamic assumptions, an
exclusion restriction (Corollary~\ref{thm:exclusion}) and a
homothetic regularity condition
(Theorem~\ref{thm:homothetic_id}).
While nonparametric sieve
estimation could in principle implement the identification results
directly, the high-dimensional numerical integration is
computationally prohibitive for census-scale panels. I develop a
Cobb--Douglas GMM estimator designed for applied use, and establish
its consistency and asymptotic normality
(Theorem~\ref{thm:asymptotics}); the extension to translog
production is developed in Appendix~\ref{app:translog}. The conditional independence
assumption yields a pairwise convergence diagnostic
(Remark~\ref{rem:testable_exclusion}) with no analogue under the
Markov assumption: within the proxy variable framework, no
restriction distinguishes a correctly specified AR(1) from an AR(2)
or a potential outcome process. In 502 industries, this diagnostic
converges to zero for capital but not for labor, providing direct
evidence on the differential applicability of the exclusion
restriction.
Second, I document that the Markov assumption generates a
systematic upward bias in measured market power. Monte Carlo
simulations show that ACF's bias in $\hat{\beta}_m$ does not
vanish as sample size grows: $+0.026$ under AR(2) dynamics and
$+0.266$ under potential outcome dynamics. In the empirical
application, ACF produces higher materials elasticities and higher
markups at every percentile across 502 industries. The gap crosses
the competitive threshold and reverses the policy-relevant
conclusion about market structure. The recovered productivity
measures also show stronger associations with economic fundamentals
than those from the standard method, consistent with a higher
signal-to-noise ratio from separating input-specific demand shocks
(Section~\ref{sec:empirical}).
Third, I show that productivity measures recovered from the
proposed method are valid under the potential outcomes framework
(Proposition~\ref{prop:omega_hat_D}), resolving the inconsistency
identified by \textcite{chen2024identifying}. Because the estimator
uses no transition equation, the recovered productivity is
invariant to how a treatment operates on productivity dynamics.
The earthquake event study illustrates the practical consequence:
the proposed estimate is $-1.28$ percent while ACF yields $-1.68$
percent, a gap that arises because the ACF estimate lacks the
theoretical guarantee that the production function parameters are
consistently estimated under treatment-induced dynamics.
The same static, $\omega$-conditional structure also renders
the estimator robust to endogenous exit: under the standard
timing convention where exit precedes input choice, conditioning
on $\omega$ absorbs survival selection, and no survival probability
correction is needed (Remark~\ref{rem:selection}).
\paragraph{Related literature.}
Table~\ref{tab:comparison} positions my identification strategy
within the recent literature. The most closely related work is
\textcite{gandhi2020onthe} (GNR). GNR's Theorem~1 establishes
that proxy variable methods alone cannot identify the gross
output production function; additional within-period,
cross-sectional information is required. Both approaches
supply such information: GNR through the structural link between the
production function and the firm's first-order condition, yielding
a nonparametric share regression that directly identifies the
flexible input elasticity; my approach through the measurement
error structure of HS08, using conditional independence across
intermediate inputs to recover the distribution of unobserved
productivity.
The two approaches rest on different assumptions regarding input
markets. GNR's first-order condition requires competitive input
markets with common prices and that any unobserved component in the
share equation is non-persistent (their Appendix~O6, Assumption~7);
when input-specific markdowns or procurement frictions are
persistent, the FOC-based estimation equation does not hold and the
share regression is misspecified. My framework
permits persistent, input-specific demand shocks arising from
procurement relationships, supply contracts, or input-specific
markdowns; identification requires only mutual independence across
inputs at each time point, accommodating arbitrary serial
dependence within each shock. GNR's second stage recovers capital
and labor elasticities using the Markov structure; my approach
requires no dynamic assumption at any stage. The scalar
unobservability case is a special case of my model, obtained when
the input-specific shocks are degenerate
(Section~\ref{sec:simulation}).
Alternative approaches that exploit static first-order conditions
\parencite{grieco2016production,caselli2025productivity} avoid
dynamic assumptions but generally require parametric restrictions
on functional forms and the demand system. Additional related work
is summarized in Table~\ref{tab:comparison}. Several recent papers
apply the HS08 framework to production functions
\parencite{brandestimating,hu2020estimating,dotyadynamic},
but all use lagged variables as instruments and therefore retain the
Markov assumption. \textcite{zeng2023identification} avoid the
Markov restriction at the estimation stage but presuppose it for
the investment policy function. A growing literature on
non-Hicks-neutral identification
\parencite{navarrononparametric,ackerberg2022nonparametric,pan2022identification,kasahara2023identification,dotyadynamic},
including factor-augmenting approaches
\parencite{doraszelski2018measuring,demirerproduction,raval2019themicro},
retains first- or higher-order Markov assumptions; my identification
results extend to these models without dynamic restrictions
(Appendix~\ref{app:prod_recovery}), though the implemented
estimator uses the Hicks-neutral Cobb--Douglas specialization
(Section~\ref{sec:gmm_spec}).
The remainder of the paper is organized as follows.
Section~\ref{sec:model_identification} presents the model and the
nonparametric identification results.
Section~\ref{sec:estimation} develops the GMM estimator.
Section~\ref{sec:simulation} presents Monte Carlo evidence.
Section~\ref{sec:empirical} applies the estimator to 502 Japanese
manufacturing industries.
Section~\ref{sec:conclusion} concludes.
\section{Model and Identification}
\label{sec:model_identification}
This section establishes the identification strategy in three steps.
First, I show that three conditionally independent input demands
identify the joint distribution of productivity and inputs within
each capital-labor cell (Theorem~\ref{thm:density_id}); any two
observationally equivalent structures differ only by a location
shift $\Delta(k,l)$ (Theorem~\ref{thm:obs_equiv}).
Second, I provide two conditions that eliminate this indeterminacy:
an exclusion restriction (Corollary~\ref{thm:exclusion}) and a
homothetic regularity condition
(Theorem~\ref{thm:homothetic_id}). The exclusion restriction
carries a testable implication
(Remark~\ref{rem:testable_exclusion}). The formal statement of
density identification (Theorem~\ref{thm:density_id}) and the technical regularity
conditions (Assumptions~\ref{ass:injectivity}--\ref{ass:labeling}) are in
Appendix~\ref{app:identification_details}.
These identification results translate into three groups of moment
conditions in the GMM estimator of
Section~\ref{sec:estimation}: proxy moments (Block~A),
covariance moments (Block~B), and curvature moments (Block~C).
When these terms appear below, they refer forward to
Section~\ref{sec:gmm_approach}.
\subsection{Model Setup}
I define the general gross output production function for
firm $j$ at time $t$ as follows:
\begin{equation}
y_{jt}=f_{t}(k_{jt},l_{jt},m_{jt},e_{jt},w_{jt},\omega_{jt})+\varepsilon_{jt}\label{eq:prod_func}
\end{equation}
Here, $y_{jt}$ is the logarithm of output, $k_{jt}$ and $l_{jt}$
are the logarithms of capital and labor. Following the production function literature
\parencite{olley1996thedynamics,ackerberg2015identification,bond2005adjustment},
capital and labor are treated as dynamic or quasi-fixed inputs whose
current values are predetermined relative to intermediate input
decisions.
The model requires at least three distinct intermediate inputs:
$m_{jt}$ (raw materials), $e_{jt}$ (electricity), and $w_{jt}$ (industrial water).
Three inputs are the minimum required by the
\textcite{hu2008instrumental} spectral decomposition: it identifies
the latent productivity distribution from three mutually independent
measurements of a common latent variable; two measurements do not
suffice for nonparametric identification without additional
restrictions.\footnote{When labor adjusts rapidly to current
productivity, it may serve as a third measurement of $\omega_{jt}$,
reducing the required number of flexible intermediate inputs from
three to two; see Footnote~\ref{rem:static_labor} for details.}
$\omega_{jt}$ is the firm's productivity, unobserved by the
econometrician but known to the firm when making input decisions.
$\varepsilon_{jt}$ denotes ex-post production shocks (measurement
error or unexpected disruptions), unobserved by both the firm and
the econometrician at the time of input choice.
The state variable vector $x_{jt}=(k_{jt},l_{jt},z_{jt})$
determines input demand. Here, $k_{jt}$ and $l_{jt}$ are the
primary inputs, while $z_{jt}$ represents additional firm-specific
state variables such as inventory levels, input prices, or market
conditions that do not directly enter the production function but
influence input demand. Given $x_{jt}$, the demand for each
intermediate input is determined as follows:
\begin{align}
m_{jt} & =g_{m}(x_{jt},\omega_{jt},\tau_{jt})\label{eq:demand_m}\\
e_{jt} & =g_e(x_{jt},\omega_{jt},\nu_{jt})\label{eq:demand_mp}\\
w_{jt} & =g_w(x_{jt},\omega_{jt},\eta_{jt})\label{eq:demand_mpp}
\end{align}
The functions $g_{m}(\cdot)$, $g_e(\cdot)$, and $g_w(\cdot)$
are unknown and potentially nonlinear. $\tau_{jt}$,
$\nu_{jt}$, and $\eta_{jt}$ are unobserved shock terms specific
to each input demand, following
\textcites{hu2020estimating}{brandestimating}{dotyadynamic}. These
shocks capture optimization errors, supply disruptions, and
adjustment frictions not explained by productivity and state
variables. Appendix~\ref{sec:micro_fnd} derives the demand system from
a cost-minimization problem under imperfect input markets and shows
that these shocks correspond to input-specific markdowns, prices,
and wedges; specifically, the components of markdowns and wedges
orthogonal to observable state variables.
The presence of input-specific shocks represents a departure from
the scalar unobservability assumption maintained in
\textcite{olley1996thedynamics},
\textcite{levinsohn2003estimating},
\textcite{ackerberg2015identification}, GNR, and others, which
requires productivity to be the sole unobservable affecting input
demand. When scalar unobservability fails because firm-level input
prices, markdowns, or wedges are unobserved, standard proxy variable
estimators are inconsistent
\parencite{jaumandreu2021reexamining,doraszelski2025production}. In
my framework, all unobserved firm-specific heterogeneity beyond
productivity is absorbed into $\tau_{jt}, \nu_{jt}, \eta_{jt}$, and
identification requires only that these shocks be mutually
independent across inputs, not that they be absent. Scalar
unobservability is nested as the special case $\tau = \nu = \eta = 0$
at the model level; the identification strategy requires
non-degenerate demand shocks and is therefore complementary to,
rather than a generalization of, scalar inversion methods.
From the standpoint of the cost-minimization model in
Appendix~\ref{sec:micro_fnd}, $\tau = \nu = \eta = 0$ requires
that all firms in an industry face identical input prices, identical
markdowns in every input market, and make no optimization errors in
input choice. In practice, firms negotiate procurement contracts
individually, face supplier-specific delivery terms, and adjust
input quantities with heterogeneous frictions. The presence of
input-specific demand shocks is the empirically relevant case; the
proposed framework treats these shocks as a source of identifying
information rather than a nuisance to be assumed away.
This formulation also addresses the collinearity problem identified
by \textcite{gandhi2020onthe}: under scalar unobservability, flexible
inputs determined by static optimization lack sufficient residual
variation to identify the gross production function
\parencite{ackerberg2015identification,bond2005adjustment}. GNR
resolve this problem by exploiting the first-order condition for the
flexible input, which identifies its output elasticity from the
revenue share. My approach resolves the collinearity through
independent input-specific shocks, which supply the cross-sectional
variation needed for identification via the measurement error
structure of HS08, without relying on the first-order condition or
dynamic moment conditions. The practical difference is that the
share regression requires the first-order condition to hold with
common input prices, whereas my approach permits firm-specific
input prices and markdowns (Appendix~\ref{sec:micro_fnd}).
\subsection{Assumptions for Identification}
\label{sec:assumptions}
The identification theory rests on two substantive assumptions
stated here, together with three regularity conditions
(Assumptions~\ref{ass:injectivity}--\ref{ass:labeling}) collected in
Appendix~\ref{app:identification_details}.
\begin{assumption}[Additive Error Structure]
\label{ass:additive_error}
The production function has an additive error structure:
\begin{equation}
y_{jt} = f_t(k_{jt}, l_{jt}, m_{jt}, e_{jt}, w_{jt},
\omega_{jt}) + \varepsilon_{jt},
\label{eq:prod_func_additive}
\end{equation}
where the ex-post shock $\varepsilon_{jt}$ satisfies
\[
\mathbb{E}\bigl[\varepsilon_{jt} \mid k_{jt}, l_{jt}, m_{jt},
e_{jt}, w_{jt}, \omega_{jt}\bigr] = 0.
\]
\end{assumption}
\textit{Role and economic content.} This is standard in the production
function literature \parencite{olley1996thedynamics,ackerberg2015identification}.
The shock $\varepsilon_{jt}$ captures ex-post deviations (measurement
error, unexpected disruptions) that are realized after input choices
are made and are therefore uncorrelated with all inputs and
productivity. It acts as classical measurement error in the
dependent variable and inflates standard errors but does not bias
the production function estimates
(Theorem~\ref{thm:asymptotics}).
\begin{assumption}[Conditional Independence]
\label{ass:cond_indep}
The demand shocks $(\tau_{jt}, \nu_{jt}, \eta_{jt})$ for the three
intermediate inputs are mutually independent, conditional on
productivity $\omega_{jt}$ and state variables
$x_{jt} = (k_{jt}, l_{jt}, z_{jt})$:
\[
f_{\tau, \nu, \eta \mid \omega, x}
= f_{\tau \mid \omega, x} \cdot f_{\nu \mid \omega, x}
\cdot f_{\eta \mid \omega, x}.
\]
Mutual independence is required; pairwise independence does not
suffice for the spectral decomposition of HS08.
\end{assumption}
\textit{Role.} This is the substantive identifying condition. Together
with the regularity conditions in
Appendix~\ref{app:identification_details}
(Assumptions~\ref{ass:injectivity}--\ref{ass:labeling}), it enables
the unique spectral decomposition of the integral
equation~\eqref{eq:integral_eq}. Conditional independence is the
economically substantive condition; it restricts the data generating
process rather than regularity of the operators.
\textit{Economic content.} The assumption posits that, for a firm
with given state variables and productivity level, an unexpected
shock to raw material demand (e.g., a supply chain disruption) is
independent of a shock to electricity demand (e.g., an unscheduled
rate surcharge). This is natural when input markets are
segmented: raw materials, electricity, and water are procured
through distinct channels, under separate contracts, with
different suppliers. The common components of demand variation (product demand
fluctuations, aggregate markup changes) are captured by $x_{jt}$;
$\tau_{jt}, \nu_{jt}, \eta_{jt}$ represent the residual,
input-specific components. The microfoundations in
Appendix~\ref{sec:micro_fnd} make this structure precise.
\subsection{Interpretation and Robustness of the Conditional Independence Assumption}
\label{subsec:relax_indep}
The general principle is as follows. Common shocks that affect all three input
demands (product demand fluctuations, markup variation, aggregate
input price movements) can be absorbed by projecting onto
observable control variables $z_{jt}$; the shock terms
$\tau_{jt}, \nu_{jt}, \eta_{jt}$ are then defined as the
orthogonal residuals of this projection
(Appendix~\ref{sec:micro_fnd}). The independence assumption
therefore requires only that the \textit{residual}, input-specific
components of demand variation are mutually independent.
Several potential threats illustrate this principle.
\textit{Unobserved demand shocks} generate common variation across
all inputs, but can be proxied by inventory fluctuations
\parencite{kumar2019productivity} or recovered from revenue data
\parencite{kasahara2020nonparametric}, included in $z_{jt}$.
\textit{Product market power} affects all input demands through
marginal revenue; following
\textcite{ackerbergproduction,jaumandreu2025robustproduction},
low-dimensional sufficient statistics for the markup (e.g.,
competitors' output, average variable cost) can be included in
$z_{jt}$.\footnote{Under Cournot competition,
\textcite{ackerbergproduction} show that the total output of
competitors serves as a sufficient statistic.}
\textit{Input market power} (markdowns) may generate common
bargaining advantages across inputs, but the common component
depends on firm attributes (size, liquidity) captured by
$(k_{jt}, l_{jt}, z_{jt})$; what remains in the shock terms are
idiosyncratic variations from individual supplier relationships.
It is economically reasonable that the outcome of negotiations
with raw material suppliers is independent of electricity rate
negotiations, conditional on firm size and other
observables.\footnote{When an intermediate input is traded on
competitive commodity markets, the firm is a price-taker and
the markdown on that input vanishes.
\textcite{avignon2025markups} exploit this property for
globally traded dairy commodities to separately identify
markups and markdowns on other inputs.}
\textit{Common input price shocks} (e.g., oil price hikes)
affect multiple inputs symmetrically and are controlled by time
fixed effects or industry-specific deflators in $z_{jt}$.
Firm-specific price variations are absorbed as part of the
structural shock terms and need only be independent across inputs.
\subsection{Identification of the Production Function}
\label{sec:prod_func_id}
The identification proceeds in two stages: first, I recover the
production function and productivity distribution within each
capital-labor cell $(k_0, l_0)$; second, I characterize and resolve
the residual indeterminacy that arises when linking these
cell-specific results across different values of $(k, l)$.
\footnote{In the following, firm subscripts $j$ are suppressed as
I discuss population-level arguments. The time subscript $t$ is
retained only to indicate time-variation in the production function
$f_t$.}
\subsubsection{Identification within Each $(k, l)$}
\label{sec:within_kl}
The foundational identification result applies the spectral
decomposition of HS08, whose conditions I verify under the
present assumptions.
\begin{thm}[Identification of Densities]
\label{thm:density_id}
Under Assumptions~\ref{ass:additive_error}--\ref{ass:cond_indep}
and Assumptions~\ref{ass:injectivity}--\ref{ass:labeling},
the observable conditional joint density
$f_{m_{jt}, e_{jt} \mid x_{jt}, w_{jt}}$ uniquely identifies
the three unknown conditional density functions:
$f_{m_{jt} \mid \omega_{jt}, x_{jt}}$,
$f_{e_{jt} \mid \omega_{jt}, x_{jt}}$, and
$f_{\omega_{jt} \mid x_{jt}, w_{jt}}$.
\end{thm}
The proof, which verifies the conditions of HS08's Theorem~1
for the integral equation~\eqref{eq:integral_eq}, is in
Appendix~\ref{app:density_id_proof}.
As a consequence of Theorem~\ref{thm:density_id} and
equation~\eqref{eq:bayes_omega}, for each fixed $(k_0, l_0)$, the
following are nonparametrically identified: the conditional densities
$f_{m \mid \omega, k_0, l_0}$, $f_{e \mid \omega, k_0, l_0}$,
$f_{w \mid \omega, k_0, l_0}$, and
$f_{\omega \mid k_0, l_0, m, e, w}$.
Using these identification results, I recover the structure of
$f_t$ as a function of $(m, e, w, \omega)$. I focus on the Hicks-neutral specification, widely adopted in the
empirical literature, and defer the general case to
Appendix~\ref{app:prod_recovery}. Under this specification
$y = g_t(k, l, m, e, w) + \omega + \varepsilon$,
Assumption~\ref{ass:additive_error} implies
\begin{equation}
g_t(k_0, l_0, m, e, w)
= \mathbb{E}[y \mid x, m, e, w]
- \mathbb{E}[\omega \mid x, m, e, w].
\label{eq:hicks_neutral_id}
\end{equation}
Here $g_t$ represents the component of the production technology that
depends on intermediate inputs, with productivity $\omega$ separated
out. The first term on the right-hand side is a conditional
expectation identified directly from the data, and the second is
computable from the posterior density in
equation~\eqref{eq:bayes_omega}. Thus $g_t$ is identified as a
function of $(m, e, w)$ without additional assumptions. For the
general non-Hicks-neutral model,
$f_t(k_0, l_0, m, e, w, \omega)$ is identified as a function of
$(m, e, w, \omega)$ under additional regularity conditions on the
distribution of $\varepsilon$; see Appendix~\ref{app:prod_recovery}
for details.
For each fixed $(k_0, l_0)$, the conditional distribution
$f_{\omega \mid k_0, l_0, m, e, w}$ is fully characterized,
and the conditional expectation
\begin{equation}
\hat{\omega}_{jt} \equiv
\mathbb{E}[\omega_{jt} \mid x_{jt}, m_{jt}, e_{jt}, w_{jt}]
= \int \omega \, f_{\omega \mid x, m, e, w}
(\omega \mid \cdot)\, d\omega
\label{eq:omega_hat}
\end{equation}
provides a firm-level productivity measure for each firm $j$ and
period $t$. The empirical applications of this within-$(k,l)$
identification, including markup estimation and policy evaluation,
are developed in Section~\ref{sec:implications_empirical} after the
identification theory is completed.
However, to identify $f_t$ as a function of $(k, l)$ as well,
additional structure is needed. (When labor adjusts rapidly to
current productivity, it serves as an additional measurement,
reducing the required intermediate inputs from three to
two.\footnote{\label{rem:static_labor}When labor adjusts within the
production period, it serves as a third measurement of
$\omega_{jt}$, and the HS08 identification procedure
(Theorem~\ref{thm:density_id}) applies to the triple
$(l_{jt}, m_{jt}, e_{jt})$, reducing the required flexible
intermediate inputs from three to two. This extension applies when
adjustment costs are small enough that $l_{jt}$ responds to
within-period productivity innovations; industries with high
turnover or temporary staffing (e.g., food processing, garment
manufacturing) are natural candidates. When labor is quasi-fixed,
$l_{jt}$ reflects past rather than current productivity, and the
conditional independence conditions for $(l, m, e)$ do not hold.
See Appendix~\ref{app:prod_recovery} for details.})
$\omega$ must be defined on a common scale across different values of
$(k, l)$. Since Theorem~\ref{thm:density_id} applies the HS08
procedure independently for each $(k, l)$, there is no automatic
correspondence between the $\omega$ values identified at
$(k_1, l_1)$ and those identified at $(k_2, l_2)$. I now formalize
this problem.
\subsubsection{Observational Equivalence and Limits of
Identification}
\label{sec:obs_equiv}
Theorem~\ref{thm:density_id} identifies the production function
within each $(k_0, l_0)$, but a practitioner needs parameters that
are comparable across different capital-labor combinations. The next
result shows exactly what remains unresolved and rules out the
possibility that the indeterminacy takes a nonlinear form.
Under Assumptions~\ref{ass:additive_error}--\ref{ass:cond_indep} and the
regularity conditions in Appendix~\ref{app:identification_details}
(Assumptions~\ref{ass:injectivity}--\ref{ass:labeling}), the conditional densities
$f_{m|\omega}$, $f_{e|\omega}$, $f_{w|\omega}$ and the marginal
density $f_\omega$ are nonparametrically identified from the joint
density of $(m,e,w)$ conditional on $(k,l,z)$
(Theorem~\ref{thm:density_id}). This pins down the shape of each
conditional distribution but leaves a common location shift
$\Delta(k,l)$ applied to the latent variable unresolved. The
following theorem characterizes this residual indeterminacy
completely.
\begin{thm}[Complete Characterization of Observational Equivalence]
\label{thm:obs_equiv}
Under Assumptions~\ref{ass:additive_error}--\ref{ass:cond_indep}
and Assumptions~\ref{ass:injectivity}--\ref{ass:labeling}, a
necessary and sufficient condition for two structures
$(f_t, \omega)$ and $(\tilde{f}_t, \tilde{\omega})$ to generate the
same joint distribution of observables is that there exists a
continuous function $\Delta(k, l)$ such that
\begin{equation}
\tilde{\omega} = \omega + \Delta(k, l), \qquad
\tilde{f}_t(k, l, m, e, w, \tilde{\omega})
= f_t(k, l, m, e, w, \tilde{\omega} - \Delta(k, l)).
\label{eq:obs_equiv}
\end{equation}
\end{thm}
The proof is given in Appendix~\ref{app:proof_obs_equiv}; the key
steps are as follows. The HS08 eigenvalue-eigenfunction
decomposition uniquely determines the functional form of each
conditional density within each $(k_0, l_0)$, ruling out nonlinear
transformations of $\omega$. Any remaining degree of freedom must
therefore be a location shift that varies across $(k, l)$,
yielding~\eqref{eq:obs_equiv}. The continuity of $\Delta(k,l)$
follows from the continuous dependence of
$f_{m \mid \omega, k, l}$ on $(k, l)$ (stated after
Assumption~\ref{ass:labeling}) together with the perturbation
theory of compact operators under simple eigenvalues
(Assumption~\ref{ass:distinct_eigen}; see
Appendix~\ref{app:proof_obs_equiv} for details).
Nonlinear transformations (including scale transformations) are ruled
out because the eigenvalue--eigenfunction decomposition in HS08
uniquely determines the functional form of each conditional density
within each $(k_0, l_0)$. Second, the $\Delta(k, l)$ indeterminacy
arises inherently from the fact that Theorem~\ref{thm:density_id}
applies the HS08 procedure independently for each $(k, l)$. Within
each $(k_0, l_0)$, Assumption~\ref{ass:labeling} fixes the level of
$\omega$, but the reference point of this normalization may depend on
$(k_0, l_0)$. The data on conditional distributions of intermediate
input demands do not contain information to unify $\omega$ levels
across different $(k, l)$.\footnote{
\textcite{hahn2023identification} show that in dynamic approaches
such as \textcite{olley1996thedynamics}, the identification of
dynamic input elasticities relies on an index restriction that
collapses state variables into a one-dimensional scalar.
Theorem~\ref{thm:density_id} does not provide such an index
restriction, and hence the indeterminacy with respect to the dynamic
elasticities persists.}
Economically, the $\Delta(k, l)$ indeterminacy means that the effect
of $(k, l)$ on the production function and
$\mathbb{E}[\omega \mid k, l]$ cannot be separated without
additional restrictions. As a direct consequence,
$f_t$ is identified up to the specification of
$\mathbb{E}[\omega \mid k, l]$: fixing
$\mathbb{E}[\omega \mid k, l]$ pins down $\Delta = 0$
(Theorem~\ref{thm:id_up_to_E},
Appendix~\ref{app:identification_details}).
The $\Delta(k, l)$ indeterminacy
also arises in the existing literature:
\textcite{gandhi2020onthe} resolve it in the Hicks-neutral setting
by combining first-order conditions with a Markov assumption, which
reduces $\Delta(k, l)$ to a constant; for non-Hicks-neutral models,
this strategy fails because $\omega$ cannot be separated from the
first-order condition.\footnote{In the Hicks-neutral model,
$\Delta(k, l)$ shifts the $(k,l)$ component of the production
function:
$\tilde{g}_t(k, l, \cdot) = g_t(k, l, \cdot) - \Delta(k, l)$.
In non-Hicks-neutral models, the FOC
$P_t \cdot (\partial f_t / \partial M) = \rho_t$ retains $\omega$
on the left-hand side, precluding a share regression.
\textcite{li2024identification} show that heterogeneous output
elasticities with respect to flexible inputs remain identifiable
under a scalar unobservable assumption on the proxy variable.}
I provide two alternative methods that close the identification
gap without dynamic assumptions. Theorem~\ref{thm:obs_equiv}
guarantees that $\Delta(k, l)$ is a continuous function of $(k,l)$
alone, which both methods exploit. Section~\ref{sec:exclusion}
imposes exclusion restrictions on intermediate input demands that
directly constrain the functional form of $\Delta(k, l)$, achieving
nonparametric point identification. Section~\ref{sec:homothetic}
parametrically specifies the $(k, l)$ component and introduces a
regularity condition on the shape of $\mathbb{E}[\omega \mid k, l]$,
achieving parametric identification through the non-constant
curvature of the homothetic transformation.
\subsubsection{Closing the Identification Gap}
\label{sec:closing_gap}
The $\Delta(k,l)$ indeterminacy is the cost of dispensing with the
Markov assumption. I now show this cost is payable: two conditions,
each operating without dynamic restrictions, eliminate the
indeterminacy and deliver point identification.
\subsubsection{Nonparametric Identification via Exclusion Restrictions}
\label{sec:exclusion}
The $\Delta(k, l)$ indeterminacy arises because the location
normalization in Assumption~\ref{ass:labeling} is applied
independently for each $(k, l)$ (Theorem~\ref{thm:obs_equiv}). If
the HS08 location normalization
$M[f_{m \mid \omega, k, l}(\cdot \mid \omega)] = \omega$ could be
applied \textit{uniformly} across all $(k, l)$, then
$\Delta(k, l) = 0$ would follow immediately. However, for this
uniform normalization to hold,
$M[f_{m \mid \omega, k, l}(\cdot \mid \omega)]$ must not depend on
$(k, l)$; that is, the conditional demand for the intermediate
input, given $\omega$, must be independent of $(k, l)$. This
observation suggests that exclusion restrictions on intermediate
input demands directly constrain
$\Delta(k, l)$.
\begin{corollary}[Identification via Exclusion Restrictions]
\label{thm:exclusion}
In addition to
Assumptions~\ref{ass:additive_error}--\ref{ass:cond_indep}
and~\ref{ass:injectivity}--\ref{ass:labeling}, suppose one of
the following conditions holds:
\begin{enumerate}
\item[(i)] The demand for some intermediate input (e.g., $w$) does
not depend on $(k, l)$:
$f_{w \mid \omega, k, l} = f_{w \mid \omega}$.
\item[(ii)] The demand for one input (e.g., $m$) does not depend on
$k$, and the demand for another (e.g., $e$) does not depend on
$l$: $f_{m \mid \omega, k, l} = f_{m \mid \omega, l}$ and
$f_{e \mid \omega, k, l} = f_{e \mid \omega, k}$.
\end{enumerate}
Then, under the normalization $\mathbb{E}[\omega] = 0$, the
production function $f_t$ is nonparametrically point-identified.
Condition~(i) is a special case of condition~(ii).
\end{corollary}
\begin{proof}
By Theorem~\ref{thm:obs_equiv}, observationally equivalent
structures are parameterized by
$\tilde{\omega} = \omega + \Delta(k, l)$. Requiring that the
exclusion restriction be maintained in the alternative structure:
\textit{Condition~(i):}
$f_{w \mid \omega, k, l} = f_{w \mid \omega}$ implies that in
the alternative structure,
$f_{w \mid \tilde{\omega}, k, l}(w \mid \tilde{\omega})
= f_{w \mid \omega}(w \mid \tilde{\omega} - \Delta(k, l))$,
which is independent of $(k, l)$ only if $\Delta(k, l)$ is constant.
\textit{Condition~(ii):}
$f_{m \mid \omega, k, l} = f_{m \mid \omega, l}$ implies $\Delta$
does not depend on $k$.
$f_{e \mid \omega, k, l} = f_{e \mid \omega, k}$ implies $\Delta$
does not depend on $l$. Together, $\Delta$ is constant.
In both cases, $\mathbb{E}[\omega] = 0$ pins down $\Delta = 0$.
\end{proof}
Economically, condition~(i) requires that the demand for some
intermediate input (e.g., electricity) depends on productivity alone
and not on capital or labor intensity; this
may hold in energy-intensive industries where electricity consumption
is driven by production volume rather than by the composition of
capital equipment. Condition~(ii) requires that different inputs
exclude different primary inputs from their demand: for example, raw
material demand does not depend on labor intensity, and fuel demand
does not depend on capital intensity. These exclusion restrictions limit the scope of application to
industries where institutional knowledge supports them. For
settings where
such restrictions cannot be justified, I provide a parametric
alternative in the next subsection.
\begin{remark}[Testability of the Exclusion Restriction]
\label{rem:testable_exclusion}
Under the linear demand specification~\eqref{eq:gmm_demand_m}--\eqref{eq:gmm_demand_mpp},
let $a_{k}^h$, $a_{l}^h$, and $a_{\omega}^h$ denote the slope coefficients
on $k$, $l$, and $\omega$ in the demand for input $h$:
$(a_{k}^m, a_{l}^m, a_{\omega}^m) = (\gamma_k, \gamma_l, \gamma_\omega)$,
$(a_{k}^e, a_{l}^e, a_{\omega}^e) = (\delta_k, \delta_l, \delta_\omega)$,
$(a_{k}^w, a_{l}^w, a_{\omega}^w) = (\zeta_k, \zeta_l, \zeta_\omega)$.
The exclusion restriction of Corollary~\ref{thm:exclusion} for a
single input $h$ is not separately testable from Block~A+B
estimates. Under the normalization $\beta_k = \beta_l = 0$
(Section~\ref{sec:gmm_approach}), the estimated demand
coefficient $\hat{a}_{k}^{h *}$ converges to
$a_{k}^{h} - a_{\omega}^{h}\beta_k$, confounding the structural
exclusion parameter $a_{k}^{h}$ with the indeterminacy
$a_{\omega}^{h}\beta_k$ from Theorem~\ref{thm:obs_equiv}.
The \emph{joint} restriction across inputs, however, yields a
diagnostic test, a necessary condition for consistency with the
exclusion restriction, but not a sufficient one. Under
Proposition~\ref{prop:excl_ols}, the OLS estimate
$\hat{\beta}_k^{(h)}$ from input $h$ converges to
$\beta_k - a_{k}^{h}/a_{\omega}^{h}$. Define the pairwise
discrepancy
\begin{equation}
d_k^{(h_1, h_2)}
\equiv
\frac{\hat{a}_{k}^{h_2 *}}{\hat{a}_{\omega}^{h_2}}
- \frac{\hat{a}_{k}^{h_1 *}}{\hat{a}_{\omega}^{h_1}}
\xrightarrow{p}
\frac{a_{k}^{h_2}}{a_{\omega}^{h_2}}
- \frac{a_{k}^{h_1}}{a_{\omega}^{h_1}},
\label{eq:overid_diff}
\end{equation}
which is free of the $\Delta(k,l)$ indeterminacy since
$\beta_k$ cancels in the difference. Under the joint
exclusion restriction $a_{k}^{h_1} = a_{k}^{h_2} = 0$,
$d_k = 0$; the converse does not hold. The test statistic
$d_k = 0$ is a \emph{necessary condition} for the full joint exclusion
restriction, not a sufficient one: $d_k = 0$ also obtains in the
knife-edge case where $a_{k}^{h}/a_{\omega}^{h}$ is equal across inputs
but nonzero. This configuration has no structural basis when the three
inputs involve distinct procurement channels, but the possibility
cannot be ruled out on the basis of $d_k$ alone
(Appendix~\ref{app:testability_exclusion}).
The test is therefore best interpreted as a \emph{diagnostic}: a
rejection of $d_k = 0$ is evidence against the exclusion restriction,
while non-rejection is consistent with, but does not establish, it.
Since $d_k$ is a smooth function of the
Block~A+B parameters, its standard error is obtained by the
delta method from the GMM variance-covariance matrix,
yielding a Wald test without the generated-regressors
problem that would arise from testing OLS estimates
directly. With three inputs, the formal test has two
degrees of freedom ($d_k = d_l = 0$)
(Appendix~\ref{app:testability_exclusion}).
I apply this test in Section~\ref{subsec:exclusion_test}.
\end{remark}
The formal statement and proof are given in
Proposition~\ref{prop:excl_ols}
(Appendix~\ref{app:identification_details}).\footnote{Replacing the
linear subtraction of $\hat{\omega}^h$ in
Proposition~\ref{prop:excl_ols} with a polynomial regression
is not consistent in general; see
Appendix~\ref{app:nonlinear_specs} for details.}
\subsubsection{Parametric Identification via Homothetic Regularity}
\label{sec:homothetic}
As an alternative when exclusion restrictions cannot be justified, I
parametrically specify the $(k, l)$ component and introduce a
regularity condition on $\mathbb{E}[\omega \mid k, l]$. Consider the
additively separable model
\begin{equation}
y = g(k, l;\, \theta) + q(m, e, w) + \omega + \varepsilon,
\label{eq:parametric_model}
\end{equation}
where $g$ is parametric with known functional form and $q$ is
nonparametric. From Section~\ref{sec:within_kl}, $q$ is
nonparametrically recoverable for each fixed
$(k_0, l_0, \omega_0)$.
Specializing to $g(k, l;\, \theta) = \beta_k k + \beta_l l$,
Theorem~\ref{thm:obs_equiv} reduces the identification indeterminacy
to
\begin{equation}
\Delta(k, l) = c_k k + c_l l,
\quad (c_k, c_l) \in \mathbb{R}^2.
\label{eq:linear_ambiguity}
\end{equation}
To eliminate this two-dimensional indeterminacy, I introduce the
following regularity condition.
\begin{assumption}[Homothetic Weak Separability]
\label{ass:homothetic}
The conditional expectation of TFP in the cross-section has a
homothetic structure: there exist continuously differentiable
functions $h \colon \mathbb{R} \to \mathbb{R}$ and
$v \colon \mathbb{R}^2 \to \mathbb{R}$ such that
\[
\bar{\omega}(k, l) \equiv \mathbb{E}[\omega \mid k, l]
= h(v(k, l)),
\]
where:
\begin{enumerate}
\item[(A)] \textit{Nonlinear transformation:} $h'$ is not a constant
function.
\item[(B)] \textit{Translation homogeneity:} $v$ satisfies $v(k+c, l+c) = v(k,l) + c$ for all $c \in \mathbb{R}$.
\item[(C)] \textit{Imperfect substitutability:} The isoquants of $v$
are strictly convex, and the marginal rate of substitution
$v_k / v_l$ is not constant on $(k, l)$.
\end{enumerate}
\end{assumption}
All three conditions are necessary for
Theorem~\ref{thm:homothetic_id}: (A) prevents observational
equivalence with linear functions; (B) ensures the counterfactual
index is also translation homogeneous, so that the MRS of
$\tilde{v}$ is translation invariant; (C) excludes Cobb--Douglas,
where $v_k/v_l$ is constant and a one-dimensional indeterminacy
persists. Economically, (A) requires nonlinear returns to the input
bundle, (B) corresponds to constant returns to scale in the level
variables (since translation homogeneity on the log scale is
equivalent to degree-one homogeneity in levels), and (C)
requires a finite and non-unit elasticity of substitution, satisfied
by CES, translog, and normalized quadratic forms.
Assumption~\ref{ass:homothetic} can be checked from Blocks~A
and~B alone (Section~\ref{sec:diagnostics}); detailed necessity
arguments and testability procedures are in
Appendix~\ref{app:assumption6_details}.
To illustrate, consider the CES specification where
$v(k, l) = \frac{1}{\rho_v}\log\bigl(\alpha e^{\rho_v k} + (1-\alpha) e^{\rho_v l}\bigr)$
is translation homogeneous on the log scale: $v(k+c, l+c) = v(k,l) + c$.
With $h(v) = \gamma v$ (for $\gamma \neq 0$ and higher-order terms
$\rho_2 v^2 + \rho_3 v^3$ with $\rho_2 \neq 0$ or $\rho_3 \neq 0$),
$h'$ is non-constant (satisfying (A)), $v$ is translation homogeneous
(satisfying (B)),
and the MRS $v_k/v_l = [\alpha/(1-\alpha)] e^{\rho_v(k-l)}$ is
non-constant for $\rho_v \neq 0$ (satisfying (C)). The
Cobb--Douglas case ($\rho_v \to 0$, so $v \to \alpha k + (1-\alpha) l$)
yields a linear $v$ and a constant MRS, violating conditions~(A) and~(C)
simultaneously; the rank condition in Theorem~\ref{thm:homothetic_id}
fails, and $(\beta_k, \beta_l)$ cannot be separately identified.
More generally, when $\rho_v$ is close to zero, identification of
$(\beta_k, \beta_l)$ through Block~C becomes weak: the marginal rate
of substitution $v_k/v_l$ approaches a constant as $\rho_v \to 0$,
so the cross-sectional variation in $(k_{jt}, l_{jt})$ provides little
leverage on the curvature parameters. In the empirical analysis, the
$t$-statistics for $\hat{\rho}_2$ and $\hat{\rho}_3$
(Section~\ref{sec:diagnostics}) provide a direct diagnostic for this
failure; industries where both are statistically insignificant should
not be relied upon for separate identification of $\beta_k$ and $\beta_l$
through Block~C alone.
When $\rho_v = 0$, the exclusion restriction of
Corollary~\ref{thm:exclusion} provides an alternative identification
route.
\begin{thm}[Static Parametric Identification]
\label{thm:homothetic_id}
Under Assumptions~\ref{ass:additive_error}--\ref{ass:cond_indep},
\ref{ass:injectivity}--\ref{ass:labeling},
and~\ref{ass:homothetic}, the structural parameters $\beta_k$
and $\beta_l$ in model~\eqref{eq:parametric_model} are
point-identified from static data alone.
\end{thm}
\begin{proof}
By contradiction. Suppose an observationally equivalent
$\tilde{\beta}_i = \beta_i + c_i$ ($i = k, l$) exists with
$(c_k, c_l) \neq (0, 0)$. By~\eqref{eq:linear_ambiguity}, the
alternative TFP function satisfies
$\mathbb{E}[\tilde{\omega} \mid k, l]
= h(v(k, l)) - c_k k - c_l l$. Requiring that
$\mathbb{E}[\tilde{\omega} \mid k, l]
= \tilde{h}(\tilde{v}(k, l))$ for some translation homogeneous
$\tilde{v}$ and differentiable $\tilde{h}$, the translation
invariance of the marginal rate of substitution of $\tilde{v}$
requires
\[
(c_l\, v_k - c_k\, v_l)
\bigl(h'(v + c) - h'(v)\bigr) = 0
\]
for all $(k, l) \in \mathbb{R}^2$ and $c \in \mathbb{R}$. By
condition~(A), $h'$ is non-constant, so the second factor is nonzero
for some $(v_0, c_0)$. Hence
$c_l\, v_k - c_k\, v_l = 0$ everywhere, so $v_k / v_l$ equals
the constant $c_k/c_l$. Under translation homogeneity, a constant
MRS forces $v(k,l) = \alpha k + (1-\alpha)l$, which is linear in
$(k,l)$, contradicting condition~(C). Therefore
$(c_k, c_l) = (0, 0)$.
Condition~(B) (translation homogeneity) enters the argument through
the translation invariance of the MRS of $\tilde{v}$: since
$\tilde{v}_k + \tilde{v}_l = 1$ (implied by translation homogeneity),
without it $\tilde{v}$ need not be translation homogeneous, and the
equality $c_l\, v_k - c_k\, v_l = 0$ does not follow.
\end{proof}
Theorem~\ref{thm:homothetic_id} is stated and proved for the CES
specification of $v(k,l)$; the argument extends to other parametric
forms (e.g., translog) subject to verifying the rank condition
specific to each functional form.\footnote{For instance, with a translog specification
$g = \beta_k k + \beta_l l + \beta_{kk} k^2 + \beta_{ll} l^2
+ \beta_{kl} kl$, $\Delta$ is restricted to the corresponding
polynomial class and the homothetic regularity condition eliminates
the indeterminacy by a similar argument, but the conditions on the
MRS differ from the CES case.}
\subsection{Implications for Empirical Applications}
\label{sec:implications_empirical}
The within-$(k_0, l_0)$ identification results have direct empirical
applications that differ in what they require. Markup estimation
requires only $\beta_m$, which is identified by Blocks~A and~B alone.
Event studies and difference-in-differences designs similarly require
only Block~A+B: because the estimator uses no transition equation for
$\omega$, the recovered $\hat{\omega}_{jt}$ is valid under any
productivity dynamics, including treatment-induced non-Markov paths
(Proposition~\ref{prop:omega_hat_D}, Appendix~\ref{app:identification_details}).
Full productivity-level analysis (including the identification of
$\beta_k$ and $\beta_l$) requires Block~C in addition.
\paragraph{Applications.}
Because estimation does not employ a transition process for
$\omega$, the estimates are invariant to how a policy $D_{jt}$
affects productivity dynamics
(Proposition~\ref{prop:omega_hat_D},
Appendix~\ref{app:identification_details}).
For markup estimation, the within-$(k_0, l_0)$ results suffice:
output elasticities $\partial f_t / \partial m$ are identified for
each fixed $(k_0, l_0)$, which recovers markups as the ratio of the
output elasticity to the revenue share
\parencite{deloecker2012markups}.
\begin{remark}[Functional Form Generality]
\label{rem:functional_form}
The identification results of this paper rest on the conditional
independence of intermediate input demands (Assumption~\ref{ass:cond_indep}),
not on the functional form of production. Theorems~\ref{thm:density_id}
and~\ref{thm:obs_equiv} establish nonparametric identification via
the HS08 spectral decomposition for any production function satisfying
Assumptions~\ref{ass:additive_error}--\ref{ass:cond_indep}. The GMM
estimator of Section~\ref{sec:estimation} implements this under
Cobb--Douglas, where input demands are linear in productivity and the
moment conditions take a tractable linear form.
Appendix~\ref{app:translog} shows that the same identification
source (conditional independence) yields nonlinear moment conditions
under translog production. The empirical implementation focuses on
Cobb--Douglas to maintain computational tractability and to isolate
the effect of relaxing the Markov assumption from functional form
complexities.
\end{remark}
\begin{remark}[Robustness to Endogenous Exit]
\label{rem:selection}
Standard proxy variable estimators require a survival probability
correction \parencite{olley1996thedynamics} because the innovation
shock $\xi_{jt}$ in the Markov transition equation
$\omega_{jt} = g(\omega_{j,t-1}) + \xi_{jt}$ is left-truncated
conditional on survival: firms with $\omega_{jt}$ below the exit
threshold do not appear in the data, biasing $\mathbb{E}[\xi_{jt}
\mid \omega_{j,t-1}, S_{jt}=1]$ away from zero.
The proposed estimator does not use the transition equation and
therefore does not involve $\xi_{jt}$. Identification of
$(\beta_m, \beta_e, \beta_w)$ rests on the within-period
conditional independence of demand shocks
(Assumption~\ref{ass:cond_indep}), which conditions on $\omega_{jt}$.
Under the standard timing convention that exit decisions are made
at the start of period $t$ based on the state $(\omega_{jt}, k_{jt})$
before input-specific demand shocks $(\tau_{jt}, \nu_{jt}, \eta_{jt})$
are realized, survival is a deterministic function of
$(\omega_{jt}, k_{jt})$. Conditioning on $\omega_{jt}$ therefore
absorbs the selection:
\[
f(\tau, \nu, \eta \mid \omega, x, S=1)
= f(\tau, \nu, \eta \mid \omega, x),
\]
and the moment conditions that identify $(\beta_m, \beta_e, \beta_w)$
hold on the surviving population without any survival probability
correction. No assumption on the productivity process is required for
this result; it follows from the static, $\omega$-conditional
structure of the identification strategy.
Two qualifications apply. First, the recovered distribution of
$\omega_{jt}$ is the survivor distribution, not the population
distribution; aggregate productivity statistics based on the
recovered $\hat{\omega}_{jt}$ reflect surviving firms only. Second,
the argument does not extend to parameters identified from the
transition equation (e.g., the persistence of productivity), which
the proposed method does not estimate.
\end{remark}
\section{Estimation Methods}
\label{sec:estimation}
The nonparametric identification results of
Section~\ref{sec:model_identification} establish that the production
function and productivity distribution are identified from the joint
density of intermediate inputs; nonparametric sieve estimation could
in principle implement this directly, but the high-dimensional
numerical integration required is computationally prohibitive for
census-scale panels spanning hundreds of industries.
I therefore develop a GMM estimator that specializes to a linear
production function and linear demand functions. Under this
parametric restriction, the observational equivalence class of
Theorem~\ref{thm:obs_equiv} reduces to a two-dimensional
indeterminacy $(c_k, c_l)$ (equation~\eqref{eq:linear_ambiguity}),
and the identification results of Corollary~\ref{thm:exclusion}
and Theorem~\ref{thm:homothetic_id} carry through directly.
\subsection{Estimation Based on the Generalized Method of Moments}
\label{sec:gmm_approach}
As noted in Remark~\ref{rem:functional_form}, the identification
results of Section~\ref{sec:model_identification} apply to general
production functions. The parametric implementation below specializes
to the Cobb--Douglas case, where input demand functions are linear in
productivity (Appendix~\ref{sec:micro_fnd}). This linearity yields the
tractable linear GMM system of Blocks~A--B. Extension to flexible
functional forms such as translog is developed in
Appendix~\ref{app:translog}; the identification source remains the
conditional independence of demand shocks.
\subsubsection{Overview}
The GMM estimator jointly recovers the production function and demand
parameters from three blocks of moment conditions:
\begin{enumerate}
\item[(i)] \textbf{Block~A} (Proxy moments): orthogonality conditions
derived from eliminating $\omega_{jt}$ across pairs of demand
residuals and the production residual, using an asymmetric
instrument strategy;
\item[(ii)] \textbf{Block~B} (Covariance moments): cross-covariance
restrictions among demand and production residuals, exploiting the
mutual independence of demand shocks;
\item[(iii)] \textbf{Block~C} (Curvature moments): conditional
moment restrictions derived from the homothetic regularity
condition on $\mathbb{E}[\omega_{jt} \mid k_{jt}, l_{jt}]$
(Assumption~\ref{ass:homothetic}), which closes the $\Delta(k,l)$
identification gap characterized in Theorem~\ref{thm:obs_equiv}.
\end{enumerate}
Blocks~A and~B identify the intermediate input elasticities
$(\beta_m, \beta_e, \beta_w)$, the demand function
parameters~$(\theta_g, \psi_\omega)$, and certain composite
functions of $(\beta_k, \beta_l)$ and the demand slopes. However, as
shown in Section~\ref{sec:obs_equiv}, these blocks alone cannot
separate $\beta_k$ and $\beta_l$ from the demand function slopes on
$(k, l)$ due to the $\Delta(k,l)$ observational equivalence
(Theorem~\ref{thm:obs_equiv}). Block~C resolves this indeterminacy through the nonlinear curvature
of $\mathbb{E}[\omega \mid k, l]$ imposed by
Assumption~\ref{ass:homothetic}, thereby achieving point
identification of all structural parameters
(Theorem~\ref{thm:homothetic_id}). When its identifying conditions
are weak, the exclusion restriction of
Corollary~\ref{thm:exclusion} provides an alternative route.
Figure~\ref{fig:flowchart} (Appendix~\ref{app:flowchart}) provides
a visual overview of the full estimation and inference pipeline,
including the diagnostic branches that determine which identification
route applies.
\subsubsection{Model Specification and Parameters}
\label{sec:gmm_spec}
The parametric specialization below implements the identification
results of Section~\ref{sec:model_identification} under additive
separability; this restriction reduces the nonparametric problem
to a finite-dimensional GMM system while preserving all theoretical
properties of Theorems~\ref{thm:obs_equiv}--\ref{thm:homothetic_id}.
To apply GMM, I impose additive separability on both the production
and demand functions.
\paragraph{Production function.}
Following the parametric model of
Section~\ref{sec:homothetic}, the production function is specified
as:
\begin{equation}
y_{jt} = \beta_k k_{jt} + \beta_l l_{jt}
+ \beta_m m_{jt} + \beta_e e_{jt} + \beta_w w_{jt}
+ \omega_{jt} + \varepsilon_{jt}.
\label{eq:gmm_prod}
\end{equation}
Here $g(k,l;\theta) = \beta_k k + \beta_l l$ is the parametric
$(k,l)$ component and
$q(m,e,w) = \beta_m m + \beta_e e + \beta_w w$ is the
(linear) intermediate input component, corresponding to the
additively separable model~\eqref{eq:parametric_model}.
\paragraph{Demand functions.}
The intermediate input demands take the additively separable form:
\begin{align}
m_{jt} &= \gamma_k\, k_{jt} + \gamma_l\, l_{jt} + h_m(z_{jt})
+ \gamma_\omega\, \omega_{jt} + \tau_{jt},
\label{eq:gmm_demand_m} \\
e_{jt} &= \delta_k\, k_{jt} + \delta_l\, l_{jt} + h_e(z_{jt})
+ \delta_\omega\, \omega_{jt} + \nu_{jt},
\label{eq:gmm_demand_mp} \\
w_{jt} &= \zeta_k\, k_{jt} + \zeta_l\, l_{jt} + h_w(z_{jt})
+ \zeta_\omega\, \omega_{jt} + \eta_{jt},
\label{eq:gmm_demand_mpp}
\end{align}
where the functions $h_m, h_e, h_w$ are left unrestricted and
$\psi_\omega = (\gamma_\omega, \delta_\omega, \zeta_\omega)$ are the
productivity loading coefficients.
\footnote{The Cobb--Douglas first-order condition
(Appendix~\ref{sec:micro_fnd}) structurally constrains the demand
function to be linear in $(k, l, \omega)$, but imposes no restriction
on the functional form of the dependence on $z$. The state variables
$z_{jt}$ enter through input prices $\ln P_{h,jt}$, the common
market factor $\ln(P_{jt}/\mu_{jt})$, markdowns $\ln \psi_{h,jt}$,
and wedges $\ln \Upsilon_{h,jt}$
(equation~\eqref{eq:structural_decomposition}), each of which may
depend nonlinearly on $z$.}
The demand slope parameters
$\theta_g = (\gamma_k, \gamma_l, \delta_k, \delta_l, \zeta_k,
\zeta_l)$ and the productivity loadings $\psi_\omega$ are estimated
jointly by GMM together with the $3\,d_z$ coefficients of
$h_m, h_e, h_w$ on the polynomial basis in $z$.
\paragraph{Homothetic structure of
$\mathbb{E}[\omega \mid k, l]$.}
Under Assumption~\ref{ass:homothetic} (Homothetic Weak
Separability), the conditional expectation of productivity admits the
representation
$\mathbb{E}[\omega_{jt} \mid k_{jt}, l_{jt}] = h(v(k_{jt}, l_{jt}))$.
The economic motivation is discussed in Section~\ref{sec:homothetic}.
I parametrize the index function using a CES aggregator:
\begin{equation}
v_{jt}(\alpha, \rho_v) = \frac{1}{\rho_v}\,\log\!\bigl(
\alpha\, e^{\rho_v\, k_{jt}} + (1 - \alpha)\, e^{\rho_v\, l_{jt}}
\bigr),
\label{eq:index_v}
\end{equation}
which, in levels, corresponds to the CES aggregator
$V = \bigl(\alpha\, K^{\rho_v} + (1 - \alpha)\, L^{\rho_v}\bigr)^{1/\rho_v}$.
This nests the Cobb--Douglas case ($\rho_v \to 0$, where $v \to \alpha\,k + (1-\alpha)\,l$)
as a special case and satisfies the degree-one
homogeneity requirement (Assumption~\ref{ass:homothetic}(B)) and the
strict convexity of isoquants
(Assumption~\ref{ass:homothetic}(C)) for $\alpha \in (0,1)$ and any $\rho_v \neq 0$.
The transformation function $h$ is approximated by a cubic polynomial:
\begin{equation}
h(v;\, \rho) = \rho_1\, v + \rho_2\, v^2 + \rho_3\, v^3,
\label{eq:h_poly}
\end{equation}
where the constant $\rho_0$ is absorbed by de-meaning prior to estimation. Under the normalization
$\mathbb{E}[\omega] = 0$, the constant satisfies
$\rho_0 = -\mathbb{E}[\rho_1 v + \rho_2 v^2 + \rho_3 v^3]$; this
constant is not separately identified from the production function
intercept and is recovered post-estimation. Condition~(A) of
Assumption~\ref{ass:homothetic} ($h'$ non-constant) requires
$\rho_2 \neq 0$ or $\rho_3 \neq 0$; this is a necessary condition
for the identification of $\beta_k$ and $\beta_l$
(Theorem~\ref{thm:homothetic_id}). I report results for polynomial orders 3 through 5 as a
robustness check; computational details including the
parametrization of $h$ are in Appendix~\ref{app:computation}.
\paragraph{Parameter classification.}
The full parameter vector is
$\Theta = (\theta_1', \theta_2')'$, where:
\begin{align*}
\theta_1 &= (\beta_m,\, \beta_e,\, \beta_w,\,
\theta_g,\, \psi_\omega)
&&\text{(intermediate input and demand parameters),} \\
\theta_2 &= (\beta_k,\, \beta_l,\,
\alpha,\, \rho_1,\, \rho_2,\, \rho_3)
&&\text{(primary input and homothetic parameters).}
\end{align*}
\paragraph{Residuals.}
Define the observable residuals, where the nuisance functions
$h_m(z), h_e(z), h_w(z)$ are estimated jointly as described below:
\begin{align}
\tilde{m}_{jt} &\equiv m_{jt} - \gamma_k\, k_{jt}
- \gamma_l\, l_{jt} - h_m(z_{jt})
= \gamma_\omega\,\omega_{jt} + \tau_{jt},
\label{eq:resid_m} \\
\tilde{e}_{jt} &\equiv e_{jt} - \delta_k\, k_{jt}
- \delta_l\, l_{jt} - h_e(z_{jt})
= \delta_\omega\,\omega_{jt} + \nu_{jt},
\label{eq:resid_mp} \\
\tilde{w}_{jt} &\equiv w_{jt} - \zeta_k\, k_{jt}
- \zeta_l\, l_{jt} - h_w(z_{jt})
= \zeta_\omega\,\omega_{jt} + \eta_{jt},
\label{eq:resid_mpp} \\
\tilde{y}_{jt} &\equiv y_{jt} - \beta_m m_{jt}
- \beta_e e_{jt} - \beta_w w_{jt}
= \beta_k k_{jt} + \beta_l l_{jt}
+ \omega_{jt} + \varepsilon_{jt}.
\label{eq:resid_y}
\end{align}
The equalities following the definition signs hold at the true
parameter values. The nuisance functions $h_m(z), h_e(z), h_w(z)$
are approximated by second-degree polynomials in $z$ and estimated
jointly with the structural parameters; details are in
Appendix~\ref{app:demeaning_details}.
\subsubsection{Moment Conditions}
\label{sec:gmm_moments}
Under Block~A+B estimation, the normalization $\beta_k = \beta_l = 0$
is adopted; this is without loss of generality because the
$\Delta(k,l)$ observational equivalence
(Theorem~\ref{thm:obs_equiv}) implies that $\beta_k$ and $\beta_l$
are not separately identified from the demand function slopes on
$(k, l)$ without Block~C. Under this normalization,
$\tilde{y}_{jt} = \omega_{jt} + \varepsilon_{jt}$.
\paragraph{Block~A: Proxy Moments.}
\footnote{The moment conditions require
Assumption~\ref{ass:gmm_uncorrelated}
(Appendix~\ref{app:identification_details}), which is implied by
the zero conditional mean condition together with
Assumption~\ref{ass:cond_indep}.}
By eliminating $\omega_{jt}$ across pairs of residuals, I
construct three error terms that depend only on the structural
shocks:
\begin{align}
u_{1,jt} &\equiv \delta_\omega\,\tilde{m}_{jt}
- \gamma_\omega\,\tilde{e}_{jt}
= \delta_\omega\,\tau_{jt}
- \gamma_\omega\,\nu_{jt},
\label{eq:e1} \\
u_{2,jt} &\equiv \zeta_\omega\,\tilde{m}_{jt}
- \gamma_\omega\,\tilde{w}_{jt}
= \zeta_\omega\,\tau_{jt}
- \gamma_\omega\,\eta_{jt},
\label{eq:e2} \\
u_{3,jt} &\equiv \gamma_\omega\,\tilde{y}_{jt}
- \tilde{m}_{jt}
= \gamma_\omega\,\varepsilon_{jt} - \tau_{jt}.
\label{eq:e3}
\end{align}
An asymmetric instrument strategy assigns different instruments to
each error based on the shock composition. Since $u_{i,jt}$ excludes
certain shocks, the corresponding intermediate inputs serve as valid
additional instruments
(Appendix~\ref{app:block_a_details}):
\begin{align}
\mathbb{E}\bigl[(Z_{\mathrm{base}},\, w_{jt})
\otimes u_{1,jt}(\Theta)\bigr] &= \mathbf{0},
\label{eq:momentA1} \\
\mathbb{E}\bigl[(Z_{\mathrm{base}},\, e_{jt})
\otimes u_{2,jt}(\Theta)\bigr] &= \mathbf{0},
\label{eq:momentA2} \\
\mathbb{E}\bigl[(Z_{\mathrm{base}},\, e_{jt},\, w_{jt})
\otimes u_{3,jt}(\Theta)\bigr] &= \mathbf{0},
\label{eq:momentA3}
\end{align}
where $Z_{\mathrm{base},jt} = (k_{jt},\, l_{jt},\, z_{jt})$.
Block~A is invariant to the $\Delta(k,l)$
transformation of Theorem~\ref{thm:obs_equiv} and therefore cannot
separately identify $\beta_k$ from the demand slopes on $(k, l)$
(Appendix~\ref{app:block_a_details}).
\paragraph{Block~B: Covariance Moments.}
Let $\phi_h$ denote the productivity loading of residual $\tilde{h}$:
$\phi_m \equiv \gamma_\omega$, $\phi_e \equiv \delta_\omega$,
$\phi_w \equiv \zeta_\omega$, and $\phi_y \equiv 1$.
The mutual exogeneity of shocks
(Assumption~\ref{ass:gmm_uncorrelated}(3)) implies that
$\mathrm{Cov}(\tilde{h}_1, \tilde{h}_2)
= \phi_{h_1}\,\phi_{h_2}\,\mathrm{Var}(\omega)$ for
each pair $(h_1, h_2) \in \{y, m, e, w\}$, $h_1 \neq h_2$.
Eliminating $\mathrm{Var}(\omega)$ across the six distinct pairs
yields six covariance relations of the form
\begin{equation}
\mathbb{E}\bigl[\tilde{h}_1\,\tilde{h}_2
- \phi_{h_2}\,\tilde{y}\,\tilde{h}_1\bigr] = 0,
\label{eq:momentB_compact}
\end{equation}
for each pair (Appendix~\ref{app:block_b_derivation} lists the
individual conditions). Of these six relations, four are
algebraically implied by the Block~A instrumental variable
moments: the conditions involving cross-products of the demand
residuals $\tilde{e}$ and $\tilde{w}$ with the proxy equation
errors are already encoded in the Block~A moment conditions
through the instruments $Z_3 = (k, l, \tilde{e}, \tilde{w})$.
Consequently, Block~B contributes only two independent moment
conditions beyond Block~A, and the combined Block~A+B system is
just-identified. The concentrated covariance-ratio formulas
derived in Appendix~\ref{app:block_b_derivation} remain useful
for obtaining closed-form scale parameter estimates,
improving computational efficiency.
As with Block~A, Block~B is invariant to the
$\Delta(k,l)$ transformation.
\paragraph{Block~C: Curvature Moments.}
\label{sec:blockC}
Block~C resolves the $\Delta(k,l)$ indeterminacy by implementing
the homothetic regularity condition
(Assumption~\ref{ass:homothetic},
Theorem~\ref{thm:homothetic_id}).
Define the net output residual
$\tilde{y}_{jt}(\theta_1) = y_{jt} - \beta_m m_{jt}
- \beta_e e_{jt} - \beta_w w_{jt}$. Evaluating at the
true parameter vector $\Theta_0$, the production
function~\eqref{eq:gmm_prod} gives
$\tilde{y}_{jt} = \beta_k k_{jt} + \beta_l l_{jt} + \omega_{jt}
+ \varepsilon_{jt}$. Taking the conditional expectation with
respect to $(k_{jt}, l_{jt})$:
\begin{equation}
\mathbb{E}[\tilde{y}_{jt} \mid k, l]
= \beta_k\, k + \beta_l\, l + h(v(k,l)).
\label{eq:cond_exp_ytilde}
\end{equation}
The first step uses
$\mathbb{E}[\varepsilon_{jt} \mid k, l] = 0$, which follows from
Assumption~\ref{ass:additive_error} by the law of iterated
expectations. The second step uses
Assumption~\ref{ass:homothetic}. No structural
decomposition of $\omega_{jt}$ is postulated;
equation~\eqref{eq:cond_exp_ytilde} follows entirely from the
definition of conditional expectation and the regularity condition
on its functional form.
Define the structural error:
\begin{equation}
u_{jt}(\Theta) \equiv \tilde{y}_{jt}(\theta_1)
- \beta_k\, k_{jt} - \beta_l\, l_{jt}
- h\bigl(v(k_{jt}, l_{jt};\, \alpha);\, \rho\bigr).
\label{eq:u_structural}
\end{equation}
Equation~\eqref{eq:cond_exp_ytilde} implies
$\mathbb{E}[u_{jt} \mid k_{jt}, l_{jt}] = 0$ at $\Theta_0$, which
yields valid moment conditions with any function of $(k, l)$ as
instruments. I use the polynomial instrument vector:
\begin{equation}
Z_{2,jt} = \bigl(k_{jt},\; l_{jt},\;
k_{jt}^2,\; l_{jt}^2,\; k_{jt}\,l_{jt},\;
k_{jt}^3,\; l_{jt}^3,\; k_{jt}^2\,l_{jt},\;
k_{jt}\,l_{jt}^2 \bigr)',
\label{eq:Z2}
\end{equation}
giving the moment conditions:
\begin{equation}
\mathbb{E}\bigl[Z_{2,jt} \cdot u_{jt}(\Theta)\bigr]
= \mathbf{0}.
\label{eq:momentC}
\end{equation}
As with Block~A, the constant term is excluded from $Z_{2,jt}$
and $\rho_0$ (the intercept of $h$) is recovered post-estimation
from the de-meaned residuals.
\paragraph{Identification mechanism.}
The structural error $u_{jt}$ depends on
$\theta_2 = (\beta_k, \beta_l, \alpha, \rho_1, \rho_2, \rho_3)$.
Theorem~\ref{thm:homothetic_id} establishes that under
Assumption~\ref{ass:homothetic}, the $\Delta(k,l) = c_k k + c_l l$
transformation is incompatible with the homothetic structure unless
$(c_k, c_l) = (0,0)$. Operationally, this identification works
through the higher-order instruments in $Z_{2,jt}$: the nonlinear
terms $v^2$ and $v^3$ in $h$ interact with the homogeneity of $v$
in a manner that uniquely pins down $\beta_k$ and $\beta_l$.
If $\rho_2 = \rho_3 = 0$ (i.e., $h$ is linear), then $\beta_k$ and
$\rho_1 \alpha$ are linearly confounded and identification fails.
The significance of $\hat{\rho}_2$ and/or $\hat{\rho}_3$ therefore
serves as a diagnostic for the strength of identification. I report
estimates and standard errors of these parameters in both the
simulation and the empirical analysis.\footnote{
In practice, even when $\rho_2$ and $\rho_3$ are nonzero,
the near-collinearity between $\rho_1 v(k,l)$ and
$(\beta_k k, \beta_l l)$ can impede numerical optimization.
I orthogonalize the polynomial basis $(v, v^2, v^3)$
against the linear span of $(1, k, l)$ before constructing
$h$, so that only the nonlinear component of $h(v)$ (the
source of identification,
Theorem~\ref{thm:homothetic_id}) enters the Block~C
moment conditions. This is a reparametrization: the
structural parameters $(\beta_k, \beta_l, \alpha)$ are
invariant, while the polynomial coefficients
$(\rho_1, \rho_2, \rho_3)$ are redefined as loadings on
the orthogonalized basis.}
\subsubsection{De-Meaning, Intercepts, and Estimation Procedure}
\label{sec:gmm_procedure}
\paragraph{De-meaning and estimation procedure.}
All variables are de-meaned prior to estimation and the constant
is excluded from all instrument vectors. All parameters
$\Theta = (\theta_1, \theta_2)$ are estimated simultaneously by
two-step GMM:
\begin{equation}
\hat{\Theta} = \arg\min_\Theta\;
g_N(\Theta)'\,\hat{W}\,g_N(\Theta),
\qquad
g_N(\Theta) = \frac{1}{N}\sum_{j=1}^N
\bar{g}_j(\Theta),
\label{eq:gmm_objective}
\end{equation}
where $\bar{g}_j(\Theta) = T^{-1}\sum_{t=1}^T g_{jt}(\Theta)$
stacks all moment conditions, and $\hat{W}$ is the optimal
weighting matrix estimated from a first-step identity-weighted
GMM. Post-estimation intercepts and further implementation
details are in Appendix~\ref{app:demeaning_details}.
\subsubsection{Recovering Productivity}
\label{sec:omega_recovery_gmm}
Given the estimated parameters $\hat{\Theta}$, the firm-level
productivity measure is computed as:
\begin{equation}
\hat{\omega}_{jt} = y_{jt} - \hat{\beta}_k\, k_{jt}
- \hat{\beta}_l\, l_{jt} - \hat{\beta}_m\, m_{jt}
- \hat{\beta}_{e}\, e_{jt}
- \hat{\beta}_{w}\, w_{jt}.
\label{eq:omega_recovery}
\end{equation}
If $\varepsilon_{jt} = 0$, this equals $\omega_{jt}$. Otherwise,
$\hat{\omega}_{jt} = \omega_{jt} + \varepsilon_{jt}$; the ex-post
shock acts as classical measurement error when $\hat{\omega}_{jt}$
is used in subsequent regressions
(Theorem~\ref{thm:asymptotics}).
\paragraph{Practical treatment of the $\Delta(k, l)$ indeterminacy.}
When Block~C is not imposed, $\hat{\omega}_{jt}$ includes a location
shift $c(k_{jt}, l_{jt})$ (Theorem~\ref{thm:obs_equiv}). Since $c$
depends only on $(k, l)$, flexible controls in $(k, l)$ absorb this
shift in regression analysis; in difference-in-differences designs
with parallel $(k, l)$ trends, $c$ is automatically differenced out.
When the identifying restrictions of Section~\ref{sec:closing_gap}
are imposed, $\Delta(k, l)$ reduces to a constant absorbed by fixed
effects. The proposed estimator therefore supports event studies and
productivity regressions without requiring Block~C: the
$\Delta(k,l)$ component is controlled via polynomial $(k,l)$
regressors in all subsequent regressions
(Section~\ref{sec:empirical}).
The formal justification is provided by
Proposition~\ref{prop:omega_hat_D} and the ATT identification
result in Appendix~\ref{app:proof_omega_hat_D}.
\subsubsection{Asymptotic Properties}
\label{sec:gmm_asymptotics_revised}
Under standard regularity conditions
(Appendix~\ref{app:asymptotic_proof}), the GMM estimator satisfies:
\begin{thm}[Asymptotic Properties of the GMM Estimator]
\label{thm:asymptotics}
As $N \to \infty$ with $T$ fixed:
(a)~$\hat{\Theta} \xrightarrow{p} \Theta_0$; and
(b)~$\sqrt{N}\,(\hat{\Theta} - \Theta_0)
\xrightarrow{d} N(0, V)$, where
\begin{equation}
V = (G'WG)^{-1}\, G'W\Sigma WG\, (G'WG)^{-1}
\label{eq:asymptotic_variance}
\end{equation}
with $\Sigma = \mathbb{E}[\bar{g}_j(\Theta_0)\,
\bar{g}_j(\Theta_0)']$ and
$G = \mathbb{E}[\nabla_\Theta \bar{g}_j(\Theta_0)]$.
\end{thm}
Standard errors are clustered at the firm level to accommodate
arbitrary within-firm serial dependence. The proof and regularity
conditions are in Appendix~\ref{app:asymptotic_proof}.
Computational details are in Appendix~\ref{app:computation}.
\subsubsection{Specification Testing and Diagnostics}
\label{sec:diagnostics}
\paragraph{Identification count.}
The combined Block~A+B system is just-identified: Block~A
contributes 10 moment conditions, and Block~B contributes exactly
two independent moment conditions beyond Block~A (four of the six
Block~B covariance relations are algebraically redundant with Block~A;
Section~\ref{sec:gmm_approach}), giving 12 moment conditions matching
the 12 free parameters in $\theta_1$. The scale parameters
$(\gamma_\omega, \delta_\omega, \zeta_\omega)$ are estimated via
closed-form covariance ratios for computational efficiency.
\paragraph{Strength of identification for $\beta_k, \beta_l$.}
As discussed in Section~\ref{sec:blockC}, the identification of
$\beta_k$ and $\beta_l$ relies on the nonlinearity of $h$
($\rho_2 \neq 0$ or $\rho_3 \neq 0$). I report the estimates and
$t$-statistics of $\hat{\rho}_2$ and $\hat{\rho}_3$ as diagnostics.
If both are insignificant, the identification of primary input
elasticities may be weak, and the researcher should interpret
$\beta_k$ and $\beta_l$ with caution or consider exclusion
restrictions (Corollary~\ref{thm:exclusion}) as an alternative
identification strategy.
\paragraph{Reduced-form check of Assumption~\ref{ass:homothetic}.}
As a pre-estimation diagnostic, one may estimate $\theta_1$ from
Blocks~A and~B alone (which does not require
Assumption~\ref{ass:homothetic}), construct
$\tilde{y}_{jt}(\hat{\theta}_1)$, and examine whether
$\mathbb{E}[\tilde{y} \mid k, l]$ exhibits a homothetic structure
via nonparametric regression. A visual departure from homotheticity
would indicate a violation of the identifying assumption.
\paragraph{Polynomial degree selection.}
The cubic specification of $h$ can be extended to higher-order
polynomials. I recommend reporting results for polynomial
orders~3 through~5 and selecting via information criteria.
Appendix~\ref{subsec:blockc_results} reports the full-sample
Block~C recovery results for $(\beta_k, \beta_l)$ across all
502 industries, comparing the homothetic approach with the exclusion
restriction and ACF estimators.
\section{Monte Carlo Simulation}
\label{sec:simulation}
The identification results in
Section~\ref{sec:model_identification} show that the Markov
assumption is unnecessary; this section asks whether removing it
matters quantitatively. I use Monte Carlo simulations to measure
the bias that the Markov assumption introduces in the materials
elasticity and to trace its propagation into downstream objects.
The primary comparison is between the proposed estimator, which
imposes no restriction on productivity dynamics, and the standard
ACF estimator, which requires first-order Markov.
\subsection{Data Generating Process (DGP)}
All DGPs share a common structure for the production function, demand
functions, and dynamic input decisions, differing only in the
productivity process. Detailed parameter settings are in
Appendix~\ref{app:dgp}.
\subsubsection{Basic Structure}
The firm's production function is Cobb--Douglas in all inputs:\footnote{The
Cobb--Douglas specification is standard in Monte Carlo studies of
production function estimators
\parencite{ackerberg2015identification,gandhi2020onthe}. Evaluating
the proposed method under more flexible production functions (e.g.,
translog) is left for future work; the identification results
(Theorems~\ref{thm:density_id}--\ref{thm:homothetic_id}) do not
require Cobb--Douglas.}
\begin{equation}
y_{jt} = \beta_0 + \beta_k k_{jt} + \beta_l l_{jt}
+ \beta_m m_{jt} + \beta_e e_{jt}
+ \beta_w w_{jt} + \omega_{jt} + \varepsilon_{jt},
\label{eq:sim_prod_func}
\end{equation}
with true parameter values
$(\beta_0, \beta_k, \beta_l, \beta_m, \beta_e, \beta_w)
= (0.1, 0.2, 0.3, 0.3, 0.15, 0.1)$ and
$\varepsilon_{jt} \sim \text{i.i.d.}\ N(0, 0.05^2)$.
Intermediate input demands are log-linear in $(k, l, \omega)$ with
input-specific demand shocks following independent AR(1) processes
($\rho = 0.5$, $\sigma = 0.15$). The demand function coefficients
are calibrated from the first-order conditions of cost minimization
under input-specific markdowns (Appendix~\ref{sec:micro_fnd});
the productivity loading coefficients
$(\gamma_\omega, \delta_\omega, \zeta_\omega) = (2.2, 2.0, 1.8)$
differ across inputs, reflecting heterogeneous
markdowns.\footnote{Under perfect competition with a Cobb--Douglas
production function, the first-order condition implies
$\gamma_\omega = 1/(1 - \beta_m) \approx 1.43$; the larger values
incorporate input-specific markdowns and procurement frictions
(see Appendix~\ref{sec:micro_fnd}).} Conditional independence
(Assumption~\ref{ass:cond_indep}) is a cross-sectional condition
requiring mutual independence across inputs at each point in time;
it is unaffected by the serial correlation of individual shocks,
since each AR(1) has mutually independent innovations.
Primary inputs are endogenously determined. Capital accumulates
through dynamic investment, and labor is chosen based on forecasted productivity from an AR(1)
model. The labor demand function
is structured so that Assumption~\ref{ass:homothetic} holds: the
conditional expectation $\mathbb{E}[\omega_{jt} \mid k_{jt},
l_{jt}]$ is a function of a CES aggregator with
$(\alpha, \rho_v) = (0.4, 0.3)$. Full parameter details are provided
in Appendix~\ref{app:dgp}.
To test the robustness of the proposed method, I generate
productivity under three scenarios:
\begin{enumerate}
\item \textbf{DGP1: AR(1) Markov Process (Baseline).} The standard
case where existing methods are correctly specified.
$\omega_{jt} = 0.8\,\omega_{j,t-1} + \xi_{jt}$, with
$\sigma_\xi = 0.2$.
\item \textbf{DGP2: AR(2) Process.} Productivity depends on its own
two-period history, as when R\&D investments require two years to
affect efficiency. The first-order Markov assumption is violated.
$\omega_{jt} = 0.6\,\omega_{j,t-1} + 0.3\,\omega_{j,t-2}
+ \xi_{jt}$, $\sigma_\xi = 0.15$.
\item \textbf{DGP3: Potential Outcome Model.} A firm's realized
productivity is determined by an endogenous binary treatment
$D_{jt}$, generating a potential outcome process incompatible
with the first-order Markov assumption.
Following the diagonal reference model of
\textcite{chen2024identifying}, untreated productivity follows
$\omega^0_{jt} = 0.8\,\omega^0_{j,t-1} + \xi_{0,jt}$
($\sigma_{\xi 0} = 0.2$) and treated productivity follows
$\omega^1_{jt} = 0.5\,\omega^1_{j,t-1} + 0.15 + \xi_{1,jt}$
($\sigma_{\xi 1} = 0.25$). Observed productivity is
$\omega_{jt} = (1-D_{jt})\omega^0_{jt} + D_{jt}\omega^1_{jt}$.
Treatment is reversible and endogenous:
$D_{jt} = \mathbb{I}(\omega^0_{jt} > 0)$, so firms enter and
exit treatment as their untreated potential productivity
fluctuates above and below zero.
Full parameter settings, including the capital accumulation
and labor decision rules common to all DGPs, are provided
in Appendix~\ref{app:dgp}.
\item \textbf{DGP4: Conditional Independence Violation.}
The productivity process is AR(1) as in DGP1, but the
electricity demand shock $\nu_{jt}$ and the water demand shock
$\eta_{jt}$ are correlated via a common factor:
$\nu_{jt} = \sqrt{1-\rho_{ew}^2}\,\sigma_\nu\,\epsilon_\nu
+ \rho_{ew}\,\sigma_\nu\,\epsilon_{\text{common}}$ and similarly for
$\eta_{jt}$, where $\epsilon_{\text{common}} \sim \mathcal{N}(0,1)$
is independent of $\omega_{jt}$; the materials shock $\tau_{jt}$
remains independent.
This generates $\mathrm{Corr}(\nu_{jt}, \eta_{jt}) = \rho_{ew} \in
\{0, 0.05, 0.10, 0.20, 0.30\}$, directly violating
Assumption~\ref{ass:cond_indep} when $\rho_{ew} > 0$.
A common energy price shock or seasonal supply constraint that
simultaneously raises both electricity and water costs is one
economic interpretation, arguably the most salient threat to
conditional independence, since both are utility services subject
to common regulatory and infrastructure conditions.
This DGP tests the robustness of the proposed method to
violations of the conditional independence assumption.
\end{enumerate}
\subsection{Estimation Methods Compared}
Using the generated data, I organize the estimation into two
parts to isolate the contributions of each block of moment conditions.
\paragraph{Part~1: Flexible input parameters.}
I estimate the intermediate input elasticities
$(\beta_m, \beta_e, \beta_w)$ and compare four estimators.\footnote{The main text figures report two of the four estimators (ACF and Proposed). ACF-Mod results are in Appendix~\ref{sec:mc_additional_results}; GNR results are also reported there.}
The four estimators are:
\begin{enumerate}
\item \textbf{Proposed (Block~A+B):} The GMM estimator of
Section~\ref{sec:gmm_approach}, using the intermediate input
moment conditions (Block~A) and the covariance moments
(Block~B). This part does not identify $\beta_k$ and $\beta_l$,
which remain subject to the $\Delta(k,l)$ indeterminacy
(Theorem~\ref{thm:obs_equiv}).
\item \textbf{Standard ACF:} The two-step GMM estimator of
\textcite{ackerberg2015identification}, assuming a first-order
Markov process for productivity.
\item \textbf{Modified ACF (ACF-Mod):} A variant of the ACF estimator
in which the demand shocks $(\tau_{jt}, \nu_{jt}, \eta_{jt})$ are
treated as observed and included as controls in the first stage.
This ensures scalar unobservability by construction. Any remaining
bias in ACF-Mod can therefore be attributed solely to the violation
of the Markov assumption, isolating the dynamic misspecification
channel.
\item \textbf{GNR:} The estimator of
\textcite{gandhi2020onthe}, implemented with a polynomial
share regression of degree~2 and a degree-3 polynomial for
$g(\omega_{t-1})$, with common prices ($P = r = 1$).
In the DGP, persistent input-specific demand shocks
($\rho = 0.5$) violate both the FOC premise and the
non-persistence condition of GNR (their Appendix~O6,
Assumption~7), so GNR tests the share regression approach
under persistent input market imperfections.
GNR is included in Part~1 only, as its second stage is
structurally identical to ACF.
\end{enumerate}
\paragraph{Part~2: Fixed input parameters.}
I additionally estimate $(\beta_k, \beta_l)$ by adding the
homothetic regularity condition (Block~C) to the proposed estimator:
\begin{enumerate}
\item \textbf{Proposed (Block~A+B+C):} The full GMM estimator using
all three blocks, with the CES aggregator
$v(k,l) = \frac{1}{\rho_v}\log\bigl(\alpha e^{\rho_v k}
+ (1-\alpha) e^{\rho_v l}\bigr)$ evaluated at the true DGP
values $(\rho_v, \alpha) = (0.3, 0.4)$.\footnote{These parameters
are known by construction in the simulation; the empirical
application treats them as unknown and estimates them by profile
GMM (Section~\ref{subsec:specification}).} The comparison with
Part~1 isolates the contribution of Block~C.
\item \textbf{Standard ACF} and \textbf{ACF-Mod:} Same as above,
now evaluated on $(\beta_k, \beta_l)$ as well.
\end{enumerate}
\subsection{Evaluation Metrics}
I report bias,
$\text{Bias}(\hat{\beta})=\mathbb{E}_{R}[\hat{\beta}^{(r)}]-\beta_{\text{true}}$,
and RMSE,
$\text{RMSE}(\hat{\beta})=\sqrt{\mathbb{E}_{R}[(\hat{\beta}^{(r)}-\beta_{\text{true}})^{2}]}$,
averaged over $R$ Monte Carlo repetitions.
\subsection{Simulation Execution}
For Part~1, I run $R = 100$ replications for each combination
of DGP and estimation method, varying the number of firms
$N \in \{50, 200, 500\}$ and the observation period
$T \in \{10, 20, 50\}$ to examine the impact of sample size.
For Part~2, I run $R = 100$ replications at $(N, T) = (200, 50)$.
The parameter estimates obtained in each repetition are collected,
and mean bias and RMSE are calculated for comparison. With $R = 100$,
the simulation standard error of the estimated bias is approximately
$\text{SD}/\sqrt{R}$; for the typical standard deviation of
$\hat{\beta}_m$ ($\approx 0.005$), this yields a simulation
uncertainty of $\approx 0.0005$, which is small relative to the
reported biases.
\subsection{Results}
I report the Part~1 results in Figures~\ref{fig:bias_convergence} and~\ref{fig:boxplot_part1}
and the Part~2 results in Figure~\ref{fig:blockc_comparison}.\footnote{Additional summary tables, including GNR results, are provided in Appendix~\ref{sec:mc_additional_results}. RMSE convergence plots are in Figure~\ref{fig:rmse_convergence_app}.}
These results confirm that the proposed estimator performs well under
all DGPs considered and illustrate the sensitivity of the ACF framework
to violations of the Markov assumption.
\paragraph{Remark on GNR.}
GNR shares the static identification strategy of the proposed
method (both recover $\beta_m$ from within-period variation
without a Markov assumption) but requires competitive input
markets with non-persistent demand shocks
(their Appendix~O6, Assumption~7). The present DGP, which
features persistent input-specific shocks ($\rho_\tau = 0.5$),
is therefore outside GNR's maintained assumptions by design:
the DGP is calibrated to the proposed method's setting, not
GNR's. Under GNR's own assumptions ($\tau = \nu = \eta = 0$),
the share regression recovers $\beta_m$ consistently regardless
of the productivity process. The simulation results for GNR
(Appendix~\ref{sec:mc_additional_results}) should accordingly
be read as illustrating the sensitivity of the FOC-based approach
to input market imperfections, not as a general performance
comparison.
\paragraph{DGP1 (AR(1) Baseline):}
Under DGP1, where the Markov assumption holds, all three estimators
(ACF, ACF-Mod, and Proposed) are consistent.
The bias for each method decays toward zero as $T$ increases
(Figure~\ref{fig:bias_convergence}).
The boxplots in Figure~\ref{fig:boxplot_part1} corroborate this finding;
the proposed method remains centered on the true values. The ACF
and ACF-Mod estimators show small positive finite-sample bias that
diminishes with sample size (see Appendix~\ref{sec:mc_additional_results}
for detailed tables). However, the proposed estimator exhibits
larger variance than the ACF estimator under DGP1, resulting in
higher RMSE when the Markov assumption is correctly specified
(Appendix Table~\ref{tab:mc_part1_dgp1ar1}). This is the
efficiency cost of the static approach: the proposed method
trades time-series information for robustness to dynamic
misspecification.
Under DGP2 and DGP3, this ranking reverses: ACF's bias dominates
its variance advantage, yielding larger mean squared error.
The static identification strategy is also the
only approach in this literature that permits event study and
difference-in-differences designs, where the treatment itself
violates the Markov assumption
(Section~\ref{subsec:event_study_body}).
\paragraph{DGP2 (AR(2)) and DGP3 (Potential Outcome):}
Under DGP2 and DGP3, where the first-order Markov assumption does
not hold, Figure~\ref{fig:bias_convergence} reveals a clear
divergence. The ACF estimator exhibits positive bias in
$\hat{\beta}_m$ that does not vanish with increasing $T$. Under
DGP2, this reflects standard omitted-variable inconsistency: the
AR(2) component of productivity persistence is not captured by the
first-order transition equation. Under DGP3, the issue is more
fundamental: the Markov transition equation is structurally
incompatible with the potential outcome process
\parencite{chen2024identifying}, so the ACF moment condition lacks
a structural interpretation and the resulting estimate does not
converge to the true $\beta_m$. An infeasible oracle benchmark
(ACF-Mod) that removes scalar unobservability by treating demand
shocks as observed shows comparable bias under both DGPs,
confirming that the source is Markov misspecification rather than
demand shock heterogeneity
(Appendix~\ref{sec:mc_additional_results}).
The proposed method, by contrast, exhibits negligible bias across these
specifications. The bias remains close to zero for all values of $T$
under both DGP2 and DGP3. Because the estimator relies solely on
static conditional independence, it remains invariant to the underlying
productivity dynamics.
The main text figures report results for $N = 500$; increasing $N$
reduces variance for all estimators but does not mitigate ACF's
asymptotic bias under DGP2 or DGP3
(Appendix~Figure~\ref{fig:bias_convergence_byN}), confirming that
the bias is asymptotic rather than finite-sample.
\paragraph{Block~A+B vs.\ Block~A+B+C (Part~2):}
Part~2 supplements Part~1 by adding Block~C to recover
$(\beta_k, \beta_l)$. I use the design point $(N, T) = (200, 50)$, which matches
the Part~1 baseline, to examine whether Block~C disturbs
the Block~A+B parameters.
Figure~\ref{fig:blockc_comparison} presents the results, where
Block~C is added to identify $(\beta_k, \beta_l)$. In the baseline
DGP1, the proposed method recovers both parameters with negligible bias.
Under DGP3, where ACF estimates of $\beta_k$ and $\beta_l$ collapse
toward zero (RMSE $\approx 0.20$--$0.30$), the proposed method
achieves substantially lower error
(RMSE $\approx 0.02$). The intermediate input
elasticities $(\beta_m, \beta_e, \beta_w)$ remain stable between
Part~1 and Part~2, confirming that the addition of Block~C moments
does not contaminate the well-identified flexible input parameters.
This stability shows in finite samples that the joint GMM
system does not transmit Block~C misspecification into the
flexible input estimates: the intermediate input elasticities are identified
by Blocks~A and~B alone
(Theorem~\ref{thm:density_id}, specialized to the
Cobb--Douglas parametric model of Section~\ref{sec:gmm_approach}),
so any misspecification in Block~C affects only $(\beta_k, \beta_l)$. Because markups depend solely on
$\beta_m$ (equation~\eqref{eq:markup_cd}), the primary
empirical application is insulated from Block~C specification.
\paragraph{DGP4 (Conditional Independence Violation):}
DGP4 examines the cost of violating Assumption~\ref{ass:cond_indep} by
introducing correlation between the electricity demand shock
$\nu_{jt}$ and the water demand shock $\eta_{jt}$, arguably the
most economically salient threat to conditional independence, since
both are utility services subject to common energy prices and
infrastructure constraints. The materials shock $\tau_{jt}$ remains
independent. The correlation
$\rho_{ew} \equiv \mathrm{Corr}(\nu_{jt}, \eta_{jt})$ varies from 0 to
0.30.
The bias mechanism operates through the scale parameter
$\zeta_\omega$. Positive $\mathrm{Cov}(\nu, \eta)$ inflates
$\mathrm{Cov}(\tilde{e}, \tilde{w})$, causing the concentrated
scale estimator $\hat{\zeta}_\omega
= \mathrm{Cov}(\tilde{e}, \tilde{w}) / \mathrm{Cov}(\tilde{y},
\tilde{e})$ to overestimate $\zeta_\omega$
(Appendix~\ref{app:ci_violation}). The overestimated
$\hat{\zeta}_\omega$ introduces a positive productivity component
into the Block~A residual $u_2 = \hat{\zeta}_\omega \tilde{m}
- \hat{\gamma}_\omega \tilde{w}$, which the GMM compensates by
\emph{increasing} $\hat{\beta}_m$, yielding an \emph{upward} bias.
Table~\ref{tab:mc_dgp4_ci}
(Appendix~Figure~\ref{fig:dgp4_bias_vs_rho}) reports the results.
When $\rho_{ew} = 0$, the proposed method is approximately unbiased.
As $\rho_{ew}$ increases, $\hat{\beta}_m$ exhibits increasing upward
bias. The magnitudes suggest that the estimator is robust to
moderate violations.
The bias direction is the \emph{same} as the Markov
misspecification bias documented in DGPs~2 and~3 for ACF: both push
$\hat{\beta}_m$ upward. Therefore, the empirical finding that the
proposed estimator yields \emph{lower} $\hat{\beta}_m$ than ACF
(Section~\ref{subsec:main_results}) cannot be attributed to CI
violation; it must reflect Markov misspecification bias in
ACF.\footnote{ACF uses only the materials demand proxy and does not
exploit cross-shock variation, so it is unaffected by
$\mathrm{Corr}(\nu, \eta)$.}
Table~\ref{tab:mc_summary} summarizes the bias properties across
DGPs 1--3. The proposed method is unbiased across all three
specifications, while ACF exhibits positive bias under Markov misspecification.
\begin{table}[htbp]
\centering
\caption{Monte Carlo Summary: Bias Properties across DGPs ($N=200$, $T=50$)}
\label{tab:mc_summary}
\small
\begin{tabular}{lccc}
\toprule
& DGP~1 (AR1) & DGP~2 (AR2) & DGP~3 (PO) \\
\midrule
Proposed & Unbiased ($+0.002$) & Unbiased ($-0.001$) & Unbiased ($+0.000$) \\
ACF & Unbiased ($+0.001$) & Biased ($+0.026$) & Biased ($+0.266$) \\
GNR & Biased ($+0.589$) & Biased ($+0.589$) & Biased ($+0.592$) \\
\bottomrule
\end{tabular}
\medskip
\footnotesize\raggedright \textit{Notes: ``Biased ($+$)'' indicates positive
asymptotic bias in $\hat{\beta}_m$ that does not diminish with
sample size. See Figures~\ref{fig:bias_convergence}--\ref{fig:blockc_comparison} for detailed convergence plots and
Tables~\ref{tab:mc_part1_dgp1ar1}--\ref{tab:mc_part2_dgp3potential} (Appendix~\ref{sec:mc_additional_results}) for
full RMSE and SD by method and DGP.}
\end{table}
Taken together, the Monte Carlo simulations confirm that the proposed
method recovers production function parameters without
imposing restrictions on the productivity process. In contrast, standard
methods exhibit substantial positive bias in $\hat{\beta}_m$ when the
assumed law of motion for productivity does not match the true data
generating process. The simulations establish that Markov misspecification generates a
detectable and economically meaningful bias. The empirical
application then examines whether these patterns hold in
Japanese manufacturing data.
\begin{figure}[htbp]
\centering
\includegraphics[width=\textwidth]{Figure/MC/main_bias_convergence}
\caption{Part~1: Mean Bias Convergence ($N = 500$)}
\label{fig:bias_convergence}
\small\raggedright \textit{Notes: Mean bias of $(\hat{\beta}_m, \hat{\beta}_e, \hat{\beta}_w)$ as a function of $T$ for three DGPs ($N = 500$, $R = 100$). Under DGP~1 (baseline AR(1)), both methods are approximately unbiased. Under DGPs~2 and~3, where the first-order Markov assumption is violated, ACF exhibits persistent bias while the proposed method remains centered at zero. Three-method comparison including ACF-Mod is in Figure~\ref{fig:bias_convergence_3methods}; GNR results are in Appendix~\ref{sec:mc_additional_results}.}
\end{figure}
\begin{figure}[htbp]
\centering
\includegraphics[width=\textwidth]{Figure/MC/main_boxplot_part1}
\caption{Part~1: Distribution of Estimates ($N=500$, $T=50$)}
\label{fig:boxplot_part1}
\small\raggedright \textit{Notes: Distribution of $(\hat{\beta}_m, \hat{\beta}_e, \hat{\beta}_w)$ across $R = 100$ replications for $N = 500$, $T = 50$. Dashed lines indicate true values. The proposed method remains centered on the true values across all DGPs. Under DGP~2 and DGP~3, ACF distributions are shifted rightward, consistent with the positive Markov misspecification bias. A four-method comparison including ACF-Mod and GNR is in Figure~\ref{fig:boxplot_all4_app}.}
\end{figure}
\begin{figure}[htbp]
\centering
\includegraphics[width=\textwidth]{Figure/MC/main_boxplot_part2_kl}
\caption{Part~2: Distribution of $\hat{\beta}_k$ and $\hat{\beta}_l$ ($N=200$, $T=50$)}
\label{fig:blockc_comparison}
\small\raggedright \textit{Notes: Distribution of $(\hat{\beta}_k, \hat{\beta}_l)$ from Block~A+B+C estimation ($R = 100$). The proposed method identifies $(\beta_k, \beta_l)$ with moderate accuracy across all DGPs. Under DGP~3, the proposed method achieves substantially lower RMSE ($\approx 0.02$). ACF estimates of $\beta_k$ collapse to near zero under DGP~3 (mean $\hat{\beta}_k \approx 0.003$, true value $0.20$).}
\end{figure}
\section{Empirical Analysis}
\label{sec:empirical}
The empirical analysis has two objectives: to test whether the
conditional independence framework produces economically plausible
estimates across the manufacturing sector, and to assess the
relative plausibility of the static and dynamic identifying
assumptions through the convergence diagnostic of
Remark~\ref{rem:testable_exclusion}. I estimate the production
function for all 502 manufacturing industries using Block~A+B,
reporting analytical standard errors.
A practical consequence of this block structure: the markup
estimates, productivity determinants, and convergence
diagnostics reported below require only Blocks~A and~B.
These results do not depend on the resolution of the
$\Delta(k,l)$ indeterminacy and are available for all 502
industries. Block~A+B+C is used for a subset of industries
where $(\beta_k, \beta_l)$ recovery is needed for
productivity level analysis.
The section is organized as follows.
Sections~\ref{subsec:data_framework}
and~\ref{subsec:specification} describe the data and
estimation specifications.
Section~\ref{subsec:spec_diagnostics} presents the
exclusion restriction diagnostic.
Section~\ref{subsec:main_results} reports production function
parameters and markup estimates.
Section~\ref{subsec:determinants} examines productivity determinants.
Section~\ref{subsec:event_study_body} presents an event study
application exploiting the non-Markov validity of the estimator.
Section~\ref{subsec:capital_labor} reports estimates of
$(\beta_k, \beta_l)$ from two independent identification routes.
\subsection{Data and Analytical Framework}
\label{subsec:data_framework}
I apply the proposed method to the Japanese Census of
Manufactures and the Economic Census for Business Activity.
I estimate the production function for all manufacturing
industries with at least 50 firm-year observations in the
extended panel (2003--2020), yielding Block~A+B estimates
for 502 industries covering 559{,}381 firm-year observations.
\footnote{The identification results of
Section~\ref{sec:model_identification} require only the joint
distribution of $(m_{jt}, e_{jt}, w_{jt}, k_{jt}, l_{jt})$ at a
single point in time; no assumption on the time-series dynamics
of $\omega_{jt}$ is needed (Appendix~\ref{subsec:annual_params}).
The panel dimension is exploited solely to improve estimation
efficiency by time-averaging the sample moment conditions,
$\bar{g}_j(\Theta) = T^{-1}\sum_{t} g_{jt}(\Theta)$,
which reduces finite-sample variance without affecting consistency.}
These estimates provide markup distributions and productivity
determinants at the level of the entire manufacturing sector.
Four industries (food processing [Bread, industry code~971],
paper products (Corrugated board boxes, code~1453),
chemicals (Plastic film, code~1821), and
machinery [Industrial robots, code~2694]) serve as representative
cases for the time-varying parameter analysis in
Appendix~\ref{subsec:annual_params}, covering major manufacturing
sectors (food, paper, chemicals, machinery).
Analytical standard errors from the GMM sandwich formula are
reported for both the proposed method and the ACF benchmark.
The core variables include the logarithm of real output, $y_{jt}$,
the logarithm of real capital stock, $k_{jt}$, and the logarithm
of labor input, $l_{jt}$.
I map the theoretical input triplet
to observable data as follows. I designate the real value of primary raw
materials as $m_{jt}$, the quantity of electricity as $e_{jt}$,
and the quantity of industrial water as $w_{jt}$. This selection
exploits the fact that industrial water and electricity prices are
typically regulated, limiting firm-specific bargaining. This institutional feature reduces the risk
of unobserved common price shocks inducing correlation between
$\nu_{jt}$ and $\eta_{jt}$, thereby supporting the validity of the
conditional independence assumption
($\tau_{jt} \perp \nu_{jt} \perp \eta_{jt} \mid
(\omega_{jt}, x_{jt})$). The principal remaining threat is
commodity price shocks that jointly affect raw materials costs
and electricity generation costs. Two features mitigate this
concern: (i)~industrial electricity prices exhibit less
high-frequency variation than raw materials procurement costs,
as the fuel cost adjustment mechanism smooths commodity price
pass-through on a quarterly basis; and (ii)~even if a residual
common utility shock induces positive
$\mathrm{Corr}(\nu_{jt}, \eta_{jt})$, the resulting bias in
$\hat{\beta}_m$ is \emph{upward} (the same direction as ACF's
Markov bias), so the empirical gap between methods cannot be
attributed to CI violation (Section~\ref{sec:simulation},
Appendix~\ref{app:ci_violation}).
I augment $x_{jt}$ with control variables $z_{jt}$ consisting of
beginning-of-period total inventory ($z_{1,jt}$), its square
($z_{2,jt} \equiv z_{1,jt}^2$), plant fixed effects, and year fixed
effects.
These controls directly implement the conditioning strategy of
Section~\ref{subsec:relax_indep}, where common shocks are absorbed
by $z_{jt}$ so that the residual shock terms $\tau_{jt}, \nu_{jt},
\eta_{jt}$ satisfy the conditional independence assumption.
Inventory proxies for unobserved product demand fluctuations
\parencite{kumar2019productivity}: a firm anticipating high demand
accumulates more stock in advance, so inventory captures the
common demand component that would otherwise enter all three input
demands simultaneously (Section~\ref{subsec:relax_indep}).
The quadratic term $z_{2,jt}$ accommodates a nonlinear relationship
between inventory and unobserved demand, consistent with the
structural decomposition in equation~\eqref{eq:structural_decomposition}
where demand-related terms enter input prices nonlinearly.
Year fixed effects absorb common input price shocks (e.g., energy
price movements) that affect all inputs simultaneously, as
discussed in Section~\ref{subsec:relax_indep}.
Plant fixed effects absorb time-invariant plant-level heterogeneity
in input prices and buyer-supplier relationships, capturing the
firm-attribute component of input market power noted in
Section~\ref{subsec:relax_indep}.
The use of fixed effects exploits the panel dimension for
efficiency and enriches the conditioning set for the
conditional independence assumption, but does not impose any
restriction on the time-series dynamics of $\omega_{jt}$.
Standard proxy variable estimators use the Markov transition
equation to address exit-driven selection
\parencite{olley1996thedynamics}. The proposed method does not
require this correction: because identification conditions on
$\omega_{jt}$, endogenous exit based on $(\omega_{jt}, k_{jt})$
is absorbed by the conditioning and the moment conditions hold
on the surviving population without a survival probability
correction (Remark~\ref{rem:selection}). Plant fixed effects
further reduce the influence of systematic level differences
across plants.
\subsection{Specification of Estimation Methods}
\label{subsec:specification}
I contrast the results of my approach with those obtained from the
standard ACF framework.
First, I implement the \textbf{proposed method} using the GMM
estimator derived in Section~\ref{sec:gmm_approach}. The estimator
jointly recovers the production function and demand parameters as
described in Section~\ref{sec:gmm_approach}. The CES aggregator
parameters $(\rho_v, \alpha)$ are selected via profile GMM: for a
grid of $(\rho_v, \alpha)$ values, the remaining parameters are
estimated by minimizing the GMM objective, and the pair yielding the
smallest $J$-statistic is selected.\footnote{Under strong
identification of $(\rho_v, \alpha)$, the profile GMM procedure
yields a $J$-statistic with the standard $\chi^2$ distribution
asymptotically \parencite{newey1994chapter}. When identification of
these parameters is weak, the minimum-$J$ selection may bias the
test toward under-rejection, making the test conservative. The
block bootstrap standard errors reported below account for the
uncertainty in $(\rho_v, \alpha)$ selection by re-running the
profile grid search within each bootstrap replication.}
The nuisance functions $h_m, h_e, h_w$ are approximated by
second-degree polynomials in $(z_{1,jt}, z_{2,jt})$, giving a
polynomial basis of dimension $d_z = 2$ and thus $\dim\Theta = 24$
where Block~A+B is just-identified (Section~\ref{sec:diagnostics}).
I report the estimates of
$\hat{\rho}_2$ and $\hat{\rho}_3$ as diagnostics for the strength of
identification of $\beta_k$ and $\beta_l$
(Section~\ref{sec:diagnostics}).
As a benchmark, I estimate the ACF two-step GMM with the same
control variables $z_{jt}$ to ensure comparability. Analytical
standard errors from the GMM sandwich formula are reported for
both methods.
\paragraph{Identifying assumptions in practice.}
The ACF framework requires scalar unobservability (productivity as
the sole unobservable in input demand) and a first-order Markov
process for productivity. GNR requires scalar unobservability and
competitive input markets. The proposed method requires conditional
independence of input-specific demand shocks. Scalar
unobservability rules out procurement relationships, supply
contracts, and input-specific markdowns; the proposed method
permits these. The GNR competitive input market assumption
precludes markup estimation, since the identifying condition
coincides with the object of interest. The conditional part of
the independence assumption depends on the adequacy of the control
variables $z_{jt}$, but this dependence is shared by the ACF
proxy equation.
\subsection{Specification Diagnostics}
\label{subsec:spec_diagnostics}
\label{subsec:exclusion_test}
Two diagnostics probe different layers of the identification
strategy before any structural results are interpreted:
\begin{enumerate}[label=(\roman*)]
\item \textbf{Exclusion restriction diagnostic}: tests whether
the pairwise discrepancy $d_k = d_l = 0$
(equation~\eqref{eq:overid_diff}), a necessary condition for
the exclusion restriction of Corollary~\ref{thm:exclusion} that
resolves the $\Delta(k,l)$ indeterminacy nonparametrically.
\item \textbf{Block~C diagnostic}: assesses the strength of the
CES curvature ($\rho_v \neq 0$), the identifying condition for
separate recovery of $(\beta_k, \beta_l)$ via
Theorem~\ref{thm:homothetic_id}
(Appendix~\ref{sec:appendix_empirical},
Table~\ref{tab:blockc_diagnostics}).
\end{enumerate}
Table~\ref{tab:spec_diagnostics} summarizes these two diagnostics
and their empirical outcomes.
\begin{table}[htbp]
\centering
\caption{Identification Roadmap and Specification Diagnostics}
\label{tab:spec_diagnostics}
{\small
\begin{tabular}{lllll}
\toprule
Diagnostic & Assumption tested & Enables & Null hypothesis & Outcome \\
\midrule
Exclusion diagnostic
& Excl.\ restriction (Corollary~\ref{thm:exclusion})
& Check on $\hat{\beta}_k$
& $d_k = d_l = 0$
& Capital only \\
Block~C diagnostic
& Homotheticity + CES curvature (Theorem~\ref{thm:homothetic_id})
& $\hat{\beta}_k$, $\hat{\beta}_l$
& $\rho_v \neq 0$
& Section~\ref{subsec:capital_labor} \\
\bottomrule
\end{tabular}
}
\smallskip
\small\raggedright \textit{Notes:
The two diagnostics probe successive layers of the identification
strategy. Row~1 is tested using Blocks~A and~B alone and
requires only Cobb--Douglas and conditional independence; results
are reported in Sections~\ref{subsec:spec_diagnostics}--\ref{subsec:main_results}.
Row~2 additionally invokes Assumption~\ref{ass:homothetic} (Homothetic
Weak Separability); results are reported in Section~\ref{subsec:capital_labor}.
The exclusion diagnostic ($d_k = d_l = 0$) provides a Wald test with
2 degrees of freedom.
Block~C diagnostic details are in Table~\ref{tab:blockc_diagnostics}.}
\end{table}
\paragraph{Exclusion restriction diagnostic.}
Figure~\ref{fig:recovery} applies the exclusion-based OLS recovery
of Proposition~\ref{prop:excl_ols} to all 502 manufacturing
industries, plotting $\hat{\beta}_k^{(m)}$ against
$\hat{\beta}_k^{(e)}$ (panel~a) and $\hat{\beta}_l^{(m)}$ against
$\hat{\beta}_l^{(e)}$ (panel~b).
Under the exclusion restriction, both panels should cluster along
the 45-degree line.
Panel~(a) confirms this for capital: points concentrate tightly
around the diagonal, consistent with $a_{k}^{h} = 0$ across industries.
Panel~(b) reveals the opposite for labor: points scatter widely,
indicating that different proxy equations yield systematically
different $\hat{\beta}_l$ values.
The asymmetry between capital and labor is the central diagnostic
finding. Capital is quasi-fixed within the production period and
does not directly influence short-run intermediate input procurement,
so $a_{k}^{h} = 0$ is economically plausible.
The systematic failure for labor is consistent with labor affecting
production scheduling, shift patterns, and input utilization through
input-specific channels
(Appendix~\ref{sec:micro_fnd}).
The formal Wald test of $d_k = d_l = 0$ is rejected for 37\% of
industries at the 5\% level, while the labor-only Wald test ($d_l = 0$)
is rejected for 28\% of industries, confirming that the labor exclusion
restriction is violated for a substantial share of the sample while capital
passes in most cases.
\begin{figure}[htbp]
\centering
\includegraphics[width=\textwidth]{Figure/Empirical/Exclusion/Beta_kl_Recovery_both}
\caption{Recovery of $(\beta_k, \beta_l)$ via the Exclusion Restriction}
\label{fig:recovery}
\small\raggedright \textit{Notes: Each industry's $\beta_k$ and
$\beta_l$ are recovered via OLS from each proxy equation
using Proposition~\ref{prop:excl_ols}. Panel~(a):
$\hat{\beta}_k^{(m)}$ versus $\hat{\beta}_k^{(e)}$.
Panel~(b): $\hat{\beta}_l^{(m)}$ versus
$\hat{\beta}_l^{(e)}$. Dashed lines are the 45-degree
reference. Under the exclusion restriction, both panels
should cluster along the diagonal. Outliers $|\hat{\beta}| > 2$ are
trimmed for readability; the full distribution is
reported in Table~\ref{tab:beta_kl_summary}.}
\end{figure}
\subsection{Production Function Parameters and Markups}
\label{subsec:main_results}
The intermediate input elasticities $(\beta_m, \beta_e, \beta_w)$
are identified by Blocks~A and~B alone (Theorem~\ref{thm:density_id},
specialized to the Cobb--Douglas parametric model of
Section~\ref{sec:gmm_approach}),
without requiring Block~C or Assumption~\ref{ass:homothetic}.
The ACF estimates of $\hat{\beta}_m$ are systematically higher
than those from the proposed method
(Table~\ref{tab:all_industry_beta_dist} in
Appendix~\ref{sec:appendix_empirical}),
consistent with the Markov misspecification bias documented in the
Monte Carlo simulations (Section~\ref{sec:simulation}).
The cross-industry distribution of all Block~A+B and Block~C
parameter estimates is reported in
Table~\ref{tab:all_industry_beta_dist} in
Appendix~\ref{sec:appendix_empirical}.
Electricity and water elasticities are small across industries
(median $\hat{\beta}_e = 0.001$ and $\hat{\beta}_w = 0.006$,
respectively), consistent with these inputs serving auxiliary
rather than central production roles in Japanese manufacturing.
Their demand shocks nevertheless remain valid exclusion
restrictions for identifying $\beta_m$ in the proposed GMM.
Under perfect competition, $\beta_m$ equals the revenue share,
which is the basis of GNR's share regression.
My estimator identifies $\beta_m$ independently of the
first-order condition, permitting imperfect competition in both
product and input markets.
\paragraph{Markups.}
Markups are computed following \textcite{deloecker2012markups}.
Under the Cobb-Douglas specification maintained throughout,
the markup formula simplifies to
\begin{equation}
\hat{\mu}_{jt} = \frac{\hat{\beta}_m}{s_{m,jt}},
\label{eq:markup_cd}
\end{equation}
where $s_{m,jt}$ denotes the expenditure share of raw materials
in total revenue.
Unlike the standard production approach, in which Hicks-neutral
productivity and scalar unobservability jointly imply
$\hat{\beta}_h/s_{h,jt} = \mu_{jt}$ for every variable input
$h$ (so that materials, labor, and energy serve as interchangeable
markup proxies), this paper allows input-specific markdowns
$\psi_{h,jt}$ for each static input $h \in \{m,e,w\}$, captured by
the demand shocks $(\tau_{jt},\nu_{jt},\eta_{jt})$
(Appendix~\ref{sec:micro_fnd}).
Consequently, $\hat{\beta}_h/s_{h,jt}$ will generally differ across
inputs by design; this divergence reflects the richer structure of
the framework, not an overidentification failure.
Raw materials are selected for markup computation because competitive
commodity markets support the absence of buyer-side market power
($\psi_{m,jt}\approx 1$; \textcite{avignon2025markups}),
giving $\hat{\beta}_m/s_{m,jt}\approx\mu_{jt}$;
this is a maintained assumption.
\footnote{The empirical specification imposes Hicks-neutral
Cobb-Douglas production; if factor-augmenting productivities
differ across inputs, $\hat{\beta}_m$ may absorb non-neutral
components and bias the markup estimate
\parencite{raval2023testing}.
The identification theory accommodates non-Hicks-neutral production
(Appendix~\ref{app:prod_recovery}), but the implemented GMM
does not exploit this generality.}
These estimates require only Blocks~A and~B and are invariant to
the $\Delta(k,l)$ indeterminacy (Theorem~\ref{thm:obs_equiv}),
since equation~\eqref{eq:markup_cd} depends only on
$\hat{\beta}_m$ and the observable cost share.
I restrict the comparison to industries with at least 50 firms
($N_{\text{firms}} \geq 50$), which removes industries where the
lower bound $\hat{\beta}_m \approx 0$ reflects identification failure
rather than true low input elasticities.
\footnote{The value-added markup
$\mu^{VA} = \beta_l^{VA}/s_l^{VA}$ can differ substantially
from the gross output markup when the materials share is large.
\textcite{gandhihowheterogeneous} document that gross output and
value-added specifications yield fundamentally different
productivity estimates. I report gross output markups throughout.}
\paragraph{Comparison with ACF.}
Figure~\ref{fig:markup_cdf} plots the empirical CDF of industry-level
median markups under the proposed method and the ACF benchmark
for the $N_{\text{firms}} \geq 50$ subsample.
The two distributions are stochastically ordered: the ACF CDF lies
strictly to the right of the proposed CDF at every percentile
(Table~\ref{tab:markup_acf_comparison}).
The proposed method yields a median markup of 0.926, while ACF
yields 1.027, a gap of 0.101 at the median.
At the 90th percentile the gap widens to approximately 0.15.
Under the proposed method, 37\% of industries show markups above
unity, compared with 54\% under ACF.
The Monte Carlo evidence in Section~\ref{sec:simulation} provides
a structural interpretation.
Under DGP~3 (potential-outcome dynamics, Table~\ref{tab:mc_part1_dgp3potential}),
ACF incurs a bias of $+0.19$ in $\hat{\beta}_m$ (true value $0.30$),
a 63\% relative overestimate, while the proposed estimator is essentially
unbiased ($\text{bias} = 0.001$).
The empirical gap of $+0.10$ at the median corresponds to a relative
overestimate of roughly 11\% in $\hat{\beta}_m$, well within the range
predicted by the DGP~3 calibration.
The evidence is therefore consistent with the theoretical prediction
that ACF overestimates $\hat{\beta}_m$ when productivity dynamics
deviate from the Markov assumption. Because markups recovered from
production functions are widely used to assess the evolution of
market power \parencite{deloecker2020rise}, the systematic gap
documented here raises the question of whether existing markup
estimates are sensitive to the choice of identifying assumption.
\begin{figure}[htbp]
\centering
\includegraphics[width=0.78\textwidth]{Figure/Empirical/Markup/Markup_CDF_comparison}
\caption{Empirical CDF of Industry Markups: Proposed vs.\ ACF}
\label{fig:markup_cdf}
\small\raggedright \textit{Notes:} Empirical CDFs of industry-level median
markups $\hat{\mu} = \hat{\beta}_m / \bar{s}_m$ under the proposed method
(solid, blue) and ACF (dashed, red).
Sample restricted to industries with $N_{\text{firms}} \geq 50$
($N = 372$ industries).
Vertical dotted line at $\hat{\mu} = 1$.
Summary statistics in Table~\ref{tab:markup_acf_comparison}.
\end{figure}
\begin{table}[htbp]
\centering
\caption{\label{tab:markup_acf_comparison}Markup Distribution: Proposed vs.\ ACF ($N_{\text{firms}}\geq 50$)}
\begin{tabular}{lcc}
\toprule
& \textbf{Proposed} & \textbf{ACF} \\
\midrule
$N$ (industries) & 372 & 372 \\
Mean & 0.880 & 1.026 \\
Std.\ dev. & 0.396 & 0.360 \\
p10 & 0.296 & 0.627 \\
p25 & 0.748 & 0.841 \\
Median & 0.926 & 1.027 \\
p75 & 1.074 & 1.233 \\
p90 & 1.240 & 1.431 \\
Fraction $\geq 1$ & 0.371 & 0.543 \\
Mean gap (ACF $-$ Proposed) & 0.146 & \multicolumn{1}{c}{---} \\
\bottomrule
\multicolumn{3}{p{0.75\textwidth}}{\small\textit{Notes:}
Industry-level median markups $\hat{\mu} = \hat{\beta}_m / \bar{s}_m$,
where $\bar{s}_m$ is the industry median materials share.
Sample: $N_{\text{firms}}\geq 50$.
ACF estimates from \textcite{ackerberg2015identification}; convergence code~0 for 495 of 502 industries.
}
\end{tabular}
\end{table}
\subsection{Productivity Determinants}
\label{subsec:determinants}
The following analysis requires only Blocks~A and~B. Because the
$\Delta(k,l)$ indeterminacy (Theorem~\ref{thm:obs_equiv}) varies
only through $(k_{jt}, l_{jt})$, it is absorbed by the cubic
polynomial controls in $(k_{jt}, l_{jt})$ included in the
regression. The same argument applies to proportional common shocks
(Section~\ref{subsec:relax_indep}): if an unobserved common
component $\xi_{jt}$ loads proportionally on all intermediate
input demands, it is absorbed into the recovered productivity
$\hat{\omega}_{jt}$, and its $(k_{jt}, l_{jt})$-dependent
component is absorbed by firm fixed effects. The determinants
regression therefore identifies the association between
covariates and the total latent efficiency measure that drives
input allocation decisions, regardless of whether this measure
coincides with physical productivity.
As a validation of the recovered productivity measures, I examine
their association with observable economic fundamentals.
I regress the productivity residual jointly on three firm-level
covariates (log investment, exporter status, and log wages) with
firm and year fixed effects:
\[
\hat{\omega}_{jt} = \phi_j + \gamma_t
+ \mathbf{x}_{jt}'\boldsymbol{\beta} + u_{jt},
\]
clustering standard errors at the firm level.
For the proposed method, I additionally include a cubic polynomial
in $(k_{jt}, l_{jt})$ as nonparametric controls, since
the $\Delta(k,l)$ indeterminacy enters through capital and labor.
The ACF regression omits these controls, as the ACF residual already
subtracts $\hat{\beta}_k k + \hat{\beta}_l l$.
Log wages is included as a correlate of productivity; a maintained
caveat is that wages may be endogenous, as high-productivity firms
can share rents with workers, so the coefficient captures association
rather than a causal effect.
Table~\ref{tab:main_regressions} reports the results.
Under the proposed method, log wages are strongly positively
associated with estimated productivity
($\hat{\beta} \approx 0.139$, $p < 0.01$), and log investment
is positive but small ($\hat{\beta} \approx 0.001$, $p < 0.01$),
while exporter status is negligible and imprecisely estimated.
The ACF regression yields a smaller wage coefficient
($\hat{\beta} \approx 0.109$). The mechanism is as follows:
under ACF, the upward bias in $\hat{\beta}_m$ propagates into
$\hat{\omega}^{\text{ACF}} = y - \hat{\beta}_m^{\text{ACF}} m
- \hat{\beta}_k k - \hat{\beta}_l l$, subtracting too large a
materials component and systematically depressing the recovered
productivity level for materials-intensive firms. This distortion
attenuates the association between $\hat{\omega}$ and economic
fundamentals that covary with input intensity. The magnitude of
the improvement depends on the relative variance of demand shocks
and productivity; whether the pattern generalizes beyond this
application requires further investigation.
\begin{table}[htbp]
\caption{\label{tab:main_regressions} Productivity Determinants: Proposed vs.\ ACF}
\bigskip
\centering
\begin{tabular}{lcc}
\toprule
& (1) Proposed & (2) ACF \\
\midrule
log(Investment) & 0.0011$^{***}$ & 0.0003$^{**}$\\
& (0.0003) & (0.0001)\\
Exporter Status & 0.0131 & -0.0010\\
& (0.0174) & (0.0068)\\
log(Wage) & 0.1388$^{***}$ & 0.1085$^{***}$\\
& (0.0185) & (0.0127)\\
\\
Observations & 433,425 & 433,308\\
R$^2$ & 0.86968 & 0.84551\\
\\
Firm FE & $\checkmark$ & $\checkmark$\\
Time FE & $\checkmark$ & $\checkmark$\\
\bottomrule
\end{tabular}
\par \raggedright
All three covariates enter jointly in a single specification. Firm and year fixed effects included. Standard errors clustered at the firm level in parentheses.\\
Exporter Status is a binary indicator equal to one if the firm exported in that year, consistent with the learning-by-exporting literature. log(Wage) is included as a correlate of productivity; wage endogeneity is a maintained caveat, as high-productivity firms may pay higher wages through rent-sharing.\\
Column~(1): proposed method with $\text{poly}(k,l,\text{degree}=3)$ nonparametric controls (coefficients suppressed).\\
Column~(2): ACF residual $\hat{\omega}^{\text{ACF}} = y - \hat{\beta}_k k - \hat{\beta}_l l - \hat{\beta}_m m - \hat{\beta}_e e - \hat{\beta}_w w$.\\
Significance: $^{*}$ $p<0.10$, $^{**}$ $p<0.05$, $^{***}$ $p<0.01$.
\end{table}
\subsection{Event Study: 2011 Tohoku Earthquake}
\label{subsec:event_study_body}
Because the proposed estimator recovers productivity from static
covariances alone, its estimates are valid under any productivity
dynamics, Markov or otherwise
(Proposition~\ref{prop:omega_hat_D}).
Standard proxy variable estimators embed a Markov transition equation
that is structurally incompatible with the potential outcomes
framework when a treatment alters the transition path of
productivity \parencite{chen2024identifying}; the moment condition
that identifies the production function parameters has no structural
interpretation under treatment, so the resulting estimates lack
economic meaning for policy evaluation. Neither problem arises
here, since estimation does not employ a transition equation.
As an illustration, I examine the 2011 T\={o}hoku earthquake using a
difference-in-differences design.
The treatment group consists of plants in the three core prefectures
directly struck by the earthquake and tsunami (Iwate, Miyagi, and
Fukushima; seismic intensity $\geq$ 6-strong), where physical
destruction and the nuclear disaster caused severe and sustained
disruption to production.
The control group consists of plants in Kinki and western prefectures
(prefectures 25--47).
Supply chain contamination of the control group is mitigated by the
industry$\times$year fixed effects, which absorb any
industry-level aggregate shocks that propagate nationally.
Pre-treatment coefficients are flat
(max$|\hat{\delta}_t| < 0.013$ for the proposed method,
$< 0.020$ for ACF);
the full event-study figure is in
Appendix~\ref{subsec:earthquake_es}.
Table~\ref{tab:earthquake_did} reports the difference-in-differences
estimates under both methods.
For the proposed method, cubic polynomial controls in $(k,l)$ are
included to absorb the $\Delta(k,l)$ indeterminacy in the residual
$\hat{\omega}$; the ACF method requires no such controls, as
$\hat{\omega}^{\mathrm{ACF}}$ already subtracts $\hat{\beta}_k k
+ \hat{\beta}_l l$.
Both methods detect a negative and statistically significant
post-treatment effect on productivity.
Under the proposed method, the DiD estimate is $-1.28$ percent
(s.e.\ $0.52$, $p < 0.05$); under ACF, the corresponding
estimate is $-1.68$ percent (s.e.\ $0.41$, $p < 0.01$).
The gap between the two estimates is approximately
$0.40$ percentage points. To illustrate the potential economic
magnitude: if a bias of this order applied to the aggregate
manufacturing sector, it would correspond to roughly \$3.6 billion
(\textyen{}400 billion) per year, given Japan's manufacturing value
added of approximately \$0.9 trillion (\textyen{}100 trillion) at
the 2003--2020 average exchange rate (National Accounts, Cabinet
Office of Japan).\footnote{This back-of-envelope
calculation extrapolates the local DiD gap to the national level
under the assumption that the Markov misspecification bias is of
comparable magnitude across industries. The cross-industry markup
comparison (Table~\ref{tab:markup_acf_comparison}) shows that ACF
yields systematically higher $\hat{\beta}_m$ at every percentile,
consistent with the assumption, but the magnitude varies by
industry. The figure should be interpreted as indicative of the
scale at stake, not as a structural estimate of aggregate
mismeasurement.}
Proposition~\ref{prop:omega_hat_D} guarantees that the proposed
estimates recover $\mathbb{E}[\omega_{jt}\mid D_{jt}]$ under
Conditions~(i)--(ii) of that proposition.
Condition~(ii) is satisfied by construction: the earthquake is a
natural disaster whose occurrence is orthogonal to firm-level
input demand shocks $(\tau,\nu,\eta)$.
Condition~(i) requires that the earthquake does not alter the
functional form of the demand functions $g_m, g_e, g_w$.
Difference-in-differences estimates of the post-treatment change
in intermediate input shares show no significant shift in the
materials share ($t = 0.87$) or water share ($t = 1.14$).
The electricity share shows a small post-treatment increase
($t = 8.91$, $\Delta s_e \approx 0.002$); this is a mechanical
compositional effect of the simultaneous contraction in materials
usage ($t = -4.53$), which raises the electricity expenditure
share $s_e$ without altering the structural demand function
$g_e$.\footnote{The level of electricity consumption does not
show a significant post-treatment increase when measured in
physical units (kWh) rather than expenditure shares, supporting
the compositional interpretation.}
The ACF estimator does not carry this guarantee.
Its residual subtracts $\hat{\beta}_k k + \hat{\beta}_l l$, so
\begin{equation}
\mathbb{E}[\hat{\omega}^{\mathrm{ACF}}_{jt} \mid D_{jt}]
= \mathbb{E}[\omega_{jt} \mid D_{jt}]
+ (\beta_k^{\mathrm{true}} - \hat{\beta}_k^{\mathrm{ACF}})
\,\mathbb{E}[k_{jt} \mid D_{jt}]
+ (\beta_l^{\mathrm{true}} - \hat{\beta}_l^{\mathrm{ACF}})
\,\mathbb{E}[l_{jt} \mid D_{jt}].
\label{eq:acf_bias_es}
\end{equation}
The bias terms vanish only if
$\hat{\beta}^{\mathrm{ACF}} = \beta^{\mathrm{true}}$
(exact identification) or if treatment is orthogonal to $(k,l)$.
Monte Carlo evidence (Section~\ref{sec:simulation}) shows that ACF incurs
positive bias in $\hat{\beta}_m$ under Markov misspecification;
the condition of orthogonality also fails here (DiD$(l) = -0.029$, $t = -7.86$).
The proposed estimates, resting on the theoretical guarantee of
Proposition~\ref{prop:omega_hat_D}, provide a theoretically
justified point of comparison.
\begin{table}[htbp]
\caption{\label{tab:earthquake_did} 2011 T\={o}hoku Earthquake: Difference-in-Differences}
\bigskip
\centering
\begin{tabular}{lcc}
\toprule
& (1) Proposed & (2) ACF \\
\midrule
Treated $\times$ Post & -0.0128$^{**}$ & -0.0168$^{***}$\\
& (0.0052) & (0.0041)\\
\\
Observations & 219,573 & 219,573\\
R$^2$ & 0.98729 & 0.94606\\
$\text{poly}(k,\ell)$ control & $\checkmark$ & \\
\\
Firm FE & $\checkmark$ & $\checkmark$\\
Ind.$\times$Year FE & $\checkmark$ & $\checkmark$\\
\bottomrule
\end{tabular}
\par \raggedright
Treatment: Iwate, Miyagi, Fukushima (seismic intensity $\geq$ 6-strong). Control: West Japan (prefectures 25--47).\\
Firm and industry$\times$year fixed effects included. Heteroskedasticity-robust standard errors in parentheses.\\
Pre-treatment coefficients are flat (max$|\hat{\delta}_t| < 0.013$ for proposed, $< 0.020$ for ACF); year-by-year coefficient estimates in Table~\ref{tab:earthquake_es} (Appendix~\ref{subsec:earthquake_es}).\\
Column~(1): proposed method with $\text{poly}(k,\ell,\text{degree}=3)$ nonparametric controls for $\Delta(k,\ell)$ (coefficients suppressed).\\
Column~(2): ACF residual already subtracts $\hat{\beta}_k k + \hat{\beta}_l l$; no polynomial control.\\
Significance: $^{*}$ $p<0.10$, $^{**}$ $p<0.05$, $^{***}$ $p<0.01$.
\end{table}
\subsection{Capital and Labor Inputs}
\label{subsec:capital_labor}
Identifying $(\beta_k, \beta_l)$ requires closing the
$\Delta(k,l)$ indeterminacy documented in
Theorem~\ref{thm:obs_equiv}.
The paper provides two independent routes:
the exclusion restriction (Corollary~\ref{thm:exclusion}) and the
homothetic regularity condition (Theorem~\ref{thm:homothetic_id}).
The exclusion-based OLS recovery
(Proposition~\ref{prop:excl_ols}) is applied to all 502
manufacturing industries and produces mutually consistent estimates
of $\beta_k$ across the three proxy equations (materials,
electricity, water), while $\beta_l$ estimates diverge
systematically, confirming the diagnostic pattern in
Figure~\ref{fig:recovery} that the exclusion restriction holds for
capital but not labor.
To identify $(\beta_k, \beta_l)$ jointly, I apply the Block~C
homothetic CES approach (Section~\ref{sec:homothetic}).
Block~C is the primary identification route for both $\beta_k$ and
$\beta_l$; the exclusion restriction provides an independent check
on $\beta_k$ only, since the restriction fails for labor
(Figure~\ref{fig:recovery}, Panel~b).
The two strategies yield mutually consistent estimates of $\beta_k$
for the 302{} industries where the exclusion restriction
is validated for capital
(Table~\ref{tab:4group},
Appendix~\ref{sec:appendix_empirical}).
Table~\ref{tab:beta_kl_summary} summarizes the cross-industry
distributions of $\hat{\beta}_k$ and $\hat{\beta}_l$ across
three approaches: exclusion restriction (broken out by proxy input),
Block~C (homothetic CES), and ACF.
\begin{table}[htbp]
\centering
\caption{Cross-Industry Distribution of $\hat{\beta}_k$ and $\hat{\beta}_l$: Exclusion Restriction, Block~C, and ACF}
\label{tab:beta_kl_summary}
\small
\begin{tabular}{l r r r r r}
\toprule
& Excl.\ ($m$) & Excl.\ ($e$) & Excl.\ ($w$) & Block~C & ACF \\
\midrule
$\hat{\beta}_k$ Median & 0.0086 & 0.0220 & 0.0106 & 0.0350 & 0.0297 \\
$\hat{\beta}_k$ Mean & 0.0121 & 0.0010 & -0.0077 & 0.0481 & 0.0528 \\
$\hat{\beta}_k$ SD & 0.2018 & 0.2724 & 0.1935 & 0.0544 & 0.0921 \\
\midrule
$\hat{\beta}_l$ Median & 0.2106 & 0.2600 & 0.2010 & 0.3316 & 0.2753 \\
$\hat{\beta}_l$ Mean & -0.1625 & -0.2815 & -0.2317 & 0.3357 & 0.3217 \\
$\hat{\beta}_l$ SD & 3.8991 & 5.1635 & 3.7362 & 0.2239 & 0.2657 \\
\midrule
$N$ & 389 & 389 & 389 & 502 & 499 \\
\bottomrule
\end{tabular}
\vspace{0.5em}
\begin{minipage}{0.95\textwidth}\footnotesize
\textit{Notes:} Industries with $|\hat{\beta}|>2$ excluded. Excl.\ ($m$/$e$/$w$): exclusion restriction OLS using each proxy (Proposition~\ref{prop:excl_ols}). Block~C: homothetic CES approach (Theorem~\ref{thm:homothetic_id}). ACF: \textcite{ackerberg2015identification}.
\end{minipage}
\end{table}
Estimates of $\hat{\beta}_k$ are broadly consistent across all three
approaches (median $\approx 0.01$--$0.04$), corroborating the
identification cross-check in Table~\ref{tab:4group}
(Appendix~\ref{sec:appendix_empirical}).
Labor elasticity estimates diverge more substantially:
Block~C yields a median $\hat{\beta}_l = 0.33$, while ACF produces
a higher median of $0.50$.
The Monte Carlo simulations (Table~\ref{tab:mc_part2_dgp3potential})
show that under DGP~3, ACF $\hat{\beta}_l$ collapses to near zero
(bias $\approx -0.30$), while the proposed method recovers the true
value accurately (bias $\approx +0.01$).
The empirical ACF estimate lies above the proposed estimate,
which is the opposite direction from the MC collapse.
Both patterns reflect the same fragility: ACF labor elasticity
identification breaks down when the Markov assumption is violated,
with the direction of the deviation depending on the specific
dynamics of the data-generating process.
Figure~\ref{fig:beta_kl_3methods} shows the full cross-industry
density distributions for all three methods.
A four-group comparison across identification strategies is
reported in Table~\ref{tab:4group}
(Appendix~\ref{sec:appendix_empirical}).
\begin{figure}[htbp]
\centering
\includegraphics[width=\textwidth]{Figure/Empirical/Production/Beta_kl_Distribution_3Methods}
\caption{Cross-Industry Distribution of $\hat{\beta}_k$ and $\hat{\beta}_l$: Three Methods}
\label{fig:beta_kl_3methods}
\small\raggedright \textit{Notes: Panel~(a) shows the density of
$\hat{\beta}_k$ from Exclusion, Homothetic (Block~C), and
ACF. Panel~(b) shows $\hat{\beta}_l$ for all three
methods; Exclusion estimates use the materials proxy
(Proposition~\ref{prop:excl_ols}).
Industries with $|\hat{\beta}| > 2$ are excluded.
Summary statistics in Table~\ref{tab:beta_kl_summary};
four-group identification cross-check in
Table~\ref{tab:4group} (Appendix~\ref{subsec:crosscheck_4group}).}
\end{figure}
\section{Conclusion}
\label{sec:conclusion}
Can the production function be identified without restricting how
productivity evolves over time? This paper answers in the
affirmative. Replacing the Markov assumption with conditional
independence across three intermediate inputs, the paper shows that
the production function and the distribution of productivity are
nonparametrically identified from a single cross-section. No
assumption on the law of motion for $\omega_{jt}$ is required at
any stage of estimation. The empirical analysis, covering 502
Japanese manufacturing industries, confirms that the choice between
the two identification strategies has quantitative consequences for
every downstream object: input elasticities, markups, allocative
efficiency, and the measured response of productivity to economic
shocks.
The consequences are economically large. The proposed method yields
systematically lower markups than the standard proxy variable
estimator across the entire distribution (median 0.93 vs.\ 1.03;
the share of industries above unity falls from 54 to 37 percent),
shifting the measured degree of market power in the manufacturing
sector. In the earthquake event
study (Section~\ref{subsec:event_study_body}), the
difference-in-differences estimate of the productivity effect on
plants in the three most severely affected prefectures is
$-1.28\%$ under the proposed method and $-1.68\%$ under the
standard method; the 0.40 percentage point gap corresponds to
roughly \$3.6 billion (\textyen{}400 billion) per year in
mismeasured productivity when scaled to aggregate manufacturing
output. The
\textcite{olley1996thedynamics} decomposition and the productivity
determinant regressions reinforce the same pattern: the log-wage
coefficient is roughly 25\% larger under the proposed method
(Table~\ref{tab:main_regressions}), consistent with a higher
signal-to-noise ratio in the recovered productivity measure once
input-specific demand shocks are separated out. The Monte Carlo
simulations, the convergence diagnostic, and the determinant
regressions all point in the same direction, and the underlying
mechanism is general: in any setting where a policy, shock, or
institutional change alters the transition path of productivity,
the Markov transition equation is structurally misspecified and
the resulting production function parameters lack a consistent
interpretation. Trade liberalization, R\&D subsidies, natural
disasters, and mergers all generate such dynamics. The proposed
method accommodates these settings because it imposes no
restriction on how productivity evolves. The Cobb--Douglas
functional form is shared by both the proposed method and the ACF
benchmark, so the gap between estimates reflects the difference in
identifying assumptions, not in functional form.
These findings connect to two broader debates. First, the recent
literature on rising global markups
\parencite{deloecker2020rise} relies on production function
estimates that impose the Markov assumption. The present results
suggest that markup levels, and potentially trends, are sensitive to
this assumption; replication of the global markup finding under
conditional independence identification is a natural next step.
Second, \textcite{chen2024identifying} show that standard proxy
variable estimators are structurally incompatible with a potential
outcomes framework: the Markov transition equation has no structural
interpretation when a treatment alters the productivity process, so
the resulting estimates lack economic meaning under policy
evaluation. The proposed method avoids this problem because it uses
no transition equation; Proposition~\ref{prop:omega_hat_D}
establishes that the recovered productivity measure retains a
causal interpretation under treatment assignment mechanisms
satisfying conditions~(i)--(ii) of that proposition (the treatment
does not alter demand function structure and is orthogonal to
input-specific demand shocks). The same static structure accommodates time-varying
parameters without additional assumptions, since no intertemporal
link is imposed.
A broader implication concerns the nature of identifying assumptions
in production function estimation. The Markov restriction is a
constraint on the time-series behavior of an unobservable; the
conditional independence restriction is a constraint on the
structure of input markets. The latter is grounded in economic
primitives (separate suppliers, distinct procurement channels,
independent regulatory regimes), and the researcher can specify
which observable controls restore the assumption when a particular
threat is identified. This transparency provides a basis for
evaluating the credibility of the estimates that has no analogue
under the Markov framework.
On the theoretical side, Theorem~\ref{thm:obs_equiv} characterizes
the residual indeterminacy that arises once the Markov assumption is
dropped. Two routes close this indeterminacy: an exclusion
restriction (Corollary~\ref{thm:exclusion}) with a testable
necessary condition (Remark~\ref{rem:testable_exclusion}), and a
homothetic regularity condition
(Theorem~\ref{thm:homothetic_id}). The two routes yield mutually
consistent estimates in industries where both apply.
Three limitations and corresponding directions for future work
deserve mention. First, the identification strategy requires at
least three intermediate inputs with separately observable quantity
data, though this requirement is met in several settings beyond the
Japanese Census of Manufactures, including the U.S.\ EIA Form~923
\parencite{fabrizio2007dothey,cicala2015when} and emissions data
in environmental economics.\footnote{
Additional datasets satisfying this requirement include
India's Annual Survey of Industries (ASI), which reports
firm-level electricity and fuel consumption alongside materials;
Canada's Annual Survey of Manufacturing and Logging (ASML),
which covers electricity and water use at the establishment level;
and the World Bank Enterprise Survey (WBES), which collects
firm-level electricity expenditure and water source data across
over 100 countries. These datasets enable direct application
of the proposed estimator in diverse institutional settings.}
When labor adjustment is rapid, labor itself serves as an additional
productivity signal, reducing the required number of intermediate
inputs from three to two (footnote~\ref{rem:static_labor});
extending the framework to such settings is a natural direction.
Second, no targeted test of the conditional independence assumption
alone exists; the convergence diagnostic of
Remark~\ref{rem:testable_exclusion} provides a necessary condition
for the exclusion restriction. Extending the moment system to
achieve overidentification (for instance via Block~C structural
constraints or cross-equation demand restrictions under
Cobb--Douglas) would enable formal specification testing. Third,
Block~C identification of $(\beta_k, \beta_l)$ requires
non-negligible curvature in $h(v)$; when the capital-labor ratio
varies little, the exclusion restriction route becomes preferable,
and combining the static identification of flexible input
elasticities with semiparametric methods for the capital-labor
component is left for future research.
\clearpage
\printbibliography
\clearpage
\thispagestyle{empty}
\vspace*{\fill}
\begin{center}
{\LARGE\bfseries Appendix}\\[1.5em]
{\large Nonparametric Identification and Estimation of Production Functions\\
Invariant to Productivity Dynamics}\\[1em]
{\normalsize Rentaro Utamaru}
\end{center}
\vspace*{\fill}
\clearpage