EconBase
← Back to paper

Production Function Estimation without Invertibility: Imperfectly Competitive Environments and Demand Shocks

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

95,179 characters · 8 sections · 15 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Production Function Estimation without Invertibility: Imperfectly Competitive Environments and Demand Shocks

abstractWe advance the proxy variable approach to production function estimation. We show that the invertibility assumption at its heart is testable. We characterize what goes wrong if invertibility fails and what can still be done. We show that rethinking how the estimation procedure is implemented either eliminates or mitigates the bias that arises if invertibility fails. In particular, a simple change to the first step of the estimation procedure provides a first-order bias correction for the GMM estimator in the second step. Furthermore, a modification of the moment condition in the second step ensures Neyman orthogonality and enhances efficiency and robustness by rendering the asymptotic distribution of the GMM estimator invariant to estimation noise from the first step.

Introduction

Production functions are one of the oldest concepts in economics VONT:42,WICK:94,COBB:28. As a description of the relationship between inputs and output, they are of interest in and of themselves and also a vehicle for measuring productivity and technical change SOLO:57. More recently, production functions have become a key input into estimating markups and markdowns in order to assess firms' market power HALL:88,DELO:12. The production approach to markup estimation recognizes that, given an estimate of the production function, the markup can be recovered from the firm's cost minimization problem. To estimate the production function, this large and rapidly growing literature relies on the procedure initially developed by \citeasnoun{OLLE:96} and further advanced by \citeasnoun{LEVI:03} and \citeasnoun{ACKE:15} (henceforth OP, LP, and ACF).

The OP/LP/ACF procedure starts from the observation that there is an endogeneity problem in estimating production functions. As first pointed out by \citeasnoun{MARS:44}, this problem arises because the decisions that the firm makes regarding its inputs depend on its productivity, which is unobserved by the econometrician. To resolve this problem, OP turn it on its head and use the firm's decisions that are observed by the econometrician to infer its productivity. The resulting proxy variable approach to production function estimation has quickly become standard practice.

The inversion from observables to productivity at the heart of the OP/LP/ACF procedure requires that firms with the same productivity make the same choices. Yet, there is no reason to believe this is the case, e.g., if these firms face different demands in the output market. The current paper therefore revisits the OP/LP/ACF procedure if invertibility fails. It characterizes what goes wrong in the OP/LP/ACF procedure and what can still be done.

The point that the OP/LP/ACF procedure cannot accommodate unobserved demand heterogeneity has been made before. \citeasnoun{FOST:08} put it as follows:

quote\ldots\ idiosyncratic demand shocks make the proxies functions of both technology and demand shocks, thereby inducing a possible omitted variable bias. Put simply, proxy methods require a one-to-one mapping between plant-level productivity and the observables used to proxy for productivity. This mapping breaks down if other unobservable plant-level factors besides productivity drive changes in the observable proxy. (p. 403)

Indeed, to establish invertibility, OP rule out that firms face different demands and abstract from competition between firms.\footnote{OP assume that any profitability differences across firms are due to differences in their capital stocks and productivities (p. 1273), thereby ruling out that firms face different demands. Limiting the state variables in the firm's investment policy to its own capital stock and productivity moreover abstracts from competition between firms \citeaffixed{PAKE:94B}{see also Lemma 3 and Theorem 1 in}.} LP assume a perfectly competitive industry where firms act as price takers and thus face the same horizontal demand curve (p. 322 and Appendix A). Based on their results, ACF start from scalar unobservable and strict monotonicity assumptions, and almost all subsequent literature simply imposes invertibility as a high-level assumption. For this reason, it may not have fully appreciated the demanding nature of the invertibility assumption.

Unobserved demand heterogeneity is likely to be empirically important. The large literatures on demand estimation and productivity analysis highlight the considerable heterogeneity in demand that remains even after controlling for detailed product attributes BERR:95 or honing in on (nearly) homogenous products FOST:08. The problem is compounded by the fact that, in imperfectly competitive environments, a firm's decisions in equilibrium depend on its own productivity as well as on the productivities of its rivals. Hence, imperfectly competitive environments require jointly inverting the decisions of all firms for the productivities of all firms. This many-to-many inverse may not exist BION:22 or it may be too high-dimensional to be practical ACKE:24. Moreover, a firm's rivals are partially or completely unobserved in typical datasets used for production function estimation. The decisions of its rivals have thus to be thought of as shocks to the demand the focal firm faces.

Invertibility can fail for reasons other than unobserved demand heterogeneity. Changes in firm conduct due to mergers and acquisitions or switches from competition to collusion are best thought of as changes to the ownership matrix. Invertibility can fail if these changes in firm conduct are partially or completely unobserved by the econometrician. Finally, invertibility can fail if there is unobserved variation across firms or time in input prices, investment opportunities, or financial constraints.

In light of the demanding nature of the invertibility assumption, this paper makes five contributions. First, we propose tests for invertibility. Our tests exploit that if invertibility fails, then productivity becomes a hidden state in a Markov model. The literature on Kalman and particle filters in dynamic systems shows that the best guess for the hidden state uses the entire history of the observables as opposed to just their current value. Invertibility can therefore be tested by including lags of the observables in the regression in the first step of the OP/LP/ACF procedure. Implementing our tests on an unbalanced panel of Spanish manufacturing firms and a balanced panel of US manufacturing industries, we strongly reject invertibility.

We therefore characterize the consequences of a failure of invertibility for the OP/LP/ACF procedure. The first step of the OP/LP/ACF procedure regresses output on observables. If invertibility fails, then the prediction of output contains an error. Because the prediction is used to control for lagged productivity in the GMM estimation in the second step, the lagged prediction error enters into the conditional moment. Building on \citeasnoun{DORA:21}, we show that this invalidates capital as an instrument and results in biased estimates.

Turning from what goes wrong if invertibility fails to what can still be done, our second contribution is to provide a necessary and sufficient condition for the moment condition in the second step of the OP/LP/ACF procedure to hold for the true production function and some (not necessarily the true) law of motion for productivity. This condition calls for rethinking how the OP/LP/ACF procedure is implemented. In particular, it compels us to ensure that any instrument used in the second step is appropriately included in the regression in the first step. Due to the lag structure of the model, this means including the lead of capital in addition to its current value in the regression. We provide a series of examples where this simple expedient suffices to satisfy our necessary and sufficient condition.

Beyond these examples, our necessary and sufficient condition may be violated. While this results in biased estimates, our third contribution is to show that rethinking how the OP/LP/ACF procedure is implemented mitigates the bias. Specifically, we show that the moment condition in the second step implicitly incorporates a first-order bias correction provided any instrument used in the second step is appropriately included in the regression in the first step.

Our fourth contribution is to explicitly incorporate a bias correction into the second step of the OP/LP/ACF procedure. We show that the modified moment condition has a property known as Neyman orthogonality NEYM:59. While our modification remains as straightforward to implement as the original OP/LP/ACF procedure, Neyman orthogonality renders the asymptotic distribution of the GMM estimator in the second step invariant to estimation noise and the quality of the prediction of output from the first step. This is particularly advantageous if the regression in the first step includes a large number of covariates. Neyman orthogonality facilitates the use of a wide range of estimation methods in the first step, including traditional nonparametric methods as well as modern machine learning techniques such as neural networks and random forests. A Monte Carlo exercise shows that Neyman orthogonality can substantially improve the performance of the GMM estimator in finite samples.

Our fifth and final contribution is to provide a diagnostic to assess the sensitivity of the estimates to the size of the prediction error in the first step. While our diagnostic has some similarities to the sensitivity measure in \citeasnoun{ANDR:17}, it is not local to the true model and directly informative about the estimated model. Our diagnostic is neither necessary nor sufficient for no bias. However, a small value of the diagnostic provides assurance that small changes in the prediction error do not dramatically alter the estimates.

In sum, this paper examines the OP/LP/ACF procedure if the invertibility assumption fails. Invertibility fails if there are demand shocks or in imperfectly competitive environments with partially or completely unobserved rivals or changes in firm conduct. Whether invertibility fails can be tested. A failure of invertibility can have a substantial, detrimental impact on the estimates.

Fortunately, much can still be done. We provide a necessary and sufficient condition for the moment condition in the second step of the OP/LP/ACF procedure to hold for the true production function. This condition compels us to ensure that any instrument used in the second step is appropriately included in the regression in the first step. We show that this simple change either eliminates or mitigates the bias. Going a step further, we modify the moment condition in the second step to endow it with Neyman orthogonality. Finally, we provide a diagnostic to assess the sensitivity of the estimates to the size of the prediction error that arises in the first step of the OP/LP/ACF procedure if invertibility fails.

The remainder of the paper is organized as follows. In Section (ref), we recall the setup and the OP/LP/ACF procedure. In Section (ref), we develop tests for invertibility. In Section (ref), we develop a necessary and sufficient condition for the moment condition in the second step to hold for the true production function. We show that ensuring that any instrument used in the second step is appropriately included in the regression in the first step either eliminates or mitigates the bias that arises if invertibility fails. In Section (ref), we modify the moment condition in the second step to endow it with Neyman orthogonality. In Section (ref), we conduct a Monte Carlo exercise to illustrate what goes wrong in the OP/LP/ACF procedure if invertibility fails and what can still be done. In Section (ref), we provide a diagnostic to assess the sensitivity of the estimates to the size of the prediction error. We conclude in Section (ref).

Setup and OP/LP/ACF procedure

Firm $i$ in period $t$ produces output $Q_{it}$ with inputs $K_{it}$ and $V_{it}$ according to the production function

equation[equation omitted — 87 chars of source]

where lower case letters denote logs. Capital $k_{it}$ is a predetermined input that is chosen in period $t-1$ whereas $v_{it}$ is freely variable and decided on in period $t$ after the firm observes its productivity $\omega_{it}$.\footnote{While $k_{it}$ and $v_{it}$ may be vectors, we think of them as scalars for simplicity. The variable input $v_{it}$ may accordingly be interpreted as a composite of labor and materials such as cost of goods sold.} Productivity follows a first-order Markov process with law of motion

equation[equation omitted — 136 chars of source]

where the productivity innovation $\xi_{it}$ is by construction mean independent of lagged productivity $\omega_{it-1}$ and further assumed to be mean independent of any variable included in the firm's information set in period $t-1$. The disturbance $\varepsilon_{it}$ sits between the firm's output $q_{it}$ as recorded in the data and the output $q^*_{it}=q_{it}-\varepsilon_{it}=f(k_{it},v_{it})+\omega_{it}$ that the firm planned on when it decided on the variable input $v_{it}$. It can be interpreted as measurement error or as an unanticipated shock to output (OP, pp. 1273--1274) and is assumed to be mean independent of the inputs and other included variables as formalized below.\footnote{\citeasnoun{MUND:65} refer to $\omega_{it}$ and $\varepsilon_{it}$ as the transmitted, respectively, untransmitted component of productivity. The untransmitted component may include machine breakdowns, labor actions, supply chain disruptions, and power outages that are not anticipated by the firm. The interpretation as measurement error accommodates serial correlation in the disturbance $\varepsilon_{it}$.} While the econometrician observes actual output $q_{it}$ and the inputs $k_{it}$ and $v_{it}$, productivity $\omega_{it}$, the disturbance $\varepsilon_{it}$, and planned output $q^*_{it}$ remain unobserved.

\paragraph{Invertibility.}

The literature following OP relies on invertibility. Invertibility assumes that there exists a function $\omega_{it}=h(x_{it})$ that maps observables $x_{it}=\left(k_{it},v_{it},\ldots\right)$ into productivity $\omega_{it}$. This is equivalently to

equation*[equation* omitted — 81 chars of source]

The observables included in $x_{it}$ depend on the decision of the firm that is being inverted.\footnote{OP invert the firm's demand for investment whereas LP and ACF invert its demand for materials.} We remain agnostic about which variables are included in $x_{it}$, aside from imposing, without loss of generality, that the inputs $k_{it}$ and $v_{it}$ are included.

\paragraph{OP/LP/ACF procedure.}

Estimation proceeds in two steps. Step 1 assumes $\text{\normalfont E}\left[\varepsilon_{it}|x_{it}\right]=0$ and flexibly or nonparametrically estimates the conditional expectation

equation[equation omitted — 287 chars of source]

where the first equality uses equation (ref), the second-to-last equality uses invertibility, and the last equality uses the definition of planned output $q^*_{it}$. Because $\text{\normalfont E}\left[q_{it}|x_{it}\right]=q^*_{it}$, step 1 separates actual output $q_{it}$ into planned output $q^*_{it}$ and the disturbance $\varepsilon_{it}=q_{it}-q^*_{it}$.

Step 2 of the OP/LP/ACF procedure assumes $\text{\normalfont E}\left[\left.\xi_{it}+\varepsilon_{it}\right|z_{it}\right]=0$ for instruments $z_{it}=\left(k_{it},k_{it-1},v_{it-1},\ldots\right)$ and estimates the parameters $\theta=(\theta_f,\theta_g)$ in the production function $f$ and the law of motion $g$ by GMM.\footnote{For notational convenience we suppress $\theta$ in much of what follows.} Capital $k_{it}$ is a valid instrument because it is a predetermined input, and the lagged inputs $k_{it-1}$ and $v_{it-1}$ are valid instruments because the productivity innovation $\xi_{it}$ is mean independent of any variable included in the firm's information set in period $t-1$. We remain agnostic regarding additional instruments in $z_{it}$.

Using equations (ref) and (ref) and the definition of planned output $q^*_{it}$, the assumption $\text{\normalfont E}\left[\left.\xi_{it}+\varepsilon_{it}\right|z_{it}\right]=0$ implies

equation[equation omitted — 220 chars of source]

Moment condition (ref) is infeasible for estimation because lagged planned output $q^*_{it-1}$ is unobserved. Substituting $\text{\normalfont E}[q_{it-1}|x_{it-1}]$ from step 1 for $q^*_{it-1}$, step 2 therefore proceeds by estimating the parameters $\theta$ from the moment condition

equation[equation omitted — 187 chars of source]

Throughout the remainder of the paper, we follow OP, LP, and ACF and maintain $\text{\normalfont E}\left[\varepsilon_{it}|x_{it}\right]=0$ for observables $x_{it}=\left(k_{it},v_{it},\ldots\right)$ and $\text{\normalfont E}\left[\left.\xi_{it}+\varepsilon_{it}\right|z_{it}\right]=0$ for instruments $z_{it}=\left(k_{it},k_{it-1},v_{it-1},\ldots\right)$. To avoid cumbersome notation, we adopt the convention that all equalities involving random variables and conditional expectations are understood to hold almost surely.

Tests for invertibility

If invertibility fails and $\text{\normalfont E}\left[\omega_{it}|x_{it}\right]\neq\omega_{it}$, then productivity $\omega_{it}$ becomes a hidden state in a Markov model. The literature on Kalman and particle filters in dynamic systems shows that the best guess for the unobservables uses the entire history of the observables as opposed to just their current value. Our tests for invertibility are based on this intuition.

Our first proposition shows that invertibility implies a testable mean-independence restriction:

propositionIf $\text{\normalfont E}\left[\omega_{it}|x_{it}\right]=\omega_{it}$ and $\text{\normalfont E}\left[\varepsilon_{it}|x_{it},x_{it-1}\right]=\text{\normalfont E}\left[\varepsilon_{it}|x_{it}\right]$, then $\text{\normalfont E}\left[q_{it}|x_{it},x_{it-1}\right]=\text{\normalfont E}\left[q_{it}|x_{it}\right]$.

While step 1 of the OP/LP/ACF procedure assumes $\text{\normalfont E}\left[\varepsilon_{it}|x_{it}\right]=0$ for observables $x_{it}=\left(k_{it},v_{it},\ldots\right)$ and step 2 assumes $\text{\normalfont E}\left[\xi_{it}+\varepsilon_{it}|z_{it}\right]=0$ for instruments $z_{it}=\left(k_{it},k_{it-1},v_{it-1},\ldots\right)$, this does not quite imply $\text{\normalfont E}\left[\varepsilon_{it}|x_{it},x_{it-1}\right]=\text{\normalfont E}\left[\varepsilon_{it}|x_{it}\right]$. However, the interpretation of the disturbance $\varepsilon_{it}$ as an unanticipated shock to output that is outside the firm's information set in period $t$ implies $\text{\normalfont E}\left[\varepsilon_{it}|x_{it},x_{it-1}\right]=\text{\normalfont E}\left[\varepsilon_{it}|x_{it}\right] = 0$.\footnote{Under the interpretation of the disturbance $\varepsilon_{it}$ as measurement error, the assumption $\text{\normalfont E}\left[\varepsilon_{it}|x_{it},x_{it-1}\right]=\text{\normalfont E}\left[\varepsilon_{it}|x_{it}\right]$ holds if $x_{it-1}$ does not have predictive power for $\varepsilon_{it}$ beyond $x_{it}$.} The proof of Proposition (ref) is straightforward:

proofIf $\text{\normalfont E}\left[\omega_{it}|x_{it}\right]=\omega_{it}$, then \begin{equation*} q_{it}=f(k_{it},v_{it})+\omega_{it}+\varepsilon_{it}=f(k_{it},v_{it})+\normalfont E\left[\omega_{it}|x_{it}\right]+\varepsilon_{it}. \end{equation*} $\text{\normalfont E}\left[\varepsilon_{it}|x_{it},x_{it-1}\right]=\text{\normalfont E}\left[\varepsilon_{it}|x_{it}\right]$ therefore implies $\text{\normalfont E}\left[q_{it}|x_{it},x_{it-1}\right]=\text{\normalfont E}\left[q_{it}|x_{it}\right]$.

Our second proposition shows that invertibility implies a testable conditional-independence restriction:

propositionIf $\text{\normalfont E}\left[\omega_{it}|x_{it}\right]=\omega_{it}$ and $\left.\varepsilon_{it} \protect\mathpalette{\protect\independenT}{\perp} x_{it-1}\right|x_{it}$, then $q_{it} \protect\mathpalette{\protect\independenT}{\perp} x_{it-1}|x_{it}$.

Going beyond the first-order implication of $\text{\normalfont E}\left[\omega_{it}|x_{it}\right]=\omega_{it}$ in Proposition (ref) calls for a stronger assumption on the disturbance $\varepsilon_{it}$. The assumption $\left.\varepsilon_{it} \protect\mathpalette{\protect\independenT}{\perp} x_{it-1}\right|x_{it}$ in Proposition (ref) means that any possible dependence between $\varepsilon_{it}$ and $x_{it-1}$ is through $x_{it}$. It is implied by $\varepsilon_{it} \protect\mathpalette{\protect\independenT}{\perp} (x_{it},x_{it-1})$ and implies $\text{\normalfont E}\left[\varepsilon_{it}|x_{it},x_{it-1}\right]=\text{\normalfont E}\left[\varepsilon_{it}|x_{it}\right]$. The proof of Proposition (ref) parallels the proof of Proposition (ref) and is therefore omitted.

We implement our tests for invertibility on an unbalanced panel of Spanish manufacturing firms from 1990 to 2006 (Encuesta Sobre Estrategias Empresariales) and a balanced panel of US manufacturing industries from 1958 to 2018 (NBER-CES). In both cases, we reject invertibility by a wide margin. We provide further details on the data and the tests in Appendix (ref). In the next section, we turn to the consequences of a failure of invertibility for production function estimation.

Failure of invertibility and validity of OP/LP/ACF moment condition

If invertibility fails and $\text{\normalfont E}\left[\omega_{it}|x_{it}\right]\neq\omega_{it}$, then a prediction error $\zeta_{it}=\omega_{it}-\text{\normalfont E}\left[\omega_{it}|x_{it}\right]$ arises in step 1 of the OP/LP/ACF procedure that is generally non-zero. By construction, the prediction error satisfies $\text{\normalfont E}\left[\zeta_{it}|x_{it}\right]=0$.

To see how the presence of the prediction error $\zeta_{it}$ affects the OP/LP/ACF procedure, recall that the conditional expectation estimated in step 1 is

equation[equation omitted — 217 chars of source]

where the first equality uses equation (ref) and $\text{\normalfont E}[\varepsilon_{it}|x_{it}] = 0$, the second equality uses the definition of the prediction error $\zeta_{it}$, and the last equality uses the definition of planned output $q^*_{it}$. Equation (ref) provides an alternative interpretation of the prediction error as the difference between planned output and its prediction in step 1: $\zeta_{it} = q^*_{it} - \text{\normalfont E}[q_{it}|x_{it}]$.

Substituting $\text{\normalfont E}[q_{it-1}|x_{it-1}]$ from equation (ref), the left-hand side of moment condition (ref) in step 2 of the OP/LP/ACF procedure becomes

equation[equation omitted — 181 chars of source]

Because of the lagged prediction error $\zeta_{it-1}$, conditional moment (ref) is generally non-zero. To see this more clearly, consider the special case of an $AR(1)$ process for productivity. If $g(\omega_{it-1})=\rho\omega_{it-1}$, then conditional moment (ref) evaluates to

equation[equation omitted — 176 chars of source]

where the equality uses $\text{\normalfont E}\left[\xi_{it}+\varepsilon_{it}|z_{it}\right]=0$.

To further see how this affects the GMM estimation, recall that $x_{it-1}=\left(k_{it-1},v_{it-1},\ldots\right)$ and $z_{it}=\left(k_{it},k_{it-1},v_{it-1},\ldots\right)$. The lagged inputs $k_{it-1}$ and $v_{it-1}$ hence remain valid instruments because $\text{\normalfont E}\left[\zeta_{it-1}|x_{it-1}\right]=0$ by construction. However, capital $k_{it}$ may no longer be a valid instrument as $\text{\normalfont E}\left[\zeta_{it-1}|k_{it}\right]\neq 0$, as noted by \citeasnoun{DORA:21}. This is because the firm chooses $k_{it}$ in period $t-1$ with knowledge of $\omega_{it-1}$. What remains of lagged productivity after controlling for the lagged observables $x_{it-1}$ may therefore be correlated with $k_{it}$. This results in biased estimates.

The above discussion foreshadows part of the solution to the problem: adding the lead of capital $k_{it+1}$ to the observables $x_{it}$ in step 1 of the OP/LP/ACF procedure ensures that $\text{\normalfont E}\left[\zeta_{it-1}|k_{it}\right]=0$ and hence the validity of capital $k_{it}$ as an instrument in step 2. This simple expedient amounts to rethinking how the OP/LP/ACF procedure is implemented. The remaining question is to what extent it suffices beyond the special case of an $AR(1)$ process for productivity.

The following theorem provides a complete answer to this question in the form of a necessary and sufficient condition.

theoremMoment condition (ref) in step 2 of the OP/LP/ACF procedure holds for the true production function $f^0$ and some law of motion $\tilde g$ if and only if \begin{equation} \normalfont E\left[g^0(\omega_{it-1})|z_{it}\right]=\normalfont E\left[\left.\tilde g\left(\normalfont E\left[\omega_{it-1}|x_{it-1}\right]\right)\right|z_{it}\right], \end{equation} where $g^0$ is the true law of motion.

We prove Theorem (ref) before parsing it:

proofRecall that the true production function $f^0$ and the true law of motion $g^0$ satisfy moment condition (ref). {\em “Only if” part:} Suppose moment condition (ref) holds for the true production $f^0$ and some law of motion $\tilde g$. Then we have \begin{gather*} 0=\normalfont E\left[\left.q_{it}-f^0(k_{it},v_{it})-\tilde g\left(\normalfont E\left[q_{it-1}|x_{it-1}\right]-f^0(k_{it-1},v_{it-1})\right)\right|z_{it}\right] \\ =\normalfont E\left[\left.q_{it}-f^0(k_{it},v_{it})-\tilde g\left(\normalfont E\left[\omega_{it-1}|x_{it-1}\right]\right)\right|z_{it}\right], \end{gather*} where the last equality uses equation (ref). Subtracting from equation (ref) implies condition (ref). {\em “If” part:} Suppose condition (ref) holds. Substituting into equation (ref), we have \begin{gather*} 0=\normalfont E\left[\left.q_{it}-f^0(k_{it},v_{it})-\tilde g\left(\normalfont E\left[\omega_{it-1}|x_{it-1}\right]\right)\right|z_{it}\right] \\ =\text{\normalfont E}\left[\left.q_{it}-f^0(k_{it},v_{it})-\tilde g\left(\text{\normalfont E}\left[q_{it-1}|x_{it-1}\right]-f^0(k_{it-1},v_{it-1})\right)\right|z_{it}\right], \end{gather*} where the last equality uses equation (ref). Hence, moment condition (ref) holds for $f^0$ and $\tilde g$.

Theorem (ref) covers the invertibility assumption at the heart of the literature following OP as a special case. If $\text{\normalfont E}\left[\omega_{it}|x_{it}\right]=\omega_{it}$, then condition (ref) becomes $\text{\normalfont E}\left[g^0(\omega_{it-1})|z_{it}\right]=\text{\normalfont E}\left[\tilde g\left(\omega_{it-1}\right)|z_{it}\right]$ and is therefore satisfied with $\tilde g=g^0$. Importantly, however, condition (ref) can be satisfied even if invertibility fails and $\text{\normalfont E}\left[\omega_{it}|x_{it}\right]\neq \omega_{it}$.

To understand condition (ref), note that it becomes $\text{\normalfont E}\left[g^0(\omega_{it-1})|z_{it}\right]=\tilde g\left(\text{\normalfont E}\left[\omega_{it-1}|x_{it-1}\right]\right)$ if $x_{it-1}\subsetneq z_{it}$. This is difficult to satisfy as the left-hand side is a function of $z_{it}$ whereas the right-hand side is a function of $x_{it-1}\subsetneq z_{it}$.

We therefore set $x_{it-1}\supseteq z_{it}$ in what follows. This amounts to rethinking how the OP/LP/ACF procedure is implemented, as foreshadowed by the special case of an $AR(1)$ process for productivity. In particular, because $k_{it}$ is included in $z_{it}$ in step 2, we must include $k_{it+1}$ in $x_{it}$ in step 1.

With this choice of $x_{it-1}\supseteq z_{it}$, we further examine condition (ref) through a series of examples:

example[linear law of motion, $AR(1)$ process] If $g^0(\omega_{it-1})=\rho\omega_{it-1}$, then $\text{\normalfont E}\left[\rho\omega_{it-1}|z_{it}\right]=\text{\normalfont E}\left[\rho \text{\normalfont E}\left[\omega_{it-1}|x_{it-1}\right]|z_{it}\right]$ by the law of iterated expectations. Condition ((ref)) is therefore satisfied with $\tilde g(\text{\normalfont E}[\omega_{it-1}|x_{it-1}])=g^0(\text{\normalfont E}[\omega_{it-1}|x_{it-1}])$.

Remarkably, moment condition (ref) in step 2 of the OP/LP/ACF procedure holds in Example (ref) irrespective of the quality of the estimate of $\text{\normalfont E}\left[q_{it}|x_{it}\right]$ in step 1.

example[quadratic law of motion] If $g^0(\omega_{it-1})=\rho_1\omega_{it-1}+\rho_2\omega_{it-1}^2$ and $\text{\normalfont Var}\left(\zeta_{it-1}|z_{it}\right)=\sigma^2$, then \begin{equation*} \normalfont E\left[\left.\rho_1\omega_{it-1}+\rho_2\omega_{it-1}^2\right|z\right]=\normalfont E\left[\left.\rho_1\normalfont E\left[\omega_{it-1}|x_{it-1}\right]+\rho_2\normalfont E\left[\omega_{it-1}|x_{it-1}\right]^2+\rho_2\zeta_{it-1}^2\right|z_{it}\right]. \end{equation*} Condition (ref) is therefore satisfied with $\tilde g(\text{\normalfont E}[\omega_{it-1}|x_{it-1}])=g^0(\text{\normalfont E}[\omega_{it-1}|x_{it-1}])+\rho_2\sigma^2$.\footnote{We can relax $\text{\normalfont Var}\left(\zeta_{it-1}|z_{it}\right)=\sigma^2$ to $\text{\normalfont Var}\left(\zeta_{it-1}|z_{it}\right) = h(\text{\normalfont E}[\omega_{it-1}|x_{it-1}])$ for some function $h$. In this case, condition (ref) holds with $\tilde g(\text{\normalfont E}[\omega_{it-1}|x_{it-1}])=g^0(\text{\normalfont E}[\omega_{it-1}|x_{it-1}])+\rho_2 h(\text{\normalfont E}[\omega_{it-1}|x_{it-1}])$. }

Example (ref) is less restrictive on the law of motion but more restrictive on the lagged prediction error than Example (ref). The next example pushes this tradeoff further:

example[analytic law of motion] If the Taylor series of $g^0$ around $\text{\normalfont E}[\omega_{it-1}|x_{it-1}]$ converges absolutely so that $g^0(\omega_{it-1})=\sum_{j=0}^\infty \frac{g^{0,(j)}(\text{\normalfont E}[\omega_{it-1}|x_{it-1}])}{j!}\zeta^j_{it-1}$, where $g^{0, (j)}$ is the $j$th derivative of $g^0$, then \begin{gather*} \normalfont E\left[\left.\sum_{j=0}^\infty \frac{g^{0,(j)}(\normalfont E[\omega_{it-1}|x_{it-1}])}{j!}\zeta^j_{it-1}\right|z_{it}\right] =\normalfont E\left[\left.\normalfont E\left[\left.\sum_{j=0}^\infty\frac{g^{0,(j)}(\normalfont E[\omega_{it-1}|x_{it-1}])}{j!}\zeta^j_{it-1}\right|x_{it-1}\right]\right|z_{it}\right] \\ =\normalfont E\left[\left.\sum_{j=0}^\infty\frac{g^{0,(j)}(\text{\normalfont E}[\omega_{it-1}|x_{it-1}])}{j!}\text{\normalfont E}[\zeta^j_{it-1}|x_{it-1}]\right|z_{it}\right]. \end{gather*} If $\text{\normalfont E}[\zeta^j_{it-1}|x_{it-1}]=\alpha_j$ for some constant $\alpha_j$ for all $j$, then condition (ref) is satisfied in this example with $\tilde g(\text{\normalfont E}[\omega_{it-1}|x_{it-1}])=\sum_{j=0}^\infty\frac{g^{0,(j)}(\text{\normalfont E}[\omega_{it-1}|x_{it-1}])}{j!}\alpha_j$.\footnote{Condition (ref) remains valid if we relax the condition that $\text{\normalfont E}[\zeta^j_{it-1}|x_{it-1}]=\alpha_j$ for some constant $\alpha_j$ for all $j$ to $\left.\zeta_{it-1}\protect\mathpalette{\protect\independenT}{\perp} z_{it}\right| \text{\normalfont E}[\omega_{it-1}|x_{it-1}]$.}

Beyond these examples, condition (ref) may be violated even if $x_{it-1}\supseteq z_{it}$.\footnote{Testing condition (ref) is empirically challenging. A necessary condition for it to be satisfied is that moment condition (ref) in step 2 of the OP/LP/ACF procedure holds at the estimated production function and law of motion. This can be tested with an overidentification test that corrects for the plug-in nature of the OP/LP/ACF procedure and clustering in the data.} While this results in biased estimates, setting $x_{it-1}\supseteq z_{it}$ mitigates the bias. If $x_{it-1}\supseteq z_{it}$, then we have $\text{\normalfont E}\left[\omega_{it-1}|z_{it}\right]=\text{\normalfont E}\left[\text{\normalfont E}\left[\omega_{it-1}|x_{it-1}\right]|z_{it}\right]$ by the law of iterated expectations and hence $\text{\normalfont E}\left[\zeta_{it-1}|z_{it}\right]$ $=$ $\text{\normalfont E}\left[\left.\omega_{it-1}-\text{\normalfont E}\left[\omega_{it-1}|x_{it-1}\right]\right|z_{it}\right]$ $=0$. This mean independence of the lagged prediction error $\zeta_{it-1}$ from the instruments $z_{it}$ intuitively eliminates a first-order source of bias. The following theorem formalizes this intuition:

theoremIf $x_{it-1}\supseteq z_{it}$ and $g$ is differentiable, then moment condition (ref) in step 2 of the OP/LP/ACF procedure implicitly incorporates a first-order bias correction: \begin{multline} \normalfont E\left[q_{it}-f(k_{it},v_{it})-g\left(\normalfont E\left[q_{it-1}|x_{it-1}\right]-f(k_{it-1},v_{it-1})\right)|z_{it}\right] = \normalfont E\left[q_{it}-f(k_{it},v_{it}) \right.\\ - g\left(q^*_{it-1} - \zeta_{it-1} -f(k_{it-1},v_{it-1}) \right) - \left. g^\prime\left(q^*_{it-1} - \zeta_{it-1} -f(k_{it-1},v_{it-1}) \right)\zeta_{it-1} \big| z_{it}\right], \end{multline} where $g^{\prime}$ is the first derivative of $g$.

We note that equation (ref) holds for any production function $f$ and any law of motion $g$ for which the conditional moments exist. The proof of Theorem (ref) is in Appendix (ref).

To appreciate Theorem (ref), recall that moment condition (ref) is valid even if invertibility fails. Because lagged planned output $q^*_{it-1}$ is unobserved, moment condition (ref) is infeasible for estimation. Step 2 of the OP/LP/ACF procedure therefore proceeds by substituting $\text{\normalfont E}\left[q_{it-1}|x_{it-1}\right]$ from step 1 for $q^*_{it-1}$. If invertibility fails, then this substitution introduces the lagged prediction error $\zeta_{it-1}=q^*_{it-1}-\text{\normalfont E}\left[q_{it-1}|x_{it-1}\right]$ into moment condition (ref) (see conditional moment (ref)). Theorem (ref) shows that setting $x_{it-1}\supseteq z_{it}$ adjusts moment condition (ref) toward the valid but infeasible moment condition (ref).

Indeed, comparing the term $g\left(\text{\normalfont E}\left[q_{it-1}|x_{it-1}\right]-f(k_{it-1},v_{it-1})\right)$ in moment condition (ref) to the term $g\left(q^*_{it-1}-f(k_{it-1},v_{it-1})\right)$ in moment condition (ref) yields

eqnarray[eqnarray omitted — 433 chars of source]

where the last equality uses a first-order Taylor expansion of $g$ around $q^*_{it-1} - \zeta_{it-1} -f(k_{it-1},v_{it-1})$ and $o(\zeta_{it-1})$ denotes the higher-order remainder. Equation (ref) in Theorem (ref) shows that if $x_{it-1}\supseteq z_{it}$, then moment condition (ref) implicitly accounts for the term $-g^\prime\left(q^*_{it-1} - \zeta_{it-1}-f(k_{it-1},v_{it-1}) \right) \zeta_{it-1}$. Moment condition (ref) thus differs from the valid but infeasible moment condition (ref) at most by higher-order terms. Put differently, setting $x_{it-1}\supseteq z_{it}$ provides a first-order bias correction in step 2 of the OP/LP/ACF procedure.

Because setting $x_{it-1}\supseteq z_{it}$ provides a first-order bias correction in step 2 of the OP/LP/ACF procedure, a violation of condition (ref) must be of second order relative to size of the prediction error. The following theorem formalizes this intuition:

theoremLet $\mathcal{G}$ be a set of functions that includes the true law of motion $g^0$. If $x_{it-1}\supseteq z_{it}$ and $g^{0}$ is twice continuously differentiable, then \begin{equation*} \inf_{\tilde g\in \mathcal{G}}\left\Vert \normalfont E\left[g^0(\omega_{it-1})|z_{it}\right]-\normalfont E\left[\tilde g\left(\normalfont E\left[\omega_{it-1}|x_{it-1}\right]\right)|z_{it}\right]\right\Vert _{L,1}\leq \tau\normalfont Var\left(\zeta_{it-1}\right), \end{equation*} where $\left\Vert \cdot\right\Vert _{L, 1}$ is the $L_1$ norm, $\tau=\sup_{\omega_{it-1}}\left|g^{0\prime\prime}(\omega_{it-1})\right|$, and $g^{0\prime\prime}$ is the second derivative of $g^0$.

The proof of Theorem (ref) is in Appendix (ref).

Theorem (ref) ties the violation of condition (ref) to the variance of the lagged prediction error $\zeta_{it-1}$. Because $\text{\normalfont Var}(\zeta_{it-1}) = \text{\normalfont E}[ (\omega_{it-1} - \text{\normalfont E}[\omega_{it-1}|x_{it-1}])^2]$, this variance gets smaller as the observables $x_{it-1}$ get richer even beyond $z_{it}$. Theorem (ref) therefore suggests to take a kitchen sink approach to the regression in step 1 of the OP/LP/ACF procedure and add as many relevant covariates as possible. Adding covariates that speak to demand conditions and firm conduct may be especially helpful.\footnote{The recent literature in fact proceeds along this line: “The materials demand function in our setting will take as arguments all state variables of the firm (\ldots), including productivity, and all additional variables that affect a firm's demand for materials. These include firm location (\ldots), output prices (\ldots), product dummies (\ldots), market shares (\ldots), input prices (\ldots), the export status of a firm (\ldots), and the input (\ldots) and output tariffs (\ldots) that the firm faces on the product it produces” DELO:15. Whether adding covariates restores invertibility can be assessed using the tests in Section (ref).}

Adding covariates may make the assumption $\text{\normalfont E}\left[\varepsilon_{it}|x_{it}\right]=0$ that underpins step 1 of the OP/LP/ACF procedure more demanding. We advocate replacing $x_{it}=\left(k_{it},v_{it},\ldots\right)$ in the original procedure by $x_{it}=\left(k_{it+1},k_{it},v_{it},\ldots\right)$. In the original procedure, the assumption $\text{\normalfont E}\left[\varepsilon_{it}|x_{it}\right]=0$ is justified by stipulating that $\varepsilon_{it}$ is outside the firm's information set in period $t$. This justification extends to our approach as the firm is assumed to choose $k_{it+1}$ in period $t$. However, adding leads of other covariates to $x_{it}$ may be more challenging if the firm eventually learns $\varepsilon_{it}$. To guard against this possibility, we recommend adding current values or lags of other covariates to $x_{it}$.

Finally, we caution that adding covariates introduces a curse of dimensionality. The resulting larger mean squared error of the estimate of $\text{\normalfont E}[q_{it-1}| x_{it-1}]$ may counterbalance the benefit of the smaller $\text{\normalfont Var}(\zeta_{it-1})$ in finite samples. In the next section, we propose a modification of the OP/LP/ACF procedure that addresses this problem.

\paragraph{Summary.}

Our analysis calls for rethinking how the OP/LP/ACF procedure is implemented. Theorem (ref) makes a strong case for ensuring $x_{it-1}\supseteq z_{it}$. For a linear law of motion, the moment condition in step 2 of the OP/LP/ACF procedure holds for the true production function $f^0$ and the true law of motion $g^0$. For a nonlinear law of motion, recovering the true production function $f^0$ requires condition (ref) to be satisfied. A violation of condition (ref) results in biased estimates. In this case, Theorem (ref) shows that setting $x_{it-1}\supseteq z_{it}$ provides a first-order bias correction.

Theorem (ref) further suggests to be flexible in modeling the law of motion in step 2 as condition (ref) is easier to satisfy if $\tilde g$ can come from a larger set $\mathcal{G}$. Theorem (ref) finally suggests to take a kitchen sink approach to step 1.

\paragraph{Related literature.}

An alternative to relaxing the invertibility assumption is to forgo step 1 of the OP/LP/ACF procedure. Instead of substituting $\text{\normalfont E}\left[q_{it-1}|x_{it-1}\right]$ for $q^*_{it-1}$ in moment condition (ref), the dynamic panel approach to production function estimation pioneered by \citeasnoun{BLUN:00} substitutes $q_{it-1}$ and imposes the linear law of motion $g(\omega_{it-1}) = \rho \omega_{it-1}$ to obtain

equation*[equation* omitted — 236 chars of source]

Besides the assumption $\text{\normalfont E}\left[\xi_{it}+\varepsilon_{it}|z_{it}\right]=0$ maintained in moment condition (ref), the dynamic panel approach requires assuming $\text{\normalfont E}[\varepsilon_{it-1}|z_{it}] = 0$. In comparison, our Example (ref) requires setting $x_{it-1}\supseteq z_{it}$ and assuming $\text{\normalfont E}[\varepsilon_{it-1}|x_{it-1}] = 0$ in step 1 of the OP/LP/ACF procedure. Either assumption can be justified by appealing to the firm's information set and the timing of decisions, as discussed above.

\citeasnoun{HU:20}, \citeasnoun{BRAN:20}, and \citeasnoun{POND:21} generalize the dynamic panel approach from a linear to a nonlinear law of motion.\footnote{The estimation strategy in \citeasnoun{HU:20} and \citeasnoun{BRAN:20} differs from their identification strategies. The latter builds on \citeasnoun{HU:08} and relies on the existence of two conditionally independent proxies for productivity and the invertibility of an integral operator. The implied semiparametric maximum likelihood estimator involves significant computational challenges that preclude its implementation.} To illustrate the similarities and differences with our approach, consider the quadratic law of motion $g(\omega_{it-1}) = \rho_1 \omega_{it-1} + \rho_2 \omega^2_{it-1}$. Substituting $q_{it-1}$ for $q^*_{it-1}$ in moment condition (ref) yields

multline*[multline* omitted — 337 chars of source]

The generalized dynamic panel approach therefore requires assuming $\text{\normalfont E}[\varepsilon_{it-1}|z_{it}] = 0$, $\text{\normalfont E}[\varepsilon_{it-1}\omega_{it-1}| z_{it}] = 0$, and $\text{\normalfont Var}\left(\varepsilon_{it-1}|z_{it}\right) = \sigma^2$. In comparison, our Example (ref) requires setting $x_{it-1}\supseteq z_{it}$ and assuming $\text{\normalfont E}[\varepsilon_{it-1}|x_{it-1}] = 0$ and $\text{\normalfont Var}(\zeta_{it-1}|z_{it})=\sigma^2$. Both approaches are similar in that they impose exclusion restrictions on second-order moments of latent variables to ensure that the moment condition holds for the true production function.

This similarity extends to the polynomial law of motion $g(\omega_{it-1}) = \sum_{j=1}^\infty \rho_{j}\omega^j_{it-1}$. In this case, the generalized dynamic panel approach requires assuming that $\text{\normalfont E}[\varepsilon_{it-1}^j|\omega_{it-1}, z_{it}]$ is constant for all $j$. This restriction on the joint distribution of $\varepsilon_{it-1}$ and $\omega_{it-1}$ may be unpalatable if $\varepsilon_{it-1}$ and $\omega_{it-1}$ are interpreted as the untransmitted, respectively, transmitted component of productivity (see footnote (ref)).\footnote{If $x_{it-1}\subseteq z_{it}$ and invertibility holds, then $\text{\normalfont E}[\varepsilon_{it-1}^j|\omega_{it-1}, z_{it}]=\text{\normalfont E}[\varepsilon_{it-1}^j|z_{it}]$. While this facilitates interpreting the restriction, we focus on the case where invertibility fails.} In comparison, our Example (ref) requires assuming that $\text{\normalfont E}[\zeta_{it-1}^j | x_{it-1}]$ is constant for all $j$.

The key difference is that our approach imposes restrictions on the lagged prediction error $\zeta_{it-1}$ rather than the lagged disturbance $\varepsilon_{it-1}$. Whereas the magnitude and properties of $\varepsilon_{it-1}$ are inherent in the data generating process, the advantage of our approach is that the magnitude of $\zeta_{it-1}$ can be reduced by adding covariates to the regression in step 1 of the OP/LP/ACF procedure. Proceeding to add covariates may shrink $\text{\normalfont Var}(\zeta_{it-1}) = \text{\normalfont E}[ (\omega_{it-1} - \text{\normalfont E}[\omega_{it-1}|x_{it-1}])^2]$ towards zero if we approach invertibility in the limit. This makes our approach particularly attractive for datasets with a rich set of observables.

Modification of OP/LP/ACF moment condition

Step 1 of the OP/LP/ACF procedure estimates the conditional expectation $\text{\normalfont E}[q_{it-1}|x_{it-1}]$. To understand the impact that a noisy estimate in step 1 has on the GMM estimator in step 2, we use $\hat{e}(x_{it-1})$ to denote the estimate of a nuisance parameter $e(x_{it-1})$ with true value $\text{\normalfont E}[q_{it-1}|x_{it-1}]$. We remain agnostic about the estimation method used in step 1.

Replacing $\text{\normalfont E}[q_{it-1}|x_{it-1}]$ by $\hat{e}(x_{it-1})$ in moment condition (ref) that underpins step 2 yields

equation[equation omitted — 162 chars of source]

Because $\hat{e}(x_{it-1})$ converges to $\text{\normalfont E}[q_{it-1}|x_{it-1}]$ as the sample size increases, we write $\hat{e}(x_{it-1}) = \text{\normalfont E}[q_{it-1}|x_{it-1}] + \lambda \delta(x_{it-1})$, where $\lambda$ converges to zero as sample size increases and $\delta(x_{it-1})$ is the direction of the finite-sample deviation of $\hat{e}(x_{it-1})$ from $\text{\normalfont E}[q_{it-1}|x_{it-1}]$. The pathwise (or Gateaux) derivative of the left-hand side of moment condition (ref) with respect to the direction $\delta(x_{it-1})$ at $\hat{e}(x_{it-1})=\text{\normalfont E}[q_{it-1}|x_{it-1}]$ is

eqnarray*[eqnarray* omitted — 400 chars of source]

Because this derivative is generally non-zero, moment condition (ref) is impacted by the deviation of $\hat{e}(x_{it-1})$ from $\text{\normalfont E}[q_{it-1}|x_{it-1}]$ that arises from estimation noise in step 1.

To eliminate this impact and its adverse consequences for the asymptotic distribution of the GMM estimator, we modify the moment condition in step 2 as follows:

multline[multline omitted — 286 chars of source]

Our modification explicitly incorporates a feasible version of the first-order bias correction that is implicit in moment condition (ref) if $x_{it-1} \supseteq z_{it}$ (see Theorem (ref)).

If $\hat{e}(x_{it-1})=\text{\normalfont E}[q_{it-1}|x_{it-1}]$ and $x_{it-1} \supseteq z_{it}$, then we have $\text{\normalfont E}[q_{it-1} - \hat{e}(x_{it-1})| z_{it}] = 0$ so that moment condition (ref) coincides with moment condition (ref). Our modification therefore retains the identification power of the original OP/LP/ACF procedure when $x_{it-1} \supseteq z_{it}$.

However, unlike the original moment condition (ref), the modified moment condition (ref) is not sensitive to small deviations of $\hat e(x_{it-1})$ from $\text{\normalfont E}[q_{it-1}|x_{it-1}]$, a property known as Neyman orthogonality NEYM:59:

theorem[Neyman orthogonality] Define the set of functions $\tilde{\mathcal{T}} =\{ \delta(x_{it-1})=e(x_{it-1}) - \text{\normalfont E}[q_{it-1}| x_{it-1}]: e \in \mathcal{T} \}$, where $\mathcal{T}$ is the set of real-valued integrable functions of $x_{it-1}$. If $x_{it-1}\supseteq z_{it}$ and $g$ has a bounded and continuous derivative, then the pathwise (or Gateaux) derivative of the left-hand side of moment condition (ref) with respect to any direction $\delta\in\tilde{\mathcal{T}}$ at $\hat{e}(x_{it-1}) =\text{\normalfont E}[q_{it-1}| x_{it-1}]$ is zero: \begin{multline*} \frac{\partial}{\partial\lambda} \normalfont E\left[q_{it}-f(k_{it},v_{it}) - g\left( \normalfont E[q_{it-1}|x_{it-1}] + \lambda \delta(x_{it-1}) -f(k_{it-1},v_{it-1}) \right) \right.\\ - \left. g^\prime\left(\normalfont E[q_{it-1}|x_{it-1}] + \lambda \delta(x_{it-1}) -f(k_{it-1},v_{it-1}) \right)(q_{it-1} - \normalfont E[q_{it-1}|x_{it-1}] - \lambda \delta(x_{it-1})) \big| z_{it}\right] \Big |_{\lambda = 0} = 0. \end{multline*}

The proof of Theorem (ref) is in Appendix (ref).\footnote{The assumption that $g$ has a bounded derivative simplifies stating Theorem (ref). As shown in Appendix (ref), it can be replaced by a much weaker assumption.}$^,$\footnote{We recommend ensuring $x_{it-1}\supseteq z_{it}$ when implementing our modification of the OP/LP/ACF procedure. Without $x_{it-1}\supseteq z_{it}$, the modified moment condition (ref) entails a first-order bias correction toward $\text{\normalfont E}[q_{it-1}|x_{it-1}, z_{it}]$ instead of $\text{\normalfont E}[q_{it-1}|x_{it-1}]$ and the expectation of the correction term conditional on $z_{it}$ is generally non-zero even if $\hat{e}(x_{it-1}) = \text{\normalfont E}[q_{it-1}|x_{it-1}]$.}

Neyman orthogonality has received renewed attention in the double-debiased machine learning literature. Following the arguments in \citeasnoun{CHER:18} and \citeasnoun{CHER:22}, as long as $\hat{e}(x_{it-1})$ converges to $\text{\normalfont E}[q_{it-1}|x_{it-1}]$ at a rate faster than $N^{-1/4}$, the GMM estimator based on moment condition (ref) is oracle efficient. This means that its asymptotic distribution is {\em as if} the true value $\text{\normalfont E}[q_{it-1}|x_{it-1}]$ of the nuisance parameter $e(x_{it-1})$ is known. The GMM estimator based on moment condition (ref) therefore achieves the efficient asymptotic distribution.

Besides efficiency, the modified moment condition (ref) enjoys further advantages over the original moment condition (ref). First, because Theorem (ref) allows for any integrable deviation, it accommodates a wide range of estimation methods in step 1 of the OP/LP/ACF procedure. This includes traditional nonparametric methods such as sieve, kernel, and LASSO estimators, which produce smooth estimates, as well as modern machine learning techniques such as neural networks and random forests, which may produce non-smooth estimates. Neyman orthogonality means that small deviations of $\hat{e}(x_{it-1})$ from $\text{\normalfont E}[q_{it-1}|x_{it-1}]$ that can arise from estimation noise, functional approximation error in sieve estimation, local smoothing bias in kernel estimation, or regularization bias in LASSO and machine learning methods do not have a first-order impact on the GMM estimator in step 2. As long as $\hat{e}(x_{it-1})$ converges to $\text{\normalfont E}[q_{it-1}|x_{it-1}]$ at a rate faster than $N^{-1/4}$, its asymptotic distribution is robust to the choice of estimation method and its implementation in step 1.\footnote{Achieving this rate of convergence requires properly implementing a nonparametric estimation method to avoid functional form bias and the validity of smoothness assumptions specific to the selected method. Because Neyman orthogonality is a local property, the first-order bias correction in moment condition (ref) loses its theoretical justification if $\hat{e}(x_{it-1})$ deviates substantially from $\text{\normalfont E}[q_{it-1}|x_{it-1}]$.}

Second, moment condition (ref) eliminates the need to correct standard errors in step 2 to account for estimation noise in step 1 through techniques such as those in \citeasnoun{HAHN:18}.

Unlike general approaches that generate Neyman orthogonality using influence functions or score projections as described in \citeasnoun{CHER:22}, moment condition (ref) is Neyman orthogonal by construction. By avoiding introducing high-dimensional nuisance parameters, our modification remains as straightforward to implement as the original OP/LP/ACF procedure.

Finally, we note that if OLS is used in step 1 of OP/LP/ACF procedure to estimate the conditional expectation $\text{\normalfont E}[q_{it-1}|x_{it-1}]$, then explicitly including the first-order bias correction can be redundant in some cases. Converting moment condition (ref) into an unconditional moment condition yields

multline*[multline* omitted — 271 chars of source]

where $h(z_{it})$ is a vector function of the instruments $z_{it}$. The first-order bias correction is redundant in a finite sample if

equation[equation omitted — 188 chars of source]

Because $q_{it-1} - \hat{e}(x_{it-1})$ is the residual from the OLS estimation in step 1, equation (ref) holds if $h(z_{it})g^\prime(\hat{e}(x_{it-1}) -f(k_{it-1},v_{it-1}))$ can be expressed as a linear combination of the regressors used in step 1. This is the case if the law of motion $g$ is linear and the regressors include $h(z_{it})$. More generally, if $h(z_{it})g^\prime(\hat{e}(x_{it-1}) -f(k_{it-1},v_{it-1}))$ is well approximated by a linear combination of the regressors used in step 1, then the left-hand side of equation (ref) is almost zero and we expect the GMM estimator in step 2 to attain near-oracle performance. While our Monte Carlo exercise in Section (ref) illustrates that this can happen, we recommend basing the GMM estimator in step 2 on the modified moment condition (ref), as this ensures oracle efficiency irrespective of whether equation (ref) holds or not. Moreover, if estimation methods other than OLS are used in step 1, then equation (ref) generally does not hold and the modified moment condition (ref) improves efficiency.

Monte Carlo exercise

\paragraph{Data generating process.}

We specify the CES production function

equation*[equation* omitted — 117 chars of source]

and the disturbance $\varepsilon_{it}\sim N\left(0,\sigma^2_\varepsilon\right)$, where $\alpha\in(0,1)$ is a distributional parameter, $\rho\leq 0$ and $\sigma=\frac{1}{1-\rho}$ is the elasticity of substitution, and $\nu>0$ is returns to scale. We further specify the law of motion

equation*[equation* omitted — 155 chars of source]

and the productivity innovation $\xi_{it}\sim N\left(0,\sigma^2_\omega\right)$, where $\rho_\omega\in(0,1)$ and $\alpha_\omega\in[0,1]$. Our particular interest lies in comparing the Gaussian $AR(1)$ process for $\alpha_\omega=0$ with the nonlinear process for $\alpha_\omega=1$. To facilitate this comparison we calibrate $\mu_\omega$, $\rho_\omega$, and $\sigma^2_\omega$ to hold fixed $\text{\normalfont E}\left[\omega_{it}\right]$, $\text{\normalfont Var}\left(\omega_{it}\right)$, and $\text{\normalfont corr}\left(\omega_{it},\omega_{it-1}\right)$. Figure (ref) illustrates the resulting distribution of productivity and overlays the law of motion for the two processes.

figure[figure omitted — 346 chars of source]

We specify the CES demand

equation*[equation* omitted — 67 chars of source]

where $p_{it}$ is the output price and $\delta_{i}=(\delta_{1i},\delta_{2i})$ captures shocks to the demand the firm faces and unobserved rivals. This avoids having to specify the imperfectly competitive environment that the firm operates in. Abstracting from time-series variation for simplicity, we specify $\delta_{1i}\sim N\left(\mu_{\delta_1},\sigma^2_{\delta_1}\right)$ and $\delta_{2i}\sim N\left(\mu_{\delta_2},\sigma^2_{\delta_2}\right)$.

We denote the price of capital as $p^K_{i}$ and the price of the variable input as $p^V_{i}$. Abstracting from time-series variation, we specify $p^K_i\sim N\left(\mu_{p^K},\sigma^2_{p^K}\right)$ and $p^V_i\sim N\left(\mu_{p^V},\sigma^2_{p^V}\right)$. Assuming short-run profit maximization and equating marginal revenue with marginal cost determines $v_{it}$ along with $q^*_{it}$, $q_{it}$, and $p_{it}$. Note that because $v_{it}$ is a function of $k_{it}$, $p^V_{i}$, $\omega_{it}$, and $\delta_{i}$, the firm's demand for the variable input cannot be inverted to express $\omega_{it}$ as a function of observables: invertibility fails as long as $\sigma^2_{\delta_{1}}>0$ or $\sigma^2_{\delta_{2}}>0$.

Recall that capital $k_{it+1}$ is chosen in period $t$. We assume that capital fully depreciates between periods and endow the firm with static expectations regarding $\omega_{it+1}$. Hence, the firm chooses $k_{it+1}$ in period $t$ given $\omega_{it}$, $\delta_{i}$, $p^K_{i}$, and $p^V_{i}$ {\em as if} $k_{it+1}$ and $v_{it+1}$ are variable inputs.\footnote{Assuming constant returns to scale in production and quadratic adjustment costs to capital, the firm's dynamic programming problem that determines investment can be solved in closed form if the firm is a price-taker in the output market SYVE:01,VANB:07,ACKE:15. As this is no longer possible if the firm has market power, we opt for a different, tractable specification of the evolution of capital.}

table[table omitted — 703 chars of source]

Table (ref) shows our baseline parameterization. The elasticity of substitution is within the range of estimates in the literature CHIR:08, as are the 90:10 percentile ratio and persistence of productivity DELO:21. Short-run profit maximization implies the markup $\mu_{it}=\frac{P_{it}}{MC_{it}}=1+\exp(\delta_{2i})$, with $\text{\normalfont E}\left[\ln\mu_{it}\right]=0.25$ and $\text{\normalfont Var}\left(\ln\mu_{it}\right)=0.0126$. We simulate $S=1000$ datasets with $N=5,000$ firms and $T=20$ periods.\footnote{We use $T_0=5000$ burn-in periods.} We provide further details on the specification in Appendix (ref).

\paragraph{Estimation.}

We estimate $\text{\normalfont E}\left[q_{it}|x_{it}\right]$ in step 1 of the OP/LP/ACF procedure by OLS using the complete set of Hermite polynomials of total degree $4$ in the variables in $x_{it}$ as regressors. To illustrate Theorems (ref)--(ref), we vary the specification of $x_{it}$ as detailed below. As an alternative to OLS, we use a neural network estimator with two hidden layers of 128 neurons. We provide further details on the neural network estimator in Appendix (ref).

We estimate the parameters $\theta=\left(\alpha,\rho,\nu,\mu_\omega,\rho_\omega,\alpha_\omega\right)$ in step 2 using the complete set of Hermite polynomials of total degree $4$ in the variables in $z_{it}=\left(k_{it},k_{it-1},v_{it-1},p^V_{i}\right)$ as instruments.\footnote{We treat the price of the variable input $p^V_{i}$ as observable to focus on demand shocks as the reason for the failure of invertibility. The nonidentification result in \citeasnoun{GAND:20} assumes invertibility and therefore does not apply in our setting.} To illustrate Theorem (ref) and the advantages of Neyman orthogonality in finite samples, we alternatively base the GMM estimator on moment conditions (ref) and (ref). We provide further details on the GMM estimator in Appendix (ref).

We use three statistics to summarize the results. First, we conduct a Lagrange multiplier test for the true parameter value $\theta^0$. Our test corrects for the plug-in nature of the OP/LP/ACF procedure and clustering in the data. We provide further details on the Lagrange multiplier test in Appendix (ref).

Second, following the production function approach to markup estimation, we use the estimate of $\theta_f$ to estimate the markup $\mu_{it}=\frac{P_{it}}{MC_{it}}$ of firm $i$ in period $t$ as

equation*[equation* omitted — 128 chars of source]

The right-hand side is the log of the output elasticity minus the log of the expenditure share of the variable input.\footnote{\citeasnoun{DELO:12} isolate $\ln\mu_{it}$ on the left-hand side and use the residual from the regression in step 1 of the OP/LP/ACF procedure to estimate $\varepsilon_{it}$ on the right-hand side. However, using equation (ref), the residual is $q_{it}-\text{\normalfont E}\left[q_{it}|x_{it}\right]=\varepsilon_{it}+\zeta_{it}$. Having $\varepsilon_{it}$ on the left-hand side thus avoids that $\zeta_{it}$ taints the correlation of the estimated markup with a variable of interest such as the firm's export status or a measure of trade liberalization DORA:21.} Noting that the disturbance $\varepsilon_{it}$ averages out as $\text{\normalfont E}[\varepsilon_{it}]=0$, we refer to the average of $\ln\mu_{it}+\varepsilon_{it}$ across firms and time simply as the average log markup.

Third, because the law of motion is of primary interest in some applications AW:08,DELO:13,DORA:07, we use the estimate of $\theta_g$ to estimate

equation*[equation* omitted — 100 chars of source]

as a measure of persistence in the productivity process.

\paragraph{Results.}

We compare three cases for the observables $x_{it}$ in step 1 of the OP/LP/ACF procedure:

itemize• {\em Case 1:} $x_{it}=\left(k_{it},v_{it},p^V_{i}\right)$; • {\em Case 2:} $x_{it}=\left(k_{it+1},k_{it},v_{it},p^V_{i}\right)$; • {\em Case 3:} $x_{it}=\left(k_{it+1},k_{it},v_{it},p_{it},p^V_{i}\right)$.

Case 1 is our baseline and corresponds to how the OP/LP/ACF procedure is implemented in the existing literature. In light of Theorems (ref) and (ref), case 2 ensures $x_{it-1}\supseteq z_{it}$. Case 3 heeds Theorem (ref) and pursues a kitchen sink approach by adding the output price as a covariate that speaks to demand conditions.

sidewaysfigure\centerline \caption{Cases 1, 2, and 3 and moment condition (ref). Baseline parameterization with Gaussian $AR(1)$ process ($\alpha_\omega=0$). Case 1 (upper row), case 2 (middle row), and case 3 (lower row). $p$-value for Lagrange multiplier test (left column), average log markup (middle column), and measure of persistence $g^\prime(0)$ (right column).}

Figure (ref) shows the results of the OP/LP/ACF procedure with moment condition (ref) for the Gaussian $AR(1)$ process ($\alpha_\omega=0$). In case 1, the failure of invertibility causes a massive bias in the production function and in the average log markup derived from it. Indeed, the average log markup centers around $-0.24$ whereas short-run profit maximization implies $\ln\mu_{it}=\ln\left(1+\exp(\delta_{2i})\right)>0$. The failure of invertibility also causes a large bias in the law of motion as summarized by the measure of persistence $g^\prime(0)$. The Lagrange multiplier test for the true parameter value $\theta^0$ rejects by a wide margin. In short, the failure of invertibility can have a substantial, detrimental impact on the OP/LP/ACF procedure as implemented in the existing literature.

sidewaystable\begin{center} \begin{tabular}{p{10.5cm}|ccc|ccc} & \multicolumn{3}{c|}{average log markup} & \multicolumn{3}{c}{measure of persistence $g^\prime(0)$} \\ & bias & variance & MSE & bias & variance & MSE \\ \hline \uline{panel 1: baseline parameterization with $AR(1)$ process, moment condition (ref):} & & & & & & \\ case 1 & -0.4875 & 0.0016 & 0.2392 & -0.1869 & 0.0001 & 0.0351 \\ case 2 & 0.0023 & 0.0004 & 0.0004 & 0.0065 & 0.0004 & 0.0005 \\ case 3 & 0.0012 & 0.0004 & 0.0004 & 0.0042 & 0.0004 & 0.0004 \\ \hline \uline{panel 2: baseline parameterization with nonlinear process, moment condition (ref):} & & & & & & \\ case 1 & -0.3970 & 0.0015 & 0.1591 & -0.1524 & 0.0003 & 0.0235 \\ case 2 & 0.0511 & 0.0005 & 0.0031 & 0.0052 & 0.0003 & 0.0003 \\ case 3 & 0.0030 & 0.0004 & 0.0004 & 0.0041 & 0.0002 & 0.0002 \\ \hline \uline{panel 3: baseline parameterization with nonlinear process, moment condition (ref):} & & & & & & \\ case 2 & 0.0415 & 0.0005 & 0.0022 & -0.0073 & 0.0003 & 0.0003 \\ case 3 & 0.0037 & 0.0004 & 0.0004 & 0.0028 & 0.0002 & 0.0002 \\ \hline \uline{panel 4: modified parameterization with nonlinear process, moment condition (ref):} & & & & & & \\ case 2 & 0.5743 & 0.0012 & 0.3309 & 0.0960 & 0.0001 & 0.0093 \\ case 3 & -0.0314 & 0.0033 & 0.0043 & 0.0021 & 0.0000 & 0.0000 \\ \hline \uline{panel 5: modified parameterization with nonlinear process, moment condition (ref):} & & & & & & \\ case 2 & -0.1327 & 0.0260 & 0.0436 & -0.0098 & 0.0000 & 0.0001 \\ case 3 & -0.0031 & 0.0023 & 0.0023 & -0.0012 & 0.0000 & 0.0000 \\ \hline \uline{panel 6: baseline parameterization with nonlinear process, neural network estimator, moment condition (ref):} & & & & & & \\ case 2 & 0.0353 & 0.0202 & 0.0214 & -0.0209 & 0.0222 & 0.0227 \\ case 3 & -0.0196 & 0.0185 & 0.0189 & -0.0165 & 0.0081 & 0.0084 \\ \hline \uline{panel 7: baseline parameterization with nonlinear process, neural network estimator, moment condition (ref):} & & & & & & \\ case 2 & 0.0400 & 0.0006 & 0.0022 & -0.0054 & 0.0003 & 0.0004 \\ case 3 & 0.0065 & 0.0005 & 0.0005 & 0.0043 & 0.0002 & 0.0002 \\ \end{tabular} \end{center} \caption{Bias, variance, and mean squared error of average log markup and measure of persistence $g^\prime(0)$.}

In cases 2 and 3, the Lagrange multiplier test does not reject. As in Example (ref) in Section (ref), ensuring $x_{it-1}\supseteq z_{it}$ eliminates the bias. The average log markup centers around its true value of $0.25$ and the measure of persistence $g^\prime(0)$ centers around its true value of 0.7. The first panel of Table (ref) summarizes the results. In addition to the bias, it shows the variance and mean squared error of the average log markup and the measure of persistence $g^\prime(0)$.

sidewaysfigure\centerline \caption{Cases 1, 2, and 3 and moment condition (ref). Baseline parameterization with nonlinear process ($\alpha_\omega=1$). Case 1 (upper row), case 2 (middle row), and case 3 (lower row). $p$-value for Lagrange multiplier test (left column), average log markup (middle column), and measure of persistence $g^\prime(0)$ (right column).}

Turning to the nonlinear process ($\alpha_\omega=1$), Figure (ref) shows the results of the OP/LP/ACF procedure with moment condition (ref). In case 1, the failure of invertibility again causes a massive bias in the average log markup and the measure of persistence $g^\prime(0)$. The Lagrange multiplier test for the true parameter value $\theta^0$ rejects by a wide margin.

Ensuring $x_{it-1}\supseteq z_{it}$ greatly reduces the bias in cases 2 and 3. In case 2, the average log markup centers around $0.30$ compared to its true value of $0.25$. The measure of persistence $g^\prime(0)$ centers around its true value of $0.7$. The Lagrange multiplier test rejects by a wide margin. Adding the output price as a covariate that speaks to demand conditions further reduces---and here almost eliminates---the bias. In case 3, the average log markup centers around $0.25$ and the measure of persistence $g^\prime(0)$ centers around $0.7$. The Lagrange multiplier test rejects for just 72 of the 1000 simulated datasets at a significance level of 5%. The second panel of Table (ref) summarizes the results.

As discussed in Section (ref), explicitly including the first-order bias correction generally mitigates the adverse impact that a noisy estimate in step 1 has on the GMM estimator in step 2, although it can sometimes be redundant if OLS is used in step 1. The third panel of Table (ref) summarizes the results of replacing moment condition (ref) by moment condition (ref) for the nonlinear process ($\alpha_\omega=1$). The results are indeed similar for the two moment conditions.\footnote{To confirm, we disrupt the finite-sample relation in equation (ref) by running OLS on a separate dataset and using the estimates to construct $\hat{e}(x_{it-1})$ for the focal dataset in step 1. For the original moment condition (ref), the mean squared error of the average log markup and the measure of persistence $g^\prime(0)$ nearly doubles in case 3. In case 2, the mean squared error of the average log markup decreases slightly while the mean squared error of the measure of persistence $g^\prime(0)$ increases by an order of magnitude. In contrast, there are no meaningful changes for the modified moment condition (ref).}

However, explicitly including the first-order bias correction is not always redundant even if OLS is used in step 1. To illustrate the advantages of Neyman orthogonality in finite samples, we modify the parameterization to $\text{\normalfont E}\left[\omega_{it}\right]=-1.25$, $\text{\normalfont Var}\left(\omega_{it}\right)=2^2$, $\text{\normalfont corr}\left(\omega_{it},\omega_{it-1}\right)=0.85$, $\sigma^2_{\delta_1}=0.5^2$, $\mu_{\delta_2}=-2.5425$, $\sigma^2_{\delta_2}=2^2$, and $\sigma^2_{p^K}=2^2$. This implies an increase in the 90:10 percentile ratio and persistence of productivity and a mean-preserving spread of the log markup ($\text{\normalfont E}\left[\ln\mu_{it}\right]=0.25$ and $\text{\normalfont Var}\left(\ln\mu_{it}\right)=0.1986$).

sidewaysfigure\centerline \caption{Case 2 and moment conditions (ref) and (ref). Modified parameterization with nonlinear process ($\alpha_\omega=1$). Moment condition (ref) (upper row) and moment condition (ref) (lower row). $p$-value for Lagrange multiplier test (left column), average log markup (middle column), and measure of persistence $g^\prime(0)$ (right column).}
sidewaysfigure\centerline \caption{OP/LP/ACF procedure. Case 3 and moment conditions (ref) and (ref). Modified parameterization with nonlinear process ($\alpha_\omega=1$). Moment condition (ref) (upper row) and moment condition (ref) (lower row). $p$-value for Lagrange multiplier test (left column), average log markup (middle column), and measure of persistence $g^\prime(0)$ (right column).}

Figures (ref) and (ref) show the results of the OP/LP/ACF procedure with moment conditions (ref) and (ref) for the nonlinear process ($\alpha_\omega=1$). Figure (ref) pertains to case 2 and Figure (ref) to case 3. In case 2, there is a large bias with the original moment condition (ref). The average log markup centers around 0.82 compared to its true value of 0.25 and the measure of persistence $g^\prime(0)$ centers around 0.79 compared to its true value of 0.7. The bias is reduced with the modified moment condition (ref). The average log markup centers around 0.12 and the measure of persistence $g^\prime(0)$ centers around 0.69. In case 3, the bias is again reduced with the modified moment condition (ref). With the original moment condition (ref), the average log markup centers around 0.22 and the measure of persistence $g^\prime(0)$ centers around 0.70. The Lagrange multiplier tests rejects by a wide margin. With the modified moment condition (ref), the average log markup centers around 0.25 and the measure of persistence $g^\prime(0)$ centers around 0.69. The Lagrange multiplier test rejects for just 205 of the 1000 simulated datasets at a 5% significance level. The fourth and fifth panels of Table (ref) summarize the results.

Finally, we use a neural network estimator in step 1 as an alternative to OLS. Reverting to the baseline parameterization with the nonlinear process ($\alpha_\omega=1$), the sixth and seventh panels of Table (ref) summarize the results. Comparing the sixth to the second panel of Table (ref) shows that using the neural network estimator increases the variance and thereby the mean squared error of the GMM estimator in step 2 by an order of magnitude if we base the GMM estimator on the original moment condition (ref). To the best of our knowledge, there is no theory on using a neural network as a plug-in estimator. In contrast, we achieve much better results if we base the GMM estimator on the modified moment condition (ref). Comparing the seventh to the third panel shows that using the neural network estimator is comparable to using OLS. As discussed in Section (ref), this aligns with the theory on Neyman orthogonality in the double/debiased machine learning literature as our modification ensures that the GMM estimator in step 2 is oracle efficient regardless of the estimation method used in step 1.

Sensitivity analysis

While setting $x_{it-1}\supseteq z_{it}$ provides a first-order bias correction, higher-order biases may remain and adversely affect the estimates of $\theta=(\theta_f,\theta_g)$ in step 2 of the OP/LP/ACF procedure. In this section, we develop a diagnostic to assess the sensitivity of the estimates to biases arising from the failure of invertibility. Our diagnostic remains valid without $x_{it-1}\supseteq z_{it}$.

To simplify the analysis, we assume that the nuisance parameter $e(x_{it-1})$ equals its true value $\text{\normalfont E}[q_{it-1}|x_{it-1}]$ and abstract from estimation noise. We consider

equation[equation omitted — 182 chars of source]

where $\lambda\geq 0$ scales the lagged prediction error $\zeta_{it-1}$. If $\lambda = 0$, then moment condition (ref) equals the valid but infeasible moment condition (ref). If $\lambda = 1$, then moment condition (ref) equals the original moment condition (ref) because $\zeta_{it-1} = q^*_{it-1} - \text{\normalfont E}[q_{it-1}|x_{it-1}]$ from equation (ref). If, moreover, $x_{it-1}\supseteq z_{it}$, then it also equals the modified moment condition (ref) with $e(x_{it-1}) = \text{\normalfont E}[q_{it-1}|x_{it-1}]$, as discussed in Section (ref).

Define

equation*[equation* omitted — 159 chars of source]

where we make explicit the parameterization of the production function and law of motion. Moreover, define the pseudo-true value of $\theta$ as a function of $\lambda$ as

equation*[equation* omitted — 222 chars of source]

where the superscript $\top$ denotes the transpose, $h(z_{it})$ is a vector function of the instruments $z_{it}$, and $W$ is the GMM weighting matrix. By construction, $\theta(0)$ equals the true parameter value $\theta^0$ and $\theta(1)$ is the probability limit of the GMM estimator.

Although it is not possible to consistently estimate the bias $\theta(1)-\theta(0)$, it is possible to consistently estimate $\frac{\mathrm{d}\theta(\lambda)}{\mathrm{d}\lambda}\big|_{\lambda = 1}$ and thus to assess the sensitivity at the pseudo-true value $\theta(1)$. A small value of $\frac{\mathrm{d}\theta(\lambda)}{\mathrm{d}\lambda}\big|_{\lambda = 1}$ provides assurance that small changes in the prediction error that arises in step 1 of the OP/LP/ACF procedure if invertibility fails do not dramatically alter the pseudo-true value $\theta(1)$. Indeed, one can show that $\frac{\mathrm{d}\theta(\lambda)}{\mathrm{d}\lambda}\big|_{\lambda = 1}$ is zero in the examples in Section (ref) where no bias arises despite the failure of invertibility. Conversely, a large value of $\frac{\mathrm{d}\theta(\lambda)}{\mathrm{d}\lambda}\big|_{\lambda = 1}$ means that the pseudo-true value $\theta(1)$ is sensitive and thus warns of potentially large bias in the estimates.

To derive our diagnostic $\frac{\mathrm{d}\theta(\lambda)}{\mathrm{d}\lambda}\big|_{\lambda = 1}$, we show in Appendix (ref) that the pseudo-true value $\theta(\lambda)$ under suitable regularity conditions satisfies the first-order condition

equation[equation omitted — 219 chars of source]

where $\frac{\partialm_{it}(\theta(\lambda),\lambda)}{\partial\theta}$ is a $\dim(\theta)\times 1$ vector. The implicit function theorem implies

equation[equation omitted — 131 chars of source]

where $\Gamma$ is a $\dim(\theta)\times \dim(\theta)$ matrix and $\gamma$ a $\dim(\theta)\times 1$ vector. Using $\frac{\partialm_{it}}{\partial\theta}$, $\frac{\partialm_{it}}{\partial\lambda}$, $\frac{\partial^2 m_{it}}{\partial\theta\partial \theta^\top}$, and $\frac{\partial^2 m_{it}}{\partial\theta\partial \lambda}$ to abbreviate $\frac{\partialm_{it}(\theta(\lambda), \lambda)}{\partial\theta}|_{\lambda = 1}$, $\frac{\partialm_{it}(\theta(\lambda), \lambda)}{\partial\lambda}|_{\lambda = 1}$, $\frac{\partial^2 m_{it}(\theta(\lambda), \lambda)}{\partial\theta\partial \theta^\top} |_{\lambda = 1}$, and $\frac{\partial^2 m_{it}(\theta(\lambda), \lambda)}{\partial\theta\partial \lambda} |_{\lambda = 1}$, respectively, the matrix $\Gamma$ is defined as

align[align omitted — 634 chars of source]

where $e_l$ is a $\dim(\theta) \times 1$ vector with a one in the $l$th position and zeros elsewhere. The vector $\gamma$ is defined as

equation[equation omitted — 396 chars of source]

To evaluate our diagnostic $\frac{\mathrm{d}\theta(\lambda)}{\mathrm{d}\lambda}\big|_{\lambda = 1}$, we assume $\text{\normalfont E}[\varepsilon_{it-1}|x_{it-1}, z_{it}] = 0$. As we detail in Appendix (ref), consistently estimating $\Gamma$ is straightforward. Consistently estimating $\gamma$ is complicated by the fact that

equation*[equation* omitted — 199 chars of source]

and

equation*[equation* omitted — 650 chars of source]

depend on the lagged prediction error $\zeta_{it-1}$. The key insight is that assuming $\text{\normalfont E}[\varepsilon_{it-1}|x_{it-1}, z_{it}] = 0$ implies

equation*[equation* omitted — 164 chars of source]

Using the law of iterated expectations, we can therefore substitute $q_{it-1} - \text{\normalfont E}[q_{it-1}|x_{it-1}]$ for $\zeta_{it-1}$ in the above expressions and rewrite $\gamma$ as

equation[equation omitted — 322 chars of source]

where $\frac{\partial m_{it}}{\partial \lambda}$ and $\frac{\partial^2m_{it}}{\partial\theta\partial \lambda}$ in (ref) are replaced by

equation*[equation* omitted — 208 chars of source]

and

equation*[equation* omitted — 526 chars of source]

respectively. We can therefore consistently estimate $\gamma$ by the finite-sample analog to equation (ref) using the estimate of $\text{\normalfont E}[q_{it-1}|x_{it-1}]$ from step 1 and that of $\theta$ from step 2.

table[table omitted — 1,253 chars of source]

Table (ref) shows our diagnostic $\frac{\mathrm{d}\theta(\lambda)}{\mathrm{d}\lambda}\big|_{\lambda = 1}$ for the baseline parameterization with the nonlinear process ($\alpha_\omega=1$). We average the diagnostic over the $S=1000$ datasets from Section (ref). For comparison, we also show the bias $\theta(1)-\theta(0)$.

In case 1, the diagnostic warns of large biases in the production function parameter $\alpha$ and in the law of motion parameters $\rho_\omega$ and $\alpha_\omega$. While the diagnostic reliably captures the sign of the bias, it is an order of magnitude smaller than the bias in the production function parameter $\rho$. In case 2, the bias and the diagnostic are both smaller than in case 1, with the exception of the law of motion parameter $\alpha_\omega$. The diagnostic accurately reflects the remaining bias in the production function parameter $\rho$. However, it is an order of magnitude smaller than the bias in $\alpha$, $\rho_\omega$, and $\alpha_\omega$. In case 3, the small biases go hand-in-hand with small values of the diagnostic. Overall, while the diagnostic does not consistently estimate the bias $\theta(1)-\theta(0)$, it usefully signals the sensitivity of the estimates to a failure of invertibility.

Finally, we emphasize that our sensitivity analysis is conducted around the potentially misspecified model at $\lambda=1$ as opposed to the true model at $\lambda=0$ and remains valid regardless of the magnitude of misspecification. Our approach thus differs from the sensitivity analysis in \citeasnoun{ANDR:17} that is only valid locally around the true model.

Concluding remarks

The OP/LP/ACF procedure for production function estimation hinges on an invertibility assumption. We show that this assumption is testable and strongly rejected in widely used panel data. A failure of invertibility has important consequences: the prediction of planned output $q^*_{it}$ from observables $x_{it}$ in step 1 of the OP/LP/ACF procedure contains an error that invalidates capital $k_{it}$ as an instrument, leading to biased estimates in step 2.

Fortunately, much can still be done. We establish a necessary and sufficient condition for the moment condition in step 2 of the OP/LP/ACF procedure to hold for the true production function. This condition compels a rethinking of the OP/LP/ACF procedure: any instrument used in step 2 must be appropriately included in the regression in step 1 to ensure $x_{it-1} \supseteq z_{it}$. At a minimum, this calls for adding the lead of capital $k_{it+1}$ to $x_{it}$. Our condition further suggests to flexibly model the law of motion in step 2 and to take a kitchen sink approach to the regression in step 1 by adding as many relevant covariates as possible.

In case our condition is violated, we show that setting $x_{it-1} \supseteq z_{it}$ provides a first-order bias correction. Explicitly incorporating a bias correction achieves Neyman orthogonality in the modified moment condition. Neyman orthogonality ensures that the asymptotic distribution of the GMM estimator in step 2 is invariant to estimation noise from step 1 and is particularly advantageous if the regression in the step 1 includes a large number of covariates or if modern machine learning techniques such as neural networks and random forests are used. Monte Carlo simulations confirm that Neyman orthogonality can substantially enhance the performance of the GMM estimator in step 2.

Finally, we recognize that, despite ensuring $x_{it-1}\supseteq z_{it}$, higher-order biases may remain. To gauge their importance for the GMM estimator in step 2, we introduce a diagnostic that measures the sensitivity of the estimates to the size of the prediction error that arises in step 1 if invertibility fails.

In sum, the invertibility assumption is demanding and can fail because unobserved demand heterogeneity or in imperfectly competitive environments with partially or fully unobserved rivals or changes in firm conduct and for other reasons. We provide tools for testing the invertibility assumption and propose straightforward modifications of the OP/LP/ACF procedure that either eliminate or mitigate the bias that arises from a failure of invertibility. To address the challenges associated with including a large number of covariates into the regression in step 1, we provide a modified moment condition that achieves Neyman orthogonality and enhances efficiency and robustness. Finally, we hope our diagnostic proves valuable to researchers seeking to assess the potential biases from a failure of invertibility in their estimates.