Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
122,677 characters · 23 sections · 62 citation commands
Identification and Estimation in a Time-Varying Endogenous Random Coefficient Panel Data Model
\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long\global\long\global\long\global\long\global\long \global\long\global\long \global\long \global\long \global\long \global\long\global\long\global\long\global\long\global\long \global\long \global\long\def\mc#1{\mathscr{#1}}
\global\long\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long\def\abs#1{\left|#1\right|}
\global\long\def\norm#1{\left\Vert #1\right\Vert }
\global\long\def\rest#1{\left.#1\right|}
\global\long\def\bracket#1#2{\left\langle #1\middle\vert#2\right\rangle }
\global\long\def\sandvich#1#2#3{\left\langle #1\middle\vert#2\middle\vert#3\right\rangle }
\global\long\def\turd#1{\frac{#1}{3}}
\global\long \global\long\def\sand#1{\left\lceil #1\right\vert }
\global\long\def\wich#1{\left\vert #1\right\rfloor }
\global\long\def\sandwich#1#2#3{\left\lceil #1\middle\vert#2\middle\vert#3\right\rfloor }
\global\long\def\abs#1{\left|#1\right|}
\global\long\def\norm#1{\left\Vert #1\right\Vert }
\global\long\def\rest#1{\left.#1\right|}
\global\long\def\inprod#1{\left\langle #1\right\rangle }
\global\long\def\ol#1{\overline{#1}}
\global\long\def\ul#1{#1}
\global\long\def\td#1{\tilde{#1}} \global\long\def\bs#1{\boldsymbol{#1}}
\global\long \global\long \global\long \global\long \global\long\global\long
Correlated random coefficient (CRC) linear panel models have proven useful due to their ability to accommodate complex forms of unobserved heterogeneity that are empirically relevant (wooldridge2010econometric,hsiao2022analysis). A crucial consideration in these models is the correlation between random coefficients and regressors. Classical methods typically address this issue by allowing a time-invariant, individual-specific fixed effect---either in the form of an additive intercept or an individual-specific coefficient---to be correlated with the regressors (e.g., hausman1981panel,hsiao2008random). While convenient, this approach may not fully capture agents' optimization behavior. For example, even high-capability firms may cut inputs in some periods when input prices are flat because input choices also respond to transitory, time-varying shocks to output efficiency or productivity. Throughout the sample, I assume capability is time-invariant. Thus, both persistent capability and transitory shocks can affect realized output efficiency and may be reflected in input choices, thereby generating endogeneity that is not fully absorbed by a fixed effect.
In this paper, I aim to address this gap by proposing a new time-varying endogenous random coefficient (TERC) linear panel model, where regressors can be correlated with time-varying and individual-specific random coefficients not only through a fixed effect but, more importantly, through a time-varying random shock. Specifically, the baseline TERC model consists of two equations:
The vector of random coefficients $\beta_{it}=\beta\left(A_{i},\varepsilon_{it}\right)$ is modeled as a vector of unknown functions $\beta$ of a fixed effect $A_{i}$ and a time-varying random shock $\varepsilon_{it}$. $Y_{it}$ is then determined by the inner product between $X_{it}$ and $\beta_{it}$. In (ref), I model the vector of regressors $X_{it}$ as a vector-valued, time-varying function $g$ of the vector of exogenous instrumental variables (IV) $Z_{it}$, fixed effect $A_{i}$, and per-period information $\eta_{it}$ about $\beta_{it}$.
The motivation for (ref) is that the agent $i$ in period $t$ optimally chooses values of her regressors $X_{it}$ by solving an optimization problem (e.g., firm profit maximization), with $Z_{it},\ A_{i},\text{ and }\eta_{it}$ in her information set, leading to (ref). All of $A_{i}$, $\varepsilon_{it}$, and $\eta_{it}$ can be scalar- or vector-valued, and any pair of the three variables can be correlated. As an analyst, I observe $\left\{ X_{it},Y_{it},Z_{it}\right\} $ for $i=1,\ldots,n$ and $t=1,\ldots,T$ and aim to identify the pooled average partial effect $T^{-1}\sum_{t=1}^{T}\mathbb{E}\beta_{it}$ (APE, see graham2012identification) and the local average response function $\mathbb{E}\left[\rest{\beta_{it}}X_{it}\right]$ (LAR, see altonji2005cross).\footnote{Note that the joint distribution of $(\varepsilon_{it},\eta_{it})$ may vary over $t$. Consequently, any functionals of their distribution (e.g., $\mathbb{E}\beta_{it}$ and $\mathbb{E}\left[\rest{\beta_{it}}X_{it}\right]$) are in principle time-varying. I suppress the subscript $t$ whenever it is clear from context.}
A key feature of the TERC model is that $X_{it}$ can be correlated with $\beta_{it}$ through both $A_{i}$ and $\eta_{it}$, potentially in a complicated manner. I refer to this feature as “time-varying endogeneity through the random coefficients.” Allowing such correlation is important to applications when agent $i$ in period $t$ has information about $\beta_{it}$ through both $A_{i}$ and $\eta_{it}$ when optimally deciding $X_{it}$. In Section (ref), I present three empirical examples---heterogeneous production function estimation, a labor supply model, and demand estimation---that demonstrate the validity of this correlation (\citet*{deaton1986demand,blundell2007labor,li2024identification,aureo2024production}). The associated technical challenge is to control for the time-varying endogeneity through the random coefficients when both $A_{i}$ and $\eta_{it}$ enter $g$ in a nonlinear and nonseparable manner in (ref). Moreover, I do not impose parametric assumptions on the distributions of $A_{i}$, $\varepsilon_{it}$, or $\eta_{it}$, nor do I specify the functional form of $\beta$ or $g$. In this sense, the TERC model is a fixed-effects panel data model, and addressing these challenges requires a new method.
I propose a new panel data-based method for identifying the APE and LAR in the TERC model. The idea is to control for $A_{i}$ and $\eta_{it}$ via a sufficient statistic and a generalized residual, respectively, such that, conditional on these two controls, the residual variation in $X_{it}$ is driven solely by the exogenous IV $Z_{it}$. In other words, the residual variation in $X_{it}$ is causal and thus can be exploited to identify the parameters of interest. Specifically, I first impose an index exclusion assumption that supplies a sufficient statistic $W_{i}$ for $A_{i}$. I present justifications for this assumption from the literature. Next, I construct a feasible conditional cumulative distribution function (CDF) $V_{it}=F_{\rest{X_{it}}Z_{it},W_{i}}\left(\rest{X_{it}}Z_{it},W_{i}\right)$ to control for $\eta_{it}$. Given $(A_{i},W_{i})$, I establish a one-to-one mapping between $V_{it}$ and $\eta_{it}$. Finally, conditioning on these two controls---which effectively fixes both $A_{i}$ and $\eta_{it}$---(ref) implies that the residual variation in $X_{it}$ is driven solely by $Z_{it}$. This residual, instrument-induced variation can therefore be exploited to identify the APE and LAR via perturbation arguments and iterated expectations. At the end of Section (ref), I present three extensions: (i) allowing $g$ to depend on multiple components of $\eta_{it}$, (ii) identifying higher-order moments of $\beta_{it}$, and (iii) adding ex-post shocks to $\beta_{it}$ and $Y_{it}$, and including exogenous covariates into $X_{it}$.
I estimate the parameters of interest using a three-step series procedure. In the first step, I estimate the conditional CDF $V_{it}$, which serves as a control for $\eta_{it}$. In the second step, I exploit the linear structure of ((ref)) to estimate $\mathbb{E}[\rest{Y_{it}}X_{it},V_{it},W_{i}]$. In the third step, I use these estimates to recover the APE and LAR. When working with large datasets, the main computational challenge comes from the first-step estimation of $V_{it}$ and the second-step estimation of $\mathbb{E}[\rest{Y_{it}}X_{it},V_{it},W_{i}]$. I address these issues in Remark (ref).
Inference is less standard because these estimators are multi-step: they rely on intermediate nonparametric estimates, include derivatives of an earlier-step estimate, and then apply an additional unknown but estimable mapping to a conditional expectation. As a result, sampling error from the first and second steps can affect the limiting distribution of the final estimators, so valid standard errors must account for how uncertainty propagates across steps. I therefore derive convergence rates and asymptotic normality by explicitly tracking the contribution of each step, building on existing results for multi-step series estimators (don1991asymnormality,newey1997convergence,imbens2009identification). A discussion of how the errors from each step aggregate into the final standard error is provided at the end of Subsection (ref).
To illustrate the method, I estimate a Cobb--Douglas production function with heterogeneous output elasticities using data on Chinese manufacturing firms. The instruments are inspired by \citet*{blp1995io} and constructed as competitors\textquoteright leave-one-out weighted-average input prices within the same city and industry. The resulting APE estimates are broadly consistent with those obtained from classical constant-coefficient approaches applied to the same data (\citet*{olleypakes1996teltfp,levinsohnpetrin2003res,ackerberg2015identification}). I also estimate the LAR functions evaluated at each firm\textquoteright s realized input choices, and I find substantial cross-firm variation. I then relate heterogeneity in output elasticities to firm characteristics and obtain economically intuitive patterns. Finally, I conduct extensive robustness checks in Appendix (ref) and find that the results are stable.
To complement the empirical analysis, Appendix (ref) presents Monte Carlo simulations designed to mirror the empirical environment; the results show that the proposed estimator performs well and reinforce the empirical findings. In particular, in the baseline design ($n=1,000$, $T=2$, and 1,000 replications, using second-degree spline bases), the APE estimator exhibits small normalized bias (about 2.8--3.2% relative to the length of the true parameter) and low normalized rMSE (about 2.9--3.7%) across coefficients. Estimating the function-valued LAR is naturally less precise, but still yields normalized rMSEs on the order of 3.4--5.6% for all the coefficients. Performance improves with larger $n$ and additional time periods $T$. Robustness checks also indicate that second-degree splines outperform first- or third-degree bases, and that adding ex-post shocks increases rMSEs only slightly.
Before turning to the related literature, I note two limitations of the paper. First, the baseline model assumes each coordinate of $g$ is monotone in a single coordinate of $\eta_{it}$, which (as in imbens2009identification) rules out certain nonseparable supply-and-demand systems. Subsection (ref) discusses remedies under additional, empirically motivated assumptions. Second, the approach requires IVs and a tractable way to summarize the fixed effect $A_{i}$ via a sufficient statistic $W_{i}$; while I provide examples for the applications considered, finding credible instruments (and justifying the required structure) may be challenging in some settings.
This paper contributes to the literature on CRC linear panel models. Related work includes chamberlain1992efficiency on regular identification of first moments when $T>d_{X}$; graham2012identification on irregular identification when $T=d_{X}$ using movers/stayers; and \citet*{arellano2012identifying} on identifying first (and, under additional assumptions, higher) moments and distributions via within--between variation and restrictions on residual dependence. wooldridge2005fixed,wooldridge2009estimating develops estimation strategies for APEs---respectively under mean-independence conditions and with sample selection in unbalanced panels---while murtazashvili2008fixed study consistency of fixed-effects IV estimators in time-invariant CRC models. Additional contributions include identification of quantiles (\citet*{graham2018quantile}), random-coefficient simultaneous equations (masten2018random), and a two-step identification approach with additive fixed effects and time-invariant random coefficients (laage2024correlated).
My contribution to the CRC literature is to study a model, ((ref)) and ((ref)), with time-varying random coefficients that can be correlated with regressors through both fixed effects and time-varying shocks. The closest paper to my work is graham2012identification, who identify moments by exploiting movers and stayers under a time-stationarity restriction. This time-stationarity restriction conditions on the regressors for all periods and thus effectively rules out the time-varying endogeneity central here. I instead introduce controls for $A_{i}$ and importantly $\eta_{it}$ to identify the APE and LAR. The two approaches rely on non-nested assumptions; I view this paper as a step toward addressing time-varying endogeneity through random coefficients.
In addition to the aforementioned papers, my work is also related to the literature on: (i) time-varying parameter panel data models (e.g., \citet*{li2011non,li2016panel}), (ii) latent structure panel data models (e.g., \citet*{su2016identifying,su2019sieve,wang2024panel}), and (iii) linear panel data model with interactive fixed effects, which can be treated as imposing a specific structure on the random coefficient associated with the constant term (e.g., pesaran2006estimation,bai2009panel,moon2015linear). The modeling technique and identification method of my paper are different from these seminal papers.
The second line of research concerns the techniques used in this paper. For nonseparable panel models with generalized fixed effects, restrictions on the fixed effect are generally needed to identify global objects such as the APE and LAR (hoderlein2012nonparametric; see also \citet*{GaoLiXu2023LogicalDifferencing} and GaoLi2025SemiparametricPanelMultinomial on dealing with the issue of fixed effects in network formation and panel multinomial choice models, respectively). My method requires a sufficient statistic $W_{i}$ for $A_{i}$ and a control variable $V_{it}$ for $\eta_{it}$. To justify the sufficiency of $W_{i}$ for $A_{i}$, I adapt several techniques from the literature (e.g., \citet*{mundlak1978pooling,altonji2005cross,bester2009identification,WOOLDRIDGE2019JoE,arkhangelsky2022role}). See \citet*{liu2024identification} for a discussion on the index sufficiency condition in nonlinear semiparametric panel data models. To justify $V_{it}$ as a valid control for $\eta_{it}$, I generalize the technique of imbens2009identification in two nontrivial ways and discuss them in relation to ((ref)). More recently, Nagasawa2024NoisyConditioning studies treatment effect estimation with noisy conditioning variables using control variables from the joint distribution of noisy proxy measures.
The asymptotic analysis builds on foundational results for series and control-function estimators. don1991asymnormality and newey1997convergence provide conditions for convergence rates and asymptotic normality of series estimators, which I use to handle vector-valued functionals of regression functions. \citet*{newey1999nonparametric} establish asymptotic normality for two-step nonparametric estimators in triangular models with a separable first stage, while imbens2009identification extend the triangular framework to nonseparable first stages and derive rates and asymptotic normality for known functionals of conditional expectations. I build on imbens2009identification to obtain asymptotic normality for unknown but estimable functionals of conditional expectation functions.
Organization. The rest of the paper is organized as follows. Section (ref) formally introduces the TERC model and provides several empirical examples that fit within its framework. Section (ref) outlines the identification idea and presents the assumptions, key identification theorem, and three extensions. Section (ref) presents the series estimators for the APE and LAR and establishes their asymptotic properties. Section (ref) provides an empirical illustration. Finally, Section (ref) concludes. Proofs, additional empirical results, and a simulation study are included in Appendix (ref), (ref), and (ref), respectively.
Notation. Let $i\in\left\{ 1,\ldots,n\right\} $ index agents and $t\in\left\{ 1,\ldots,T\right\} $ index periods, with finite $T\geq2$. I use boldface upper case letters for random matrices, regular upper case letters for random variables and random vectors, and lower case letters for their values. Let ${\bf X}_{i}=\left(X_{i1},...,X_{iT}\right)'$, $Y_{i}=\left(Y_{i1},...,Y_{iT}\right)'$, ${\bf Z}_{i}=\left(Z_{i1},...,Z_{iT}\right)'$, $\varepsilon_{i}=\left(\varepsilon_{i1},...,\varepsilon_{iT}\right)'$, and $\eta_{i}=\left(\eta_{i1},...,\eta_{iT}\right)'$. I assume $\left({\bf X}_{i},{\bf Z}_{i},A_{i},\varepsilon_{i},\eta_{i}\right)$ are i.i.d. across $i$. Let $d_{X}$ be the dimension of $X_{it}$ and $X_{it,l}$ denote the $l^{\text{th}}$ coordinate of vector $X_{it}$. Similar notation is used for all other variables. I use $\coloneqq$ to define a new random variable, $\sim_{d}$ to indicate that two random variables are identically distributed, $F$ for CDF, $f$ for probability density function (PDF), $\mathbb{E}$ for expectation, $\mathbb{\mathbb{V}}$ for variance, $I_{d}$ for $d\times d$ identity matrix, $\overset{p}{\longrightarrow}$ for convergence in probability, and $\overset{d}{\longrightarrow}$ for convergence in distribution. All convergence results are stated under $n\to\infty$.
Consider the following baseline TERC model
where $\beta_{it}=\beta\left(A_{i},\varepsilon_{it}\right)\in\mathbb{R}^{d_{X}}$, the central object of interest, is a vector of random coefficients modeled as an unknown vector-valued function $\beta$ of $(A_{i},\varepsilon_{it})$. Here, $A_{i}$ is a fixed effect and $\varepsilon_{it}$ governs the time-varying behavior of $\beta_{it}$; both may have arbitrary dimension. The mapping $\beta$ itself can be time-varying, because $\varepsilon_{it}$ captures time variation in the functional form. $\eta_{it}\in\mathbb{R}^{d_{\eta}}$ is a continuously distributed, time-varying random vector that may be correlated with $A_{i}$ and $\varepsilon_{it}$; its coordinates may also be mutually correlated. This correlation can be important in applications; see the examples below. Let $X_{it}\in\mathbb{R}^{d_{X}}$ denote choice variables (e.g., capital and labor), $Y_{it}\in\mathbb{R}$ the outcome, and $Z_{it}\in\mathbb{R}^{d_{Z}}$ instruments. Finally, $g$ is an unknown vector-valued function of $(Z_{it},A_{i},\eta_{it})$ that determines each coordinate of $X_{it}$.
The TERC model is widely used in empirical studies. It extends the classic constant-coefficient linear panel model while preserving directly interpretable targets---such as the APE and LAR---that carry clear policy relevance. It also accommodates natural extensions, including higher-order moments of the APE and LAR. In what follows, I describe three applications of the TERC model.
The time-varying correlation between $X_{it}$ and $\beta_{it}$ illustrates a distinct source of endogeneity: optimization-induced selection into regressors through random coefficients. Classic work (\citet*{manski1987semiparametric,altonji2005cross,graham2012identification,chernozhukov2013average}) allows $X_{it}$ to correlate with $A_{i}$ (e.g., better-managed firms choose higher inputs). But empirically, $X_{it}$ can also co-move with $\beta_{it}$ beyond $A_{i}$, because time variation in choices is rarely fully spanned by instrument variation $Z_{it}$. For example, even with little movement in input prices, a well-managed firm may temporarily scale inputs up or down in response to private information about returns. This motivates a time-varying shock $\eta_{it}$ that affects $X_{it}$ and is informative about $\beta_{it}$.
A second departure from standard triangular models (\citet*{newey1999nonparametric,imbens2009identification}) is that $A_{i}$ enters both the outcome equation and the first-step choice equation: $X_{it}=g\left(Z_{it},A_{i},\eta_{it}\right)$, not $X_{it}=g\left(Z_{it},\eta_{it}\right)$. Consequently, residual- or CDF-based control-function approaches that rely on a one-to-one mapping between the control and $\eta_{it}$ no longer apply. Moreover, $g$ is nonseparable and can depend on $A_{i}$ nonlinearly, as implied by optimization, so demeaning or first-differencing cannot eliminate $A_{i}$. I instead use an index exclusion condition to handle the fixed effects.
The main obstacle to identifying the APE and the LAR function in the TERC model using standard linear identification is that the regressor is chosen by the agent using information that is also informative about the random coefficients. To see where the standard argument breaks down, take conditional expectations in ((ref)) given $X_{it}$:
In a constant-coefficient model, the term $\mathbb{E}[\rest{\beta_{it}}X_{it}]$ would be a constant vector and could be recovered from moment conditions. Here, however, $\mathbb{E}[\rest{\beta_{it}}X_{it}]$ is generally a function of $X_{it}$. As a result, the familiar argument based on multiplying both sides of ((ref)) by $X_{it}$, taking expectations, and inverting $\mathbb{E}[X_{it}X_{it}']$ does not identify $\mathbb{E}[\rest{\beta_{it}}X_{it}]$:
because $X_{it}X_{it}'$ and $\mathbb{E}[\beta_{it}\mid X_{it}]$ move together inside the expectation. Likewise, the “derivative/perturbation” argument for linear models does not isolate $\mathbb{E}[\rest{\beta_{it}}X_{it}]$, since differentiating ((ref)) yields
and the second term captures how selection into different $X_{it}$ changes the conditional distribution of $\beta_{it}$.
The identification strategy in this paper is to recreate, using observables, the thought experiment “vary $X_{it}$ while holding fixed everything the agent knows that is related to $\beta_{it}$.” Doing so requires two controls---one for the time-invariant heterogeneity $A_{i}$ and one for the time-varying information $\eta_{it}$.
The first control is a statistic $W_{i}=W({\bf X}_{i},{\bf Z}_{i})$ constructed from the individual\textquoteright s panel history. The maintained index exclusion condition implies that once I condition on $W_{i}$, the remaining variation in $(X_{it},Z_{it})$ for a fixed $t$ contains no additional information about $A_{i}$:
Intuitively, $W_{i}$ is chosen to summarize the persistent, individual-specific component of behavior that is linked to $A_{i}$. For example, in many panel applications a time average of $X_{it}$ (possibly together with time averages of instruments) is a natural candidate: firms with higher managerial ability or individuals with higher baseline productivity tend to exhibit persistently different choices, and $W_{i}$ is designed to capture that persistent component.
Even after accounting for $A_{i}$, endogeneity remains because agents may respond to time-varying information $\eta_{it}$ when choosing $X_{it}$. The second control is constructed as the conditional CDF (a “rank” or \textquotedblleft generalized residual\textquotedblright ):
The interpretation is that $V_{it}$ records where the realized $X_{it}$ sits in the distribution of $X_{it}$ among agents with the same $(Z_{it},W_{i})$. Since $A_{i}$ and $Z_{it}$ appear in the first-step equation ((ref)) and $W_{i}$ controls for $A_{i}$ by ((ref)), the residual variation of $X_{it}$ given $(Z_{it},W_{i})$ is only driven by $\eta_{it}$ which is assumed to enter ((ref)) strictly monotonically. For example, consider two firms with the same observed average inputs over all periods ($W_{i}$) and input prices in period $t$ ($Z_{it}$). If firm 1's choice of capital ($X_{it}$) in this period is higher than that of firm 2, it can be inferred that firm 1 receives a larger productivity shock ($\eta_{it}$) to its capital choice function in period $t$ than firm 2 since ((ref)) ensures that the residual variation in $A_{i}$ given $W_{i}$ is not informative about the choice of $X_{it}$.
As a result, $V_{it}$ records where the realized $\eta_{it}$ sits in the distribution of $\eta_{it}$ among agents with the same $(Z_{it},A_{i},W_{i})$. By an exogeneity assumption between $Z_{it}$ and $\eta_{it}$, it drops out from the conditioning set and hence $V_{it}$ controls for $\eta_{it}$ given $(A_{i},W_{i})$. Thus, conditioning on $(V_{it},W_{i})$ is intended to hold fixed both the individual\textquoteright s persistent type (through $W_{i}$) and the individual\textquoteright s period-$t$ information that drives endogenous adjustment (through $V_{it}$).
This two-control structure matches the economics of the motivating examples in Section (ref). In the production function estimation example, $W_{i}$ captures persistent firm heterogeneity (e.g., management skill) that affects both average input choices and average elasticities, while $V_{it}$ captures time-$t$ private information or expectations about technology that shift input choices and are correlated with realized productivity shocks. In the labor supply application, $W_{i}$ summarizes persistent ability/tastes and $V_{it}$ captures time-varying job-match information that affects wage choice. In the AIDS application, $W_{i}$ captures persistent household heterogeneity and $V_{it}$ captures time-varying wealth or liquidity information that affects expenditure plans.
Given the two controls $V_{it}$ and $W_{i}$, the central step to identification of the APE and LAR is to show that once I condition on the two controls, $X_{it}$ no longer carries additional information about $\beta_{it}$. The argument proceeds in two layers.
The first step tackles the time-varying part of the endogeneity by assuming $A_{i}$ is known and showing
Specifically, since $V_{it}$ records where the realized $\eta_{it}$ sits in the distribution of $\eta_{it}$ among agents with the same $(A_{i},W_{i})$, holding $(A_{i},V_{it},W_{i})$ fixed amounts to holding $(A_{i},\eta_{it})$ fixed. In this step, Assumption (ref) (Componentwise Monotonicity) is used to invert $\eta_{it}$ out from function $g$, Assumption (ref) (Index Exclusion) allows one to add $A_{i}$ into the conditioning set of the CDF $V_{it}$, and Assumption (ref)(ref) (Exogeneity of IV) enables one to exclude $Z_{it}$ from the conditioning set of the CDF $V_{it}$. Assumption (ref)(ref) (Monotonicity of the Conditional CDF) then implies that $V_{it}$ is a one-to-one transformation of $\eta_{it}$ given $(A_{i},W_{i})$, proving the claim that $(A_{i},V_{it},W_{i})$ uniquely pins down $(A_{i},\eta_{it},W_{i})$. Hence, equation ((ref)) implies that the remaining variation in $X_{it}$ is generated only by $Z_{it}$. Under the maintained IV exogeneity condition, this residual variation is independent of $\beta_{it}$. This is the point at which the time-varying part of the endogeneity problem is resolved: within cells defined by $(A_{i},V_{it},W_{i})$, movements in $X_{it}$ are \textquotedblleft as good as randomized\textquotedblright by $Z_{it}$. In the production function estimation example, if two firms have the same management ability ($A_{i}$) and expected per-period technology shock ($\eta_{it}$) but selects different capital or labor, such difference can only be explained by the variation in exogenous $Z_{it}$, which is causal and thus can be exploited to identify the APE and LAR.
The second step tackles the time-invariant part of the endogeneity. Since $A_{i}$ is not observed, it is not feasible to condition on it in ((ref)). I deal with unknown $A_{i}$ by integrating it out. This involves iterating expectations and the index exclusion restriction of $W_{i}$. Specifically, I show
Notice that on the left-hand side of ((ref)), $X_{it}$ affects the conditional expectation only through the conditional density $f_{\rest{A_{i}}X_{it},V_{it},W_{i}}$. The index exclusion restriction (Assumption (ref)) implies that, conditional on $W_{i}$, $A_{i}$ is independent of $(X_{it},Z_{it})$ and thus of $V_{it}$. For instance, in the production function estimation example, $W_{i}$ captures the firm's time-invariant management ability $A_{i}$, so once $W_{i}$ is held fixed, period-$t$ inputs $X_{it}$ and control for productivity shock $V_{it}$ only reflect transitory variation and add no extra information about $A_{i}$. Therefore, I have
which ensures that ((ref)) holds.
Combining ((ref)) and ((ref)) gives
Equation ((ref)) is the key identification result: after conditioning on $(V_{it},W_{i})$, the dependence of $\beta_{it}$ on $X_{it}$ is removed.
Given ((ref)), equation ((ref)) implies a conditional linear representation
With sufficient residual variation in $X_{it}$ conditional on $(V_{it},W_{i})$, $\mathbb{E}[\rest{\beta_{it}}V_{it},W_{i}]$ can be identified either from conditional second moments
or from the derivative of the conditional mean
Finally, identification of the APE and LAR follows by the law of iterated expectations
In the production function estimation example, $\ol b$ summarizes output elasticities by averaging across firms and time, so it plays the same role as the constant coefficient in standard linear panel regressions. $b_{t}(x)$ denotes the period-$t$ average elasticity conditional on operating at input level $x$, a policy-relevant local parameter when elasticities are systematically related to firms\textquoteright input choices. Analogous interpretations apply to labor supply (average and local wage elasticities) and demand estimation (average and local expenditure/price elasticities). The purpose of $(W_{i},V_{it})$ is precisely to make these interpretations credible by ensuring that the remaining variation used to identify ((ref)) is driven by the exogenous instruments rather than by unobserved, coefficient-relevant information.
Next, I provide a list of assumptions on the model primitives required for the identification analysis and discuss them in relation to model ((ref))--((ref)).
The role of Assumption (ref) is to deliver the invertibility/control-function step used in the identification argument. This assumption is mild when $\eta_{it}$ is scalar: it is automatically satisfied in additive triangular specifications (e.g., \citet*{newey1999nonparametric}), and in the nonseparable scalar case it is essentially the same as the monotonicity condition in imbens2009identification, except that I also allow for time-invariant heterogeneity $A_{i}$. When $d_{X}>1$, as in production-function estimation (\citet*{olleypakes1996teltfp,levinsohnpetrin2003res,ackerberg2015identification}), one can further weaken the requirement to monotonicity in only one endogenous input, mirroring the identifying restriction in those classic approaches.
Assumption (ref) requires justification when $\eta_{it}$ is vector-valued and each coordinate of $X_{it}$ admits one distinct coordinate of $\eta_{it}$. In the production function estimation example, this is plausible if firms make hiring and investment decisions in different units (e.g., HR and finance), each responding primarily to its own market-specific signal---labor-market conditions for hiring and financial conditions for investment---satisfying the \textquotedblleft single-coordinate\textquotedblright dependence. Moreover, positive shocks in each market naturally raise the corresponding choice (more hiring when labor conditions improve, more investment when financing conditions improve), implying that $g_{j}$ is strictly monotone in $\eta_{it,j}$ for each coordinate $j$ and thus satisfying componentwise monotonicity.
Importantly, Assumption (ref) places no restrictions on the dependence structure across coordinates of $\eta_{it}$, nor on the dependence between $A_{i}$ and $\eta_{it}$. Accordingly, without loss of generality, one may normalize $\eta_{it}$ so that $d_{\eta}=d_{X}$ and each coordinate of $g$ depends on the corresponding coordinate of $\eta_{it}$.
Assumption (ref) serves two roles. First, it makes $W_{i}$ a sufficient statistic for $A_{i}$, which is used to establish ((ref)). Second, it is used to construct $V_{it}$ and establish its connection with $\eta_{it}$ by holding $A_{i}$ fixed. Under this assumption, the channel of endogeneity operating through $A_{i}$ is handled entirely via $W_{i}$, while the remaining “time-varying endogeneity through the random coefficients” is driven by the relationship between $\eta_{it}$ and $\varepsilon_{it}$ and is addressed via $V_{it}$. Assumption (ref) is similar to Assumption 2.1 of altonji2005cross and Assumption A2 of \citet*{liu2024identification}. Considering the heterogeneity and endogeneity afforded in the TERC model, some restriction that isolates $A_{i}$ is typically needed.
To motivate $W_{i}$ in empirical settings, consider steady-state input policies with convex adjustment costs in the production application. In standard investment and hiring models with convex (e.g., quadratic) adjustment costs (cooper2006nature), optimality implies convergence to a firm-specific long-run target $X_{i}^{\star}(A_{i})$ that is increasing in ability $A_{i}$. Observed inputs fluctuate around this target due to transitory shocks and gradual adjustment. Averaging $X_{it}$ over time filters out these high-frequency deviations and provides a consistent estimator of $X_{i}^{\star}(A_{i})$. Because $X_{i}^{\star}(\cdot)$ is monotone, conditioning on $\ol X_{i}=T^{-1}\sum_{t=1}^{T}X_{it}$ induces a control for $A_{i}$, satisfying Assumption (ref).
Next, I provide three sufficient conditions for Assumption (ref).
First, Assumption (ref) can be justified by nonparametrically generalizing equation (2.4) of mundlak1978pooling to be $A_{i}=h\left(\ol X_{i},\nu_{i}\right)$, where $\nu_{i}\perp\left(X_{it},Z_{it}\right)$ for all $t$. Here, the functional form of $h$ and distribution of $\nu_{i}$ can be unknown. Then, Assumption (ref) is satisfied by taking $W_{i}$ to be $\ol X_{i}$. This is true because conditioning on $W_{i}$, any $t$-specific $X_{it}$ and $Z_{it}$ does not affect the distribution of $A_{i}$ as the residual randomness in $A_{i}$ is only driven by $\nu_{i}$.
Second, insights from treatment assignment models in panel data (e.g., arkhangelsky2022role) can be exploited to find candidate $W_{i}$. For example, if $(X_{it},Z_{it})$ conditional on $A_{i}$ is a two-dimensional normal random vector i.i.d. through time with known covariance and ${\bf Z}_{i}\perp A_{i}$, then the density of $(X_{it},Z_{it})$ given $A_{i}$ depends on $A_{i}$ only through $\ol X_{i}$. Hence, by the sufficient statistic for bi-variate normal random vectors with known covariance matrix (\citet*[ch. 7]{hogg2019introduction}), $W_{i}=(\ol X_{i},\ol Z_{i})$ suffices for Assumption (ref).
Third, the nonparametric exchangeability condition (e.g., altonji2005cross) can be adapted to justify Assumption (ref). I present its details in Proposition (ref) of Appendix (ref). The key idea is time exchangeability: conditional on $A_{i}$, the joint behavior of $(X_{it},Z_{it})$ is unchanged if I relabel periods, so the time order contains no information about $A_{i}$. Hence only order-invariant summaries of the panel matter for learning about $A_{i}$, and symmetric functions of the observed pairs $(X_{it},Z_{it})$ can be used as $W_{i}$ to satisfy Assumption (ref).
$\text{\ }$
Next, I control for the time-varying $\eta_{it}$. When $A_{i}$ is not present in ((ref)) and under Assumption (ref), imbens2009identification propose
as a control for $\eta_{it}$. The intuition is that, holding input prices $Z_{it}$ fixed, a firm choosing a higher capital $X_{it}$ must be experiencing a higher productivity shock $\eta_{it}$; otherwise the higher input choice would not be optimal. However, this argument breaks down in the TERC model because the unobserved, time-invariant component $A_{i}$ also shifts input choices: firm 1\textquoteright s higher $X_{it}$ could reflect greater managerial ability rather than a larger transitory shock $\eta_{it}$. Therefore $A_{i}$ must also be conditioned on, but this makes $F_{\rest{X_{it}}A_{i},Z_{it}}\left(\rest{X_{it}}A_{i},Z_{it}\right)$ infeasible since $A_{i}$ is unobserved.
I proceed in two steps to generalize the argument of imbens2009identification, adapting it to the TERC model. First, to make the control for $\eta_{it}$ feasible, I replace $A_{i}$ with the observable proxy $W_{i}$ and define
Then, I show in ((ref)) that, under the next Assumption, $V_{it}$ controls for $\eta_{it}$ conditional on $A_{i}$ and $W_{i}$. Second, I further extend their analysis to allow coordinates of $g$ to depend on multiple elements of $\eta_{it}$ in Subsection (ref).
Assumption (ref) enables construction of a feasible control variable $V_{it}$ for $\eta_{it}$ given $A_{i}$ and $W_{i}$. To present it clearly, suppose $d_{X}=1$ and recall that $V_{it}=F_{\rest{X_{it}}Z_{it},W_{i}}\left(\rest{X_{it}}Z_{it},W_{i}\right)$. Then,
where the first equality holds by Assumption (ref), the second equality holds by Assumption (ref) and a change of variable argument, and the last equality holds by Assumption (ref)(ref). Thus, $V_{it}$ uniquely determines $\eta_{it}$ given $A_{i}$ and $W_{i}$ by Assumption (ref)(ref). Note that ((ref)) relies on Assumption (ref); when multiple elements of $\eta_{it}$ appear in one coordinate of $X_{it}$, the one-to-one relationship does not hold in general, and additional structure or assumption is needed to recover ((ref)). I discuss them in Subsection (ref).
Assumption (ref)(ref) is similar to condition (i) of Theorem 1 in imbens2009identification, except that I further condition on $(A_{i},W_{i})$. Since $W_{i}$ can be viewed as summarizing all the time-invariant information about $A_{i}$ in the data, Assumption (ref)(ref), loosely speaking, requires $\rest{Z_{it}\perp\left(\varepsilon_{it},\eta_{it}\right)}A_{i}$, which is already implied by the standard exogeneity condition $Z_{it}\perp\left(A_{i},\varepsilon_{it},\eta_{it}\right)$. Note that $Z_{it}\perp\left(A_{i},\varepsilon_{it},\eta_{it}\right)$ is stronger than Assumption (ref)(ref) as the former rules out the possibility that $Z_{it}$ and $A_{i}$ are correlated.\footnote{For instance, Bartik instruments use national industry shocks, weighted by pre-existing local industry shares. The intuition is the national shocks are plausibly \textquotedblleft as-good-as-random\textquotedblright conditional on the fixed effects and controls, so the conditional exogeneity is plausible. But the baseline shares can reflect deep local traits---so unconditional exogeneity is harder to defend.} Nonetheless, the unconditional exogeneity condition remains a useful benchmark since it is widely used in applied work.
Assumption (ref)(ref) can be justified in many applications with the usual choice of IVs from the literature. For example, in labor-supply applications, $Z_{it}$ can be county minimum-wage changes or EITC expansions---policy-driven shocks that, conditional on time-invariant ability $A_{i}$ and its sufficient statistic $W_{i}$, vary independently of expected sectoral labor-demand shocks $\eta_{it}$ and labor-supply-elasticity shocks $\varepsilon_{it}$. In gasoline-demand estimation, $Z_{it}$ can be the head of household\textquoteright s earned income or gas-tax changes, assumed (conditional on a household fixed effect $A_{i}$ and $W_{i}$) to be independent of anticipated near-term income changes $\eta_{it}$ (e.g., an upcoming bonus) and transitory wealth shocks $\varepsilon_{it}$.
For the empirical application of Section (ref), I construct the instruments toward BLP-style (\citet*{blp1995io}) cost shifters: leave-one-out, competitor-weighted input-price (wages and interest rates) averages at the industry--city--year level. These instruments are relevant because firms compete locally for labor and capital, so competitors' input prices shift firm $i$'s cost environment and input choices. They are plausibly exogenous to firm $i$\textquoteright s idiosyncratic productivity shocks conditional on the fixed effect and its control variable because firm $i$\textquoteright s own input prices are excluded from the construction, so identification comes from competitors' input-price variation rather than firm $i$'s unobserved shocks. Robustness checks in Appendix (ref) provide evidence that these IVs are plausibly valid and satisfy Assumption (ref)(ref).
Assumption (ref)(ref) is analogous to condition (ii) of Theorem 1 in imbens2009identification. It is mild because it concerns a conditional CDF, which is typically strictly increasing for continuous random variables.
The last assumption concerns the residual variation in $X_{it}$ given $V_{it}$ and $W_{i}$, which is used to identify the APE and LAR by ((ref)).
Assumption (ref) ensures that $\mathbb{E}\left[\rest{X_{it}X_{it}'}V_{it},W_{i}\right]$ in ((ref)) is invertible. Thus, the APE and LAR can be identified by ((ref)) and the law of iterated expectations.
Assumption (ref) is best interpreted as a conditional support (non-degeneracy) requirement: for each $t$, $X_{it}$ must retain sufficient variation conditional on $(V_{it},W_{i})$ to identify the target objects. Requiring residual variation in $X_{it}$ conditional on $V_{it}$ is generally not restrictive (imbens2009identification). Essentially, it requires $Z_{it}$ to induce enough variation in $X_{it}$ so that, for each $v\in\text{Supp}(V_{it})$, the set $\left\{ X_{it}:V_{it}=v\right\} $ is not a singleton. Therefore, it corresponds to the classical IV relevance condition of $\partial g(z,a,\eta)/\partial z\neq0$.
The restriction on the support of $X_{it}$ in Assumption (ref) mainly comes from conditioning on $W_{i}$, which also concerns Assumption (ref). There is an inherent trade-off between Assumptions (ref) and (ref). To see this, suppose $W_{i}=(\mathbf{X}_{i},\mathbf{Z}_{i})$ which satisfies Assumption (ref) trivially. However, because $X_{it}$ is included in $W_{i}$, it makes the support of $X_{it}$ given $W_{i}$ a singleton. Hence, Assumption (ref) does not hold unless $X_{it}$ is a scalar and the point in the support of $X_{it}$ given $W_{i}$ is nonzero. More generally, including more elements in $W_{i}$ tends to make Assumption (ref) easier to satisfy while making Assumption (ref) harder to maintain. Because of this trade-off, I describe next how the conditions used to justify Assumption (ref) affect the plausibility of Assumption (ref), and give corresponding sufficient conditions.
When justifications for Assumption (ref) are used so that $W_{i}$ only includes averages of $X_{it}$ and/or $Z_{it}$ through time (e.g., \citet*{mundlak1978pooling,arkhangelsky2022role,liu2024identification}), Assumption (ref) is not restrictive. For instance, with $d_{X}=2$, $T=2$, and $W_{i}=\ol X_{i}$, Assumption (ref) holds whenever $X_{i1}$ and $X_{i2}$ are not perfectly linearly dependent, so that $\mathrm{Supp}(X_{it}\mid\ol X_{i})$ remains non-degenerate. By ((ref)) and the relevance condition above, it in turn requires that the serial dependence in $Z_{it}$ cannot be so high that the value of $Z_{it}$ (hence $X_{it}$) for any $t$ is uniquely determined once $W_{i}$ is fixed.
By contrast, when a high-dimensional $W_{i}$ is used to justify Assumption (ref) (e.g., via the exchangeability condition of altonji2005cross), Assumption (ref) becomes more restrictive. Intuitively, a high-dimensional $W_{i}$ imposes many constraints on the path $(X_{i1},\ldots,X_{iT})$, so one typically needs longer panel to leave enough feasible configurations and satisfy Assumption (ref).
In this case, one possibility is to explore the symmetry through time in the solution to $\left\{ {\bf X}_{i}:V_{it}=v,W_{i}=w\right\} $. In particular, when $W_{i}$ includes averages through time of the polynomials of $\left(X_{it},Z_{it}\right)$ up to the $T^{\text{th}}$ order, Assumption (ref) is satisfied under a nonparametric exchangeability condition on $f_{\rest{\eta_{i}}A_{i}}$ by adapting the proof of altonji2005cross in Proposition (ref). Then, since $W_{i}$ is symmetric in $\left(X_{it},Z_{it}\right)$ through time, one may permute the order of $\left(X_{it},Z_{it}\right)$ in time without changing the value of $W_{i}$. This creates enough linearly independent points in the support of $X_{it}$ given $W_{i}$ when $T\geq d_{X}$. When such permutation also does not change the value of $V_{it}$ which is usually satisfied under the relevance condition, Assumption (ref) is satisfied. I provide an example in Appendix (ref).
Assumption (ref) pertains to the inversion-based argument in ((ref)). To facilitate the derivative-based argument in ((ref)), I introduce the next assumption, which requires more residual variation of $X_{it}$ conditional on $V_{it}$ and $W_{i}$ than Assumption (ref).
\addtocounter{assumption}{-1} \begingroup \let\oldtheassumption\oldtheassumption\ensuremath{'} \let\oldtheHassumption\oldtheHassumption-p
\endgroup
Assumption (ref) allows one to perturb $X_{it}$ in $\mathbb{E}[\rest{Y_{it}}X_{it},V_{it},W_{i}]$ when $(V_{it},W_{i})$ is fixed so that conditional moments of $\beta_{it}$ can be recovered by ((ref)). It is similar to Assumption 2.2 of altonji2005cross.
The same trade-off between Assumptions (ref) and (ref) also applies to Assumption (ref). When Assumption (ref) is justified such that $W_{i}$ is low-dimensional (e.g., $\ol X_{i}$ as in mundlak1978pooling), Assumption (ref) is generally not restrictive. In essence, $Z_{it}$ must vary enough to generate sufficient residual variation in $X_{it}$ conditional on $(V_{it},W_{i})$.
However, when Assumption (ref) is justified using nonparametric arguments (e.g., the exchangeability condition of altonji2005cross), $d_{W}$ is large and Assumption (ref) can be restrictive. To address this issue, I present two sets of solutions in Appendix (ref). First, the unconditional variation of exogenous regressors outside of $W_{i}$ can be leveraged. Second, following the suggestions of altonji2005cross, one may restrict how $W_{i}$ enters $f_{\rest{A_{i}}W_{i}}$ so that the relevant conditional expectations only concern a subset of $W_{i}$ or a linear index of its components.
Assumption (ref) allows the straightforward perturbation-based method ((ref)) to identify the APE and LAR without requiring the computation of the inverse of $\mathbb{E}\left[\rest{X_{it}X_{it}'}V_{it},W_{i}\right]$. When residual variation is not a concern---such as when $W_{i}$ contains only the mean of $X_{it}$ through time or when exogenous regressors are included in ((ref))---Assumption (ref) is preferred for simpler analysis. I compare these two approaches in Remark (ref) below.
$\text{\ }$
The next theorem summarizes the main identification result of this paper.
I present four methods for incorporating vector-valued $\eta_{it}$ into $X_{it}$ under additional assumptions. To illustrate the idea clearly, I assume $d_{X}=d_{\eta}=2$.
First, timing assumptions on the choice of regressors can be exploited. For instance, in production function estimation, choices of later inputs such as capital typically do not depend on random shocks to earlier chosen inputs such as labor after conditioning on the level of the earlier chosen inputs; see \citet*[(Appendix 1 of 2006 Working Paper)]{ackerberg2015identification} for a similar idea. Under this timing assumption, I modify ((ref))--((ref)) to be:
Note that now $X_{it,2}$ depends on the whole vector of $\eta_{it}$ if $X_{it,1}=g_{1}\left(Z_{it},A_{i},\eta_{it,1}\right)$ is substituted into $g_{2}$. I first use $V_{it,2}\coloneqq F_{\rest{X_{it,2}}X_{it,1},Z_{it},W_{i}}\left(\rest{X_{it,2}}X_{it,1},Z_{it},W_{i}\right)$ to control for $\eta_{it,2}$ first. Then, I use $F_{\rest{X_{it,1}}Z_{it},W_{i}}\left(\rest{X_{it,1}}Z_{it},W_{i}\right)$ to control for $\eta_{it,1}$.
Second, interaction across different coordinates of $X_{it}$ can provide useful information about $\eta_{it}$. One possibility is to assume that, although the full vector $\eta_{it}$ enters each coordinate of $X_{it}$, the ratio $X_{it,1}/X_{it,2}$ depends only on a single component of $\eta_{it}$. For example, in production function estimation, $\eta_{it,1}$ represents the Hicks neutral productivity and $\eta_{it,2}$ denotes the labor-augmenting technology. Although both capital and labor choices depend on vector $\eta_{it}$, it is plausible that capital per worker depends only on the labor-augmenting technology $\eta_{it,2}$. See Proposition 2.1 of demirer2020production for more discussions about this assumption. Given this assumption, I can proceed to control for $\eta_{it,2}$ first by
and then for $\eta_{it,1}$ by
Third, general “global univalence” results---such as gale1965jacobian or palais1959natural,hadamard1906b,hadamard1906a---suffice to establish the invertibility required here. In production function estimation, the inverse isotonicity approach in \citet*{berry2013connected} appears promising to obtain global injectivity of $g$ in $\eta_{it}$. In particular, suppose each input choice increases in its own component of $\eta_{it}$, while cross-partials are weakly negative due to substitutability. This happens when, for example, inputs are substitutes in the firm\textquoteright s optimal input demand; a favorable shock to one input in the form of higher efficiency or lower effective cost leads the firm to substitute toward that input and away from others. Then, the Jacobian matrix of $g$ with respect to $\eta_{it}$, denoted by $J_{g}(\eta_{it})$, has positive diagonal entries and nonpositive off-diagonal entries, i.e., the $Z$-matrix sign pattern required for an $M$-matrix. If, in addition, the own effects are strong enough that the matrix is strictly diagonally dominant, then $J_{g}(\eta_{it})$ is an $M$-matrix. This, in turn, guarantees a unique monotone inverse mapping from observables back to the vector $\eta_{it}$.
Finally, if $\eta_{it}$ enters the structural function $g$ through a scalar index $\eta_{it}^{\prime}\tau$ and $g$ is monotone in that index (ichimura1991semiparametric), then the effective unobservable is scalar, satisfying Assumption (ref). This aligns with common practice in asset pricing, where a scalar market factor emerges as an index of aggregate shocks and affects stock returns monotonically. A similar structure arises in production function estimation, where scalar productivity summarizes richer underlying states and affects input choices strictly increasingly.
I identify the $P^{\text{th}}$-order moments of $\beta_{it}$. For clarity of illustration, suppose the regressor is $X_{it}=(K_{it},L_{it})'$ and $\beta_{it}=(\beta_{it,K},\beta_{it,L})'$. The extension to $d_{X}>2$ is straightforward. Since the residual variation in $X_{it}$ given $(V_{it},W_{i})$ is driven only by exogenous $Z_{it}$, ((ref)) holds for any measurable function $h$ of $\beta_{it}$:
In particular, let $h=(h_{1},\ldots,h_{P})'$ and $h_{q}(\beta_{it})=\beta_{it,K}^{q}\beta_{it,L}^{P-q}$ for $q=0,\ldots,P$. By ((ref)), the conditional $P^{\text{th}}$-order moment of $Y_{it}$ given $(X_{it},V_{it},W_{i})$ is
Under $P$-times differentiability of $\mathbb{E}\left[Y_{it}^{P}\mid X_{it},V_{it},W_{i}\right]$ in $X_{it}$, the coefficients on $K_{it}^{P}$ and $L_{it}^{P}$ are identified via differentiation, yielding
and hence the $P^{\text{th}}$-order moments of $\beta_{it}$ follow by the law of iterated expectations
Repeating this argument for all $P\in\mathbb{N}$ identifies the entire moment sequence of $\beta_{it}$. Under standard moment-determinacy conditions (e.g., stoyanov2000krein), this sequence uniquely determines the distribution of $\beta_{it}$.
The identification argument goes through when I include ex-post shocks $\upsilon_{it}$ to $\beta_{it}$ (i.e., $\beta_{it}=\beta\left(A_{i},\varepsilon_{it},\upsilon_{it}\right)$) and $\widetilde{\epsilon}_{it}$ to $Y_{it}$ (i.e., $Y_{it}=X_{it}'\beta_{it}+\widetilde{\epsilon}_{it}$ with $\mathbb{E}\widetilde{\epsilon}_{it}=0$), where $\upsilon_{it}$ and $\widetilde{\epsilon}_{it}$ are independent of all other variables. Note that $\beta(\cdot)$ may a priori be time-varying and the joint distribution of $(\varepsilon_{it},\eta_{it})$ may also vary with $t$. Introducing ex-post shocks $\upsilon_{it}$ provides an additional and economically distinct source of time variation (e.g., unexpected technology shock). While the identification result remains valid with the inclusion of $\upsilon_{it}$, I present the required changes to the proof at the end of the proof of Theorem (ref). I also include $\upsilon_{it}$ into $\beta_{it}$ when estimating the APE and LAR and examine its impact via simulations. As for $\widetilde{\epsilon}_{it}$, since it is independent of all other variables and enters ((ref)) in an additive way, equations ((ref))--((ref)) hold as before. Thus, the identification analysis is unaffected.
The presence of exogenous covariates in $X_{it}$ can strengthen identification. The key point is that Assumptions (ref)--(ref) pertain only to the endogenous covariates of $X_{it}$. In particular, the control variables $V_{it}$ and $W_{i}$ depend only on the endogenous covariates of $X_{it}$, which alleviates concerns that the controls may be high-dimensional and makes the required residual variation in $X_{it}$ easier to satisfy. Moreover, as discussed in Appendix (ref), unconditional variation in exogenous covariates of $X_{it}$ can be directly exploited to make Assumptions (ref) and (ref) easier to satisfy. Details on how to adapt the analysis to incorporate exogenous covariates are provided at the end of the proof of Theorem (ref).
In this section, I first describe how to estimate the parameters using three-step series estimators. I then establish the convergence rates of the proposed estimators, and finally show that they are asymptotically normal.
The parameters of interest are
$\ol b$ is the pooled APE across all firms and periods. $b_{t}\left(x\right)$ is the LAR function for a subpopulation characterized by $X_{it}=x$ in period $t$ and is also useful for answering policy-related questions. For example, plugging realizations $x_{it}$ of $X_{it}$ into $b_{t}\left(x\right)$ provides a fine approximation to $\beta_{it}$.
I propose to estimate the parameters in ((ref)) with three-step series estimators. For clarity of exposition, I set $d_{X}=1$ and highlight the modifications required for $d_{X}>1$ when needed. To fix ideas, I follow \citet*{liu2024identification} to let $W_{i}=T^{-1}\sum_{t=1}^{T}X_{it}$, which can be motivated by generalizing the method of mundlak1978pooling.
First, for each $t$, I estimate $V_{t}\left(x,z,w\right)\coloneqq F_{\rest{X_{it}}Z_{it},W_{i}}\left(\rest xz,w\right)$ by regressing $\mathbf{\mathbbm1}\left\{ X_{it}\leq x\right\} $ on the basis functions $q^{M_{1}}$ of $\left(Z_{it},W_{i}\right)$ with trimming function $\tau$:
where $\widehat{Q}_{t}:=n^{-1}\sum_{i=1}^{n}q_{it}q_{it}'$ and $q_{it}:=q^{M_{1}}\left(z_{it},w_{i}\right)$. Examples of $q^{M_{1}}$ include power series and spline functions. When $d_{X}>1$, I regress $\mathbf{\mathbbm1}\left\{ X_{it,l}\leq x_{l}\right\} $ on the basis functions $q^{M_{1}}$ of $\left(Z_{it},W_{i}\right)$ with trimming function $\tau$ for each $l$ and obtain $\widehat{V}_{t}\left(x,z,w\right)=\left(\widehat{V}_{t,1}\left(x,z,w\right),\ldots,\widehat{V}_{t,d_{X}}\left(x,z,w\right)\right)'$. I highlight two properties of $\widehat{V}_{t}\left(x,z,w\right)$. First, the regression coefficient $\widehat{\gamma}_{t}^{M_{1}}\left(x\right)$ in (ref) depends on $x$ because the dependent variable $\mathbf{\mathbbm1}\left\{ X_{it}\leq x\right\} $ is a function of $x$. This fact causes the convergence rate of $\widehat{V}_{t}\left(x,z,w\right)$ to be slower than the standard rates for series estimators (imbens2009identification). Second, a trimming function $\tau$ is needed since I am estimating a conditional CDF. An example of $\tau$ is $\tau\left(x\right)=\mathbf{\mathbbm1}\left\{ x>0\right\} \cdot\min\left(x,1\right)$.
Next, define $S\coloneqq\left(X,V,W\right)$ and let $\mathcal{X},\ \mathcal{Z},\ \mathcal{V},\ \mathcal{W},$ and $\mathcal{S}$ denote the supports of $X,\ Z,\ V,\ W,$ and $S$, respectively. Let $V_{it}\coloneqq V_{t}\left(X_{it},Z_{it},W_{i}\right),$$\widehat{V}_{it}\coloneqq\widehat{V}_{t}\left(X_{it},Z_{it},W_{i}\right)$, and $\widehat{v}_{it}\coloneqq\widehat{V}_{t}\left(x_{it},z_{it},w_{i}\right)$. For any $s=\left(x,v,w\right)\in\mathcal{S}$, I estimate $G_{t}\left(s\right)\coloneqq\mathbb{E}\left[\rest{Y_{it}}S_{it}=s\right]$ by regressing $Y_{it}$ on the basis functions $p^{M_{2}}$ of $\left(X_{it},\widehat{V}_{it},W_{i}\right)$:
where $\widehat{P}_{t}:=n^{-1}\sum_{i=1}^{n}\widehat{p}_{it}\widehat{p}_{it}'$, $\widehat{p}_{it}:=p^{M_{2}}\left(x_{it},\widehat{v}_{it},w_{i}\right)$, $\widehat{p}_{t}$ is an $n\times M_{2}$ matrix with $\widehat{p}_{it}'$ being its $i^{\text{th}}$ row, and $y_{t}$ is the vector of $y_{it}$s. Following \citet*{newey1999nonparametric}, I let
by exploiting the index structure of ((ref)). Hence, the effective degree that matters for the convergence rate is $m_{2}$ since $M_{2}=d_{X}\times m_{2}$ and $d_{X}$ is finite.
Finally, I exploit the index structure of ((ref)) again to estimate $b_{1t}\left(v,w\right)\coloneqq\mathbb{E}\left[\rest{\beta_{it}}V_{it}=v,W_{i}=w\right]$. When ((ref)) is used, I estimate $b_{1t}\left(v,w\right)$ by
where $\widehat{\mathbb{E}}\left[\rest{X_{it}Y_{it}}\widehat{V}_{it},W_{i}\right]$ is obtained by regressing each coordinate of $X_{it}Y_{it}$ on the basis functions of $\left(\widehat{V}_{it},W_{i}\right)$ and similarly for $\widehat{\mathbb{E}}\left[\rest{X_{it}X_{it}'}\widehat{V}_{it},W_{i}\right]$. When ((ref)) is used, I estimate $b_{1t}\left(v,w\right)$ by
where the second equality holds by the chain rule. I follow ((ref)) in what follows, as it simplifies both implementation and the derivation of asymptotic properties.
To estimate $\ol b$ and $b_{t}\left(x\right)$, by the law of iterated expectations I regress $\widehat{b}_{1t}\left(\widehat{V}_{it},W_{i}\right)$ on the basis function $r^{M_{3}}$ of constant one and $X_{it}$, respectively:
where $\widehat{R}_{t}:=n^{-1}\sum_{i=1}^{n}r_{it}r_{it}'$, $r_{it}:=r^{M_{3}}(x_{it})$, $r_{t}$ is an $n\times M_{3}$ matrix with $r_{it}'$ being its $i^{\text{th}}$ row, and $\widehat{B}_{t}$ is an $n\times d_{X}$ matrix with $\widehat{b}_{1t}(\widehat{v}_{it},w_{i})'$ being its $i^{\text{th}}$ row.
Since I let $n\rightarrow\infty$ for each fixed $t$ in the asymptotic analysis, the $t$-subscript is suppressed when it is clear. Let $p_{i}\coloneqq p^{M_{2}}\left(X_{i},V_{i},W_{i}\right)$ and $P\coloneqq\mathbb{E} p_{i}p_{i}'$. In the analysis, I use results from imbens2009identification to establish convergence rates for $\widehat{V}\left(x,z,w\right)$ and $\widehat{G}\left(s\right)$. The following assumption is needed.
Let $\zeta\left(m_{2}\right)\coloneqq m_{2v}^{\theta}m_{2}$ and $\zeta_{1}\left(m_{2}\right)\coloneqq m_{2v}^{\theta+2}m_{2}$. With Assumption (ref) in place, the next lemma follows directly from Theorem 12 of \citet*{imbens2009identification}. Let $\Delta_{1n}^{2}:=n^{-1}M_{1}+M_{1}^{1-2d_{1}/j_{1}}$ and $\Delta_{2n}^{2}:=\Delta_{1n}^{2}+n^{-1}m_{2}+m_{2}^{-2d_{2}/j_{2}}$ in the next lemma.
Lemma (ref) states that the mean-square convergence rate for $\widehat{G}\left(s\right)$ is the sum of the first-step rate $\Delta_{1n}^{2}$, the variance term $n^{-1}m_{2}$, and the squared bias term $m_{2}^{-2d_{2}/j_{2}}$. $d_{1}/j_{1}$ and $d_{2}/j_{2}$ are the uniform approximation rates that govern how well the unknown functions $V\left(x,z,w\right)$ and $G\left(s\right)$ can be approximated with $\widehat{V}\left(x,z,w\right)$ and $\widehat{G}\left(s\right)$, respectively; see Assumptions 3 and 5 of \citet*{imbens2009identification} for the details.
Let $\xi_{i}\coloneqq b_{1}\left(V_{i},W_{i}\right)-b\left(X_{i}\right)$ and $\xi\coloneqq\left(\xi_{1},\ldots,\xi_{n}\right)'$. To derive the convergence rates for the APE and LAR estimators, I impose the following assumption.
Assumption (ref) imposes conditions on the degree of smoothness of $b\left(x\right)$, the normalization of basis functions $r^{M_{3}}\left(x\right)$, and the boundedness of the second moment of $\xi_{i}$, similar to those in Assumption (ref). Since $\widehat{b}_{1}\left(v,w\right)$ and $\widehat{G}\left(s\right)$ share the same series regression coefficient $\widehat{\alpha}^{M_{2}}$, the convergence rate of $\widehat{b}_{1}\left(v,w\right)$ to $b_{1}\left(v,w\right)$ is the same as that of $\widehat{G}\left(s\right)$ to $G\left(s\right)$. I use this result to prove the convergence rates of $\widehat{\ol b}$ and $\widehat{b}\left(x\right)$, both of which are unknown but estimable functionals of $\widehat{G}\left(s\right)$. Let $\Delta_{3n}^{2}:=n^{-1}M_{3}+M_{3}^{-2d_{3}/j_{3}}+\Delta_{2n}^{2}$ in the next theorem.
The 2002 working paper version of imbens2009identification (henceforth IN02) has obtained asymptotic normality for estimators of known and scalar-valued linear functionals of $G\left(s\right)$. I build on their results to analyze vector-valued functionals of $G\left(s\right)$ via a Cram\'er--Wold device and prove asymptotic normality for $\widehat{b}_{1}\left(v,w\right)$. To simplify the notation, I take $\widetilde{\ol p}^{M_{2}}\left(s\right)$ and $\widetilde{r}^{M_{3}}\left(x\right)$ that satisfy Assumptions (ref) and (ref) as $\ol p^{M_{2}}\left(s\right)$ and $r^{M_{3}}\left(x\right)$ in what follows.
Assumption (ref)(ref)--(ref) is also imposed by IN02 and is a regularity condition required for the asymptotic normality of $\widehat{b}_{1}\left(v,w\right)$. Assumption (ref)(ref) concerns the asymptotic covariance matrix of $\widehat{b}_{1}\left(v,w\right)$ and is used by don1991asymnormality. It guarantees that the normality result of IN02 applies to vector-valued functionals of $G\left(s\right)$. Essentially, Assumption (ref)(ref) requires that all the coordinates of $\widehat{b}_{1}\left(v,w\right)$ converge at the same speed, which is mild because ex-ante I do not distinguish any coordinate of $\beta_{it}$ from the others.
Lemma (ref) concerns $b_{1}\left(v,w\right)$, a known functional of $G\left(s\right)$. Therefore, the results of IN02 directly apply and I omit its proof in this paper. However, the results of IN02 do not directly apply to $\widehat{\ol b}$ and $\widehat{b}(x)$ because $\ol b$ and $b(x)$ are unknown functionals of $G\left(s\right)$. To explain, notice that by the law of iterated expectations
both of which involve integrating $b_{1}\left(V,W\right)=\partial G\left(S\right)/\partial X$ with respect to the unknown but estimable distribution of $\left(V,W\right)$. Therefore, I need to estimate the unknown functionals in ((ref)) and correctly account for the additional estimation bias in the asymptotic analysis.
Assumption (ref)(ref) is needed to show that the asymptotic covariance matrix $\Omega_{2}$ of $\sqrt{n}(\widehat{b}\left(x\right)-b\left(x\right))$ is positive definite and similarly for $\Omega_{3}$. Assumption (ref)(ref) is a regularity condition imposed for the Lindeberg--Feller central limit theorem (CLT). Assumption (ref)(ref) is similar to Assumption (ref)(ref) and is needed to prove that the asymptotic normality result holds for vector-valued functionals of $G\left(s\right)$ by the Cram\'er--Wold device.
I present the main idea to the proof of Theorem (ref), focusing on how the error propagates through the three-step analysis and the roles of $\widehat{\Omega}_{2}$ and $\widehat{\Omega}_{3}$. Define the functionals
Then
This decomposition makes clear that the asymptotic covariance of $\sqrt{n}\,(\widehat{b}(x)-b(x))$, denoted by $\Omega_{2}$ as in ((ref)), must account for errors of all three sources. In the proof of Theorem (ref), I show that
where the influence functions $\psi_{1i},\ \psi_{2i},\text{ and }\psi_{3i}$ correspond exactly to the three error terms on the right-hand side of ((ref)) up to proper normalization. Once ((ref)) is established, I verify the Lindeberg-Feller condition on $\Psi_{in}:=n^{-1/2}\sum_{k=1}^{3}\psi_{ki}$, which proves its convergence to the standard normal distribution and that $\Omega_{2}$ is indeed the appropriate asymptotic covariance associated with $\sqrt{n}\,(\widehat{b}(x)-b(x))$.
For feasible inference, a consistent estimator for $\Omega_{2}$ is needed, and I show in the proof of Theorem (ref) that $\widehat{\Omega}_{2}$ is consistent. Note that $\Omega_{2}$ may diverge because $\widehat{b}(x)$ converges more slowly than $n^{-1/2}$ as shown in Theorem (ref). Hence, all convergence results are expressed in self-normalized way, for example,
The same argument applies to $\widehat{\Omega}_{3}$ with the basis function $r^{M_{3}}\left(x\right)=1$ and the influence function capturing serial correlation of estimating the APE across periods.
In this section, I apply my procedure to estimate a heterogeneous Cobb-Douglas production function for each of the five largest manufacturing sectors in China. I obtain unconditional means of the output elasticities (APE) and compare them with those derived using classic methods on the same data set. Furthermore, I estimate conditional means of the elasticities given the regressors (LAR). Results show that there are significant across-firm variations in the output elasticities. I also examine how heterogeneity in output elasticities relates to observed characteristics and obtain intuitive results.
In Appendix (ref), I conduct several robustness checks---including alternative price deflators, sample trimming, alternative IV and $W_{i}$ constructions, city--year fixed effects, and comparison with OLS and naive IV estimates---and the results remain stable. In Appendix (ref), I conduct a production-function motivated simulation study to support the findings in this application.
I use the China Annual Survey of Industrial Firms (CASIF), a longitudinal micro-level dataset collected by the National Bureau of Statistics of China that includes information on all state-owned industrial firms and non-state-owned firms with annual sales above 5 million RMB (\textasciitilde US\$770K). According to \citet*{brandt2017wto}, they account for 91% of the gross output, 71% of employment, 97% of exports, and 91% of total fixed assets in 2004, and thus are representative of industrial activities in China. Many papers on topics such as firm behavior, international trade, and growth theory have used the CASIF data (e.g., \citet*{hsieh2009misallocation}).
I focus on the five largest 2-digit sectors in terms of the number of firms between 2004 and 2007. Note that the CASIF dataset spans between 1998 and 2007. I choose year 2004 to 2007 to (i) ensure data consistency due to the change in the Chinese Industry Classification codes in 2003, (ii) avoid major structural breaks in the early 2000s (e.g., China joined the WTO in 2001), and (iii) use the most recent data.
Following \citet*{brandt2014challenges}, appropriate price deflators for inputs and outputs are applied separately. I preprocess the data so that firms with strictly positive amounts of capital, employment, value-added output, real wage expense, and real interest rate are used for estimation. The final dataset consists of a balanced panel of 11,317 firms over four years across five sectors. Summary statistics are presented in Table (ref).
The value-added heterogeneous production function is
I highlight two features of model (ref). First, output elasticities $\beta_{it}\coloneqq(\beta_{it,K},\beta_{it,L})'$ are allowed to differ across firms and through time. Second, input choices $X_{it}\coloneqq(k_{it},l_{it})'$ can be correlated with $\beta_{it}$ through their dependence on the pairwise correlated random variables $A_{i}$, $\eta_{it}$, and $\varepsilon_{it}$.
I use leave-one-out industry--city--year cell averages of real wages and real interest rates faced by other firms in the same industry and city, weighted by competitors\textquoteright employment (for wages) and total debt (for interest rates) as IVs. I write them as $Z_{it}\coloneqq(\ln\ol r_{cs(i),t}^{\text{real},(-i)},\,\ln\ol w_{cs(i),t}^{\text{real},(-i)})'$, where $cs(i)$ represents city and sector of firm $i$. These instruments are relevant because firms compete locally for labor and capital, so variation in competitors\textquoteright cost conditions shifts firm $i$\textquoteright s effective input prices and hence its optimal input choices. At the same time, they are plausibly exogenous to firm $i$\textquoteright s idiosyncratic productivity shocks and fixed effects because firm $i$\textquoteright s own input prices are excluded from the construction; identification therefore relies on competitors\textquoteright cost variation rather than firm-level wage or interest-rate realizations. In Appendix (ref), I examine the exogeneity issue by using a more exogenous IV constructed at the provincial level, and the results remain stable.
For estimation, I take $W_{i}$ to be the mean through time of each coordinate of $X_{it}$ by adapting the argument of mundlak1978pooling nonparametrically. I also examine the effect of further including the mean through time of each coordinate of $Z_{it}$ and the results reported in Table (ref) are similar. For all series estimation steps, I use second-degree polynomial spline basis functions with knots at the median, and a robustness check to this choice is included in Appendix (ref). First, I estimate each coordinate of $V_{it}\coloneqq(F_{\rest{k_{it}}Z_{it},W_{i}}(\rest{k_{it}}Z_{it},W_{i}),F_{\rest{l_{it}}Z_{it},W_{i}}(\rest{l_{it}}Z_{it},W_{i}))'$ by regressing $\mathbf{\mathbbm1}(k_{it}\leq k)$ and $\mathbf{\mathbbm1}(l_{it}\leq l)$ on the basis functions of $(Z_{it},W_{i})$, respectively. Next, I estimate $G_{it}\coloneqq\mathbb{E}[\rest{y_{it}}X_{it},V_{it},W_{i}]$ by regressing $y_{it}$ on the basis functions of $(X_{it},\widehat{V}_{it},W_{i})$. Estimation of $b_{1t}(V_{it},W_{i})\coloneqq\mathbb{E}[\rest{\beta_{it}}V_{it},W_{i}]$ is then obtained by taking the partial derivative of $\widehat{G}_{it}(X_{it},\widehat{V}_{it},W_{i})$ with respect to $X_{it}$. Finally, I estimate the pooled APE $\ol b$ by simply averaging $\widehat{b}_{1t}(\widehat{V}_{it},W_{i})$ over $i$ and $t$. For the LAR, as in ((ref)), I regress $\widehat{b}_{1t}(\widehat{V}_{it},W_{i})$ on the basis functions of $X_{it}$ for each $t$, then average over $t$ to obtain $\widehat{\beta}_{i,K}$ and $\widehat{\beta}_{i,L}$ as a measure of heterogeneity in output elasticities across firms.
I compare my TERC estimates of average output elasticities with those obtained from OP (olleypakes1996teltfp), LP (levinsohnpetrin2003res), and ACF (\citet*{ackerberg2015identification}) applied to the same dataset. Table (ref) contains the estimation results.
The $\widehat{\ol b}_{K}$ estimates from TERC fall within $\left[0.371,\,0.463\right]$ across the five sectors, broadly consistent with the results obtained from applying OP, LP, and ACF to the same dataset. In contrast, the $\widehat{\ol b}_{L}$ estimates exhibit greater discrepancies across methods. The TERC estimates of $\widehat{\ol b}_{L}$ lie in $\left[0.311,0.588\right]$ across the five sectors. OP's labor elasticity estimates are close to TERC's for the first three sectors but are smaller for the last two. LP seems to produce smaller $\widehat{\ol b}_{L}$ estimates across all sectors, whereas ACF generates larger $\widehat{\ol b}_{L}$ estimates for the general equipment and transportation equipment sectors. It is worth noting that the ACF approach tends to produce wider confidence intervals in certain sectors, which may partly reflect numerical instability in the Stata implementation, an issue also discussed by \citet*{aureo2024production}. The 95% CIs for TERC estimates are reasonably tight across all sectors.
I consider the estimation results from TERC reasonable based on two pieces of empirical evidence. First, it is well documented in the literature that output elasticities in Cobb-Douglas production function estimation usually lie within $\left[0,1\right]$. Second, using Chinese manufacturing data, hsieh2009misallocation show that roughly half of output accrues to capital. Under profit maximization with a Cobb-Douglas production function, output elasticities correspond to factor revenue shares. The magnitudes of the estimated capital and labor elasticities can therefore be interpreted through this lens, and my estimates are consistent with this result.
The broad similarity of average elasticities across methods is not unexpected, as OP, LP, ACF, and the TERC approach are all designed to address simultaneity between input choices and unobserved productivity. While OP, LP, and ACF provide consistent estimates under a constant-coefficient production function, TERC adds value by relaxing this restriction and allowing output elasticities to vary across firms while still addressing simultaneity via a control-function approach.
Next, I examine the distribution of conditional mean output elasticities given the regressors. These conditional means summarize average elasticities for firm subgroups defined by their capital and labor levels, and thus provide a useful basis for designing more targeted, sector-specific policies. They are also informative about the underlying distribution of true elasticities. For instance, by the law of total variance, the variance of these conditional means is a lower bound on the variance of the true output elasticities. demirer2020production argues that the heterogeneity in output elasticities is largely explained by across-firm variation. To study this issue, I average $\widehat{b}_{t,K}(X_{it})$ and $\widehat{b}_{t,L}(X_{it})$ for each firm through time and denote them by $\widehat{\beta}_{i,K}$ and $\widehat{\beta}_{i,L}$, respectively. These $\widehat{\beta}_{i}$ estimates can be considered as a proxy for their output elasticities. Then, I plot the histograms of $\widehat{\beta}_{i,K}$ and $\widehat{\beta}_{i,L}$ across firms for each sector.
In Figure (ref), I present the histograms of $\widehat{\beta}_{i,K}$ and $\widehat{\beta}_{i,L}$ for the textile sector and indicate the corresponding pooled APE labeled by “Mean” in the figure. The left subplot of Figure (ref) corresponds to the histogram of $\widehat{\beta}_{i,K}$. All of its probability mass lies between zero and one, with its mode at 0.4. Over 95% of its probability mass lies between 0.3 and 0.5. The distribution of $\widehat{\beta}_{i,K}$ is concentrated around the mean and symmetric, suggesting that these firms tend to be homogeneous in capital efficiency. The right subplot of Figure (ref) is the histogram of $\widehat{\beta}_{i,L}$. Again, the majority of its probability mass lies between zero and one. The distribution of $\widehat{\beta}_{i,L}$ is more dispersed than that of $\widehat{\beta}_{i,K}$ and is slightly left-skewed, with a small number of textile firms exhibiting low labor efficiency.
Figure (ref) presents the histograms of $\widehat{\beta}_{i,K}$ and $\widehat{\beta}_{i,L}$ for the other sectors. I draw the following two conclusions. First, for all sectors, the majority of $\widehat{\beta}_{i,K}$ and $\widehat{\beta}_{i,L}$ lie between zero and one, consistent with the empirical evidence in hsieh2009misallocation. Second, there is substantial across-firm variations in both capital and labor elasticities within each sector. The extent of heterogeneity, however, differs across elasticities and sectors. For example, in the left panel of Figure (ref), capital elasticity exhibits greater dispersion among chemical firms than among firms in the nonmetallic minerals sector.
I examine cross-firm heterogeneity in output elasticities by regressing the estimated $\widehat{\beta}_{i,K}$ and $\widehat{\beta}_{i,L}$ on firm characteristics. Specifically, using the CASIF dataset, I construct five covariates: firm size, leverage, export status, industrial-park location, and state ownership. All specifications include industry and city fixed effects, and I report heteroskedasticity-robust standard errors.
The results summarized in Table (ref) show that larger firms have significantly higher capital elasticities, while leverage is positively associated with labor elasticities. Exporters exhibit modestly higher capital elasticities, and state-owned enterprises have substantially lower labor elasticities. Conditional on these controls, the industrial-park indicator is small. Industry and location fixed effects are jointly significant, indicating that sectoral and regional factors account for an important share of the heterogeneity. The adjusted-$R^{2}$s are 0.58 and 0.57 for $\widehat{\beta}^{K}$ and $\widehat{\beta}^{L}$, respectively, suggesting that over half of the variation in estimated elasticities can be explained by observable firm characteristics and fixed effects.
The results are generally intuitive and broadly consistent with existing empirical evidence. Larger firms, for instance, often operate in more capital-intensive segments or adopt technologies with stronger scale economies, automation, and standardized production processes, so output tends to be more responsive to capital at the margin (Syverson2011JEL). Leverage may also be associated with more labor-intensive operating choices---for example, delaying or outsourcing capital investment, relying on older or rented equipment, or placing greater emphasis on variable inputs (\citet*{KalemliOzcanLaevenMoreno2022JEEA}). To the extent that highly leveraged firms substitute away from capital toward labor, the estimated production relationship will place greater weight on labor. Finally, exporters tend to use capital more efficiently (BernardJensen1999JIE), whereas state-owned enterprises may employ labor in excess of the efficient level due to political considerations (Wen2025RESTAutocraticControl), which is consistent with their substantially lower labor elasticities relative to private firms.
This paper proposes a new TERC model in which regressors are correlated with random coefficients through not only a fixed effect but also a time-varying shock---an empirically relevant feature consistent with optimizing behavior in many applications. I construct feasible control variables for both the fixed effect and the time-varying shock, and use the resulting residual variation in regressors to identify the APE and LAR. I then develop three-step series estimators and establish their convergence rates and asymptotic normality. In an application to Chinese manufacturing, the estimates reveal substantial cross-firm dispersion in output elasticities, part of which is explained by observable firm characteristics.
I propose two directions for future research. First, beyond the low-order moments studied here, policymakers may be interested in the full distribution of random coefficients. One---admittedly demanding---route is to identify moments of all orders by induction and recover the distribution via moment determinacy (stoyanov2000krein). Second, it remains open whether elements of the approach can be extended to dynamic linear or nonlinear panel models, such as those in \citet*{marx2024heterogeneous} and \citet*{liu2024identification}. Doing so would likely require additional structure governing how lagged outcomes (or state variables) co-move with the random coefficients.