Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
112,653 characters · 24 sections · 14 citation commands
Using Monotonicity Restrictions to Identify Models with Partially Latent Covariates
\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\def\mc#1{\mathscr{#1}} \global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\def\abs#1{\left|#1\right|} \global\long\def\norm#1{\left\Vert #1\right\Vert } \global\long\def\rest#1{\left.#1\right|} \global\long\def\bracket#1#2{\left\langle #1\middle\vert#2\right\rangle } \global\long\def\sandvich#1#2#3{\left\langle #1\middle\vert#2\middle\vert#3\right\rangle } \global\long\def\turd#1{\frac{#1}{3}} \global\long\global\long\def\sand#1{\left\lceil #1\right\vert } \global\long\def\wich#1{\left\vert #1\right\rfloor } \global\long\def\sandwich#1#2#3{\left\lceil #1\middle\vert#2\middle\vert#3\right\rfloor } \global\long\def\abs#1{\left|#1\right|} \global\long\def\norm#1{\left\Vert #1\right\Vert } \global\long\def\rest#1{\left.#1\right|} \global\long\def\inprod#1{\left\langle #1\right\rangle } \global\long\def\ol#1{\overline{#1}} \global\long\def\ul#1{#1} \global\long\def\td#1{\tilde{#1}} \global\long\global\long\global\long\global\long\global\long\global\long
This paper develops a new method for identifying econometric models with partially latent covariates. We show that a broad class of econometric models that play a large role in industrial organization and labor economics can be nonparametrically identified if the partially latent covariate variables satisfy certain monotonicity assumptions. Examples that fall into this class of models are a variety of different production, skill formation, and achievement functions.\footnote{Other potential applications in applied microeconomics are discussed in the conclusions.}
It is often plausible to assume that the different inputs or covariates are functions of a common unobserved random shock, and we consider models in which it is natural to impose strict monotonicity in this common shock.\footnote{Note that this assumption is commonly used, for example, in the production function literature as discussed by \citet*{olley1996dynamics}. In particular, this assumption does not require that inputs are “optimally” chosen by competitive firms and is consistent with a broad class of strategic and non-strategic models that may describe the agents' behavior.
} The monotonicity assumption imposes functional dependencies on the explanatory variables as pointed out in the context of production function estimation by \citet*{ackerberg2015identification}. The key insight of this paper is that we can leverage the functional dependence between inputs to achieve identification within a partially latent covariate framework. In that sense, we turn the functional dependence problem on its head to impute the partially latent covariates. Broadly speaking, our imputation is in the spirit of matching algorithms Rubin-73. In contrast to traditional matching algorithms, we propose to match on the expected dependent variable to impute missing covariates.
The partially latent data structure that we study in this paper arises quite naturally in certain matched employer-employee data sets which contain information collected from individuals as well as information collected from businesses or establishments. In the past decades economists made enormous strides in finding and using matched employer-employee data, which have provided a new empirical basis for the study of workplace organization, compensation design, mobility, and production. In this paper we focus on cross-sectional employer-employee data sets which are commonly used in applied microeconomics and econometrics. Following \citet*{Abowd-Kramaz-99}, an important dimension that distinguishes matched employer-employee data sets is the sampling design. Some sampling designs focus on the firm, while others use the employee as the primary unit of analysis. In this paper we focus on the later type of sampling designs that often only sample one employee or a small random set of employees of the firm.
This lack of sampling of all employees of the firm may not be important from a perspective of labor economics, which largely focuses on job mobility and wage determination. However, it is more problematic if the focus is on productivity measurement within a production function framework. If the survey does not sample all employees in a firm then some important inputs are latent from the perspective of the econometrician. As a consequence, we call this type of sampling design an “input-based sampling”\ approach, since the sampling unit is one of the multiple labor input factors. We provide a formal definition of this data structure in this paper. The main application that we study considers a production team in which team members perform different tasks. In our data set only one member from each team is interviewed to provide the data. It is plausible that this person knows the team's output, but does not have complete information about the other team members' input choices. By randomly sampling the teams we elicit information from all different types of team members and hence input factors.
Once we have identified the latent inputs, the estimation of the outcome or production function can proceed using standard semiparametric methods developed in the econometric literature. One key issue here is that the common shock creates an endogeneity problem.\footnote{In the context of production function estimation this endogeneity problem is referred to as the transmission bias problem since inputs are correlated with unobserved productivity shocks marschak1944random.} We show that we can combine our identification results with a variety of linear, nonlinear, and semiparametric estimation strategies. In that sense our approach is flexible and allows researchers to make appropriate functional form assumptions if necessary. We consider the scenario in which researchers only have access to a single cross-section of data and rely on instrumental variables for estimation.\footnote{Hence we cannot address this endogeneity problem using panel data with fixed effects, first advocated by \citet*{hoch1955estimation,hoch1962estimation} and \citet*{mundlak1961empirical,mundlak1963specification}. We can also not use more sophisticated timing assumptions within a control function or IV frameworks as discussed, for example, in \citet*{olley1996dynamics} and \citet*{blundell1998initial,blundell2000gmm}, \citet*{levinsohn2003estimating}, and \citet*{ackerberg2015identification}. We discuss the extension of our methods to this scenario in the conclusions.} For example, production function estimation relies on the assumption that differences in local input prices give rise to differences in input choices that are uncorrelated with productivity shocks at the local level.\footnote{Hence, local input prices can serve as valid instruments for endogenous input choices. See griliches1998 for a critical discussion of the assumption that these input prices are exogenous. Similarly, skill formation and achievement function estimation requires the choice of suitable instruments for parental inputs. For a more general discussion of the issues encountered in estimating achievement and skill formation functions see, among others, \citet*{Todd-Wolpin-03} and \citet*{Cunha-Heckman-Schennach-10}.}
Estimation proceeds in two steps. In finite samples, we first nonparametrically estimate the latent input functions. Plugging the estimators into our outcome equation, we can estimate the parameters of this function using a standard IV estimator based on the observed and imputed inputs. The second econometric challenge then arises for the need to account for the sequential nature of the estimator when deriving the correct rate of convergence and computing asymptotic standard errors. To illustrate this we consider the standard log-linear, Cobb-Douglas model. We propose two different estimators and provide both high-level and lower-level conditions under which these semiparametric two-step estimators are consistent and asymptotically normal at the usual parametric rate of convergence. The technical proofs are based on the general econometric theory on semiparametric two-step estimation as in newey1994asymptotic, \citet*{newey1994large}, and \citet*{chen2003estimation}. Finally, we show that using the conditional expectation of outcomes as the dependent variable produces efficiency gains relative to the more traditional estimator that uses the observed output instead.
To evaluate the performance of our estimator we conduct a variety of Monte Carlo experiments. Our findings suggest that our estimators are well-behaved in samples that are similar in size to those observed in our applications discussed below. We also study the behavior of our estimator when we pool observations across markets as is often necessary for many practical applications.
The type of sampling design that gives rise to a partial latent data structure arises in administrative employer-employee data sets that are designed by statistical agencies. It is also common among profession-based surveys that focus on one narrowly-defined occupation. These surveys tend to be national in scope since professions have generally a national market, the characteristics of which the survey organizers want to know. In our main application, we use data from the National Pharmacist Workforce Survey (NPWS) in 2000 which not only collects data for each pharmacist that is surveyed but also a limited amount of information at the store level including output.
We implement our new estimator using the NPWS and study differences in the marginal product of labor inputs in pharmacies. \citet*{goldin2016most} have forcefully argued that this is one of the most egalitarian and family-friendly professions in which females face little discrimination in the workforce. One potential explanation of this fact has been related to the rise of chains that have replaced independent pharmacies in many local markets. Here we estimate a team production function that distinguishes between managerial and non-managerial certified pharmacists. We can, therefore, test the hypothesis of whether managers have higher marginal products in chains than in independent pharmacies. We find that we can reject the null hypothesis that independent pharmacies and chains have the same technology. Estimates for independent pharmacies are somewhat noisy but do not suggest that there is a large difference between managers and regular employees. Estimates for chains suggest that managers have higher marginal products than regular employees. We thus conclude that chains seem to improve the effectiveness of managers which may partially explain why they have become the dominant firm type in this industry.
This paper relates to the line of literature on production function estimation by proposing a method to handle the problem of partially latent inputs. Our identification strategy is based on strict monotonicity and the consequent invertibility in a scalar unobservable, a feature also leveraged by \citet*{olley1996dynamics} and \citet*{levinsohn2003estimating}. They essentially use an auxiliary variable together with an input to control for the unobserved productivity shock: investment with capital in \citet*{olley1996dynamics} and intermediate inputs with capital in \citet*{levinsohn2003estimating}. In comparison, we use the output with the observed input to pin down the productivity shock. We emphasize that the feature of functional dependence between input variables, which was pointed out by \citet*{ackerberg2015identification} as an underlying problem in \citet*{olley1996dynamics} and \citet*{levinsohn2003estimating}, in fact, forms the basis of our imputation strategy. While most of these papers focus on value-added production functions, there is also much interest in estimating gross output production functions. \citet*{Doraszelski-Jaumandreu-13} propose a solution to the transmission bias problem that also relies on observed firm-level variation in prices. In particular, they show that by explicitly imposing the parameter restrictions between the production function and the demand for a flexible input and by using this price variation, they can recover the gross output production function. \citet*{Gandhi-Navarro-Rivers-20} provide an alternative identification strategy to estimate gross output production functions that works well in short panels. Beyond these conceptual linkages, our paper has a different focus from these papers cited above: they focus more on the dynamic nature of capital inputs, while we focus on the problem of partially latent inputs. Moreover, the estimation of production functions is just one of many applications of our general identification result. This paper shows that our methods may be even more useful for applications outside of IO where these data structures are more prevalent as we discuss below.
Also, we should point out that this paper is both conceptually and technically different from previous work on missing data in linear regression and, more generally, GMM estimation settings, such as \citet*{rubin1976inference}, \citet*{little1992regression}, \citet*{robins1994estimation}, \citet*{wooldridge2007inverse}, \citet*{graham2011efficiency}, \citet*{chaudhuri2016gmm}, \citet*{abrevaya2017gmm} and \citet*{mcdonough2017missing}. This line of literature usually exploits two types of conditions: first, observations with no missing data occur with positive probability, and second, data are “missing at random” (potentially with conditioning). Neither condition is satisfied in our setting: every observation contains missing data, and missing can be correlated with other observables as well as the unobserved productivity shock. Instead, we rely on monotonicity in a scalar unobservable shock to identify and impute the latent input.
Similarly, our monotonicity conditions also differentiate our paper from the econometric literature on data combination as surveyed by \citet*{ridder2007econometrics}, which mostly involves conditional independence assumptions. In particular, our paper shares a similar flavor with the “moment-matching” approach in this literature, such as \citet*{angrist1992effect}, \citet*{ridder2007econometrics}, and \citet*{buchinsky2022estimation}: these papers utilize conditional independence to achieve “moment matching”, while we exploit monotonicity to do so. Hence, our proposed method may also be useful as a complementary data combination method for scenarios where our monotonicity conditions are interpretable and justifiable.
Broadly speaking, our imputation is in the spirit of matching algorithms Rubin-73. In contrast to traditional matching algorithms, we propose to match on the expected dependent variable to impute missing covariates. Hence, we do not apply the matching approach within the standard potential outcome framework of program evaluation based on the potential outcome model developed by \citet*{Fisher-35}.\footnote{For a discussion of the properties of matching estimators in that context see, among others, \citet*{Rosenbaum-Rubin-83}, \citet*{Heckman-Ichimura-Smith-Todd-98}, and \citet*{Abadie-Imbens-06}.}
The rest of the paper is organized as follows. Section (ref) presents our main identification result. Section (ref) discusses the problems associated with estimation. Section (ref) reports the results from a Monte Carlo Study. Section (ref) introduces our application focusing on the production functions of pharmacies. It discusses our data sources and presents our main empirical findings. Section (ref) provides a discussion of other potential applications and presents our conclusions.
Consider the following cross-sectional econometric model:
where $i=1,...,N$ indexes a generic observation from a random sample, $y_{i}$ denotes an observable scalar-valued output variable, and $x_{i}:=\left(x_{i1},x_{i2}\right)$ denotes a two-dimensional vector of covariates.\footnote{See Corollary (ref) for the extension of our identification method to settings with covariates of higher dimensions.} Both $u_{i}$ and $\epsilon_{i}$ are scalar-valued unobserved errors, with $u_{i}$ taken to be a structural error (such as a productivity shock) that is endogenous with respect to $x_{i}$, while $\epsilon_{i}$ is a “measurement error” that is assumed to be exogenous. The unknown outcome function $F$ may be either parametric or nonparametric.
First, we need to define what we mean by partially latent covariates, a key data structure that we seek to handle in this paper.
Essentially, one of the two covariates $\left(x_{i1},x_{i2}\right)$ is latent in each observation in the data. In the following, it will be convenient to write \[ d_{i}:=
\] so that effectively $\left(d_{i},\left(2-d_{i}\right)x_{i1},\left(d_{i}-1\right)x_{i2}\right)$ is observed for $i$.
The next assumption imposes a monotonicity condition on the outcome function.
This assumption essentially states that the inputs $\left(x_{i1},x_{i2}\right)$ and the productivity shock $u_{i}$ have nonnegative effects on the output variable $y_{i}$. Moreover, the monotonicity is strict in, at least, one of the three arguments $x_{i1},x_{i2}$, and $u_{i}$. The restriction of monotonicity with respect to $\left(x_{i1},x_{i2}\right)$ is substantive: it requires that the inputs cannot negatively affect the output variable holding everything else fixed. In contrast, the restriction of monotonicity with respect to $u_{i}$ is largely innocuous given the interpretation of $u_{i}$ as a (weakly) “positive shock”.
Next, we turn to the assumptions on the unobserved errors $u_{i}$ and $\epsilon_{i}$ in equation (ref). First, we assume that the endogenous inputs $x_{i}$ are strictly monotone functions of the scalar productivity shock $u_{i}$, potentially after conditioning on a set of observed covariates $z_{i}$, that may affect the inputs $x_{i}$.
We note that the functions $h_{1}$ and $h_{2}$ can be unknown and nonparametric. The only requirement here is that, after conditioning on $z_{i}$, the covariates $x_{i1}$ and $x_{i2}$ can be written as deterministic monotone functions of the error $u_{i}$. Such a “monotonicity-in-a-scalar-error” assumption has been widely used in the econometric literature on identification analysis.\footnote{See matzkin2007nonparametric for a general survey, and see \citet*{ackerberg2015identification} in the specific context of production function identification, which fits into our working example (ref).}
Assumption (ref) can be further micro-founded in a variety of settings based on efficiency or equilibrium criteria, which we elaborate in Subsection (ref), given that Assumption (ref) is the key assumption in this paper.
Note that, under Assumption (ref), conditioning on $\left(x_{i},z_{i},d_{i}\right)$ is equivalent to conditioning on $\left(u_{i},z_{i},d_{i}\right)$. In the production function estimation literature without the partial latency problem, $\mathbb{E}\left[\left.\epsilon_{i}\right|u_{i},z_{i}\right]=0$ is a standard assumption imposed on $\epsilon_{i}$. In our current setting, we are requiring that $\epsilon_{i}$ is furthermore exogenous with respect to the partial latency indicator variable $d_{i}$.
It is worth noting that this paper is both conceptually and technically different from previous work on missing data in linear regression and, more generally, GMM estimation settings, such as \citet*{rubin1976inference}, \citet*{little1992regression}, \citet*{robins1994estimation}, \citet*{wooldridge2007inverse}, \citet*{graham2011efficiency}, \citet*{chaudhuri2016gmm}, \citet*{abrevaya2017gmm} and \citet*{mcdonough2017missing}. This line of literature usually exploits two types of assumptions to handle missing values: first, observations with no missing data occur with positive probability, and second, data are “missing at random (MAR)”: the indicator for missingness is exogenous to or independent of certain observable covariates or constructed conditioning variables. Neither condition is satisfied in our setting: here every observation contains “missing values”, and the partial latency indicator $d_{i}$ is allowed to be correlated with other observables as well as the unobserved productivity shock. Instead, we will be relying on monotonicity conditions to identify and impute the latent input.
Specifically, Assumption (ref) here is simply requiring that $\epsilon_{i}$ is a “measurement error” term that is exogenous with respect to the observables and consequently the productivity shock $u_{i}$, but does not impose any restriction on the dependence structure between the partial latency indicator $d_{i}$ and other structural components of the model $\left(u_{i},x_{i},z_{i}\right)$.
That said, we do require the following very mild condition on the variable $d_{i}$.
Assumption (ref) guarantees that conditioning on realizations of $\left(u_{i},z_{i}\right)$ we will observe $x_{i1}$, and $x_{i2}$, with strict positive probabilities. Again, this assumption is much weaker than “missing-at-random” assumptions, which would usually require that $\mathbb{P}\left\{ \rest{d_{i}=1}u_{i},z_{i}\right\} $ is constant in $u_{i}$, $z_{i}$, or some other variables. In contrast, here we do not impose any restrictions on the dependence of $\mathbb{P}\left\{ \rest{d_{i}=1}u_{i},z_{i}\right\} $ on $\left(u_{i},z_{i}\right)$ beyond non-degeneracy.
As discussed in the introduction, allowing for $d_{i}$ to depend on $\left(u_{i},z_{i}\right)$, and thus $x_{i}$, is also a feature that differentiates our monotonicity-based method from the “moment-matching method” in the literature on data combination, such as angrist1992effect,ridder2007econometrics,buchinsky2022estimation, which are based on conditional independence of $d_{i}$.
We are now ready to present our main identification result.
Given that the identification strategy underlying the Theorem (ref) is the key novelty of this paper, we prove Theorem (ref) in the main text below.
It should be pointed out that (ref) is an explicit representation of the “functional dependence" between the two input variables as in ackerberg2015identification: $x_{i1}$ is a deterministic function of $x_{i2}$, and vice versa, conditional on instruments $z_{i}$. While functional dependence was a concern in the context of olley1996dynamics, levinsohn2003estimating and ackerberg2015identification, here we are exactly leveraging the functional dependence between input variables to solve the partially latency problem.
The monotonicity of input choices in the unobserved productivity shock (Assumption (ref)) can be further micro-founded in a variety of settings based on efficiency or equilibrium criteria.
On a general level, one may use the theory of monotone comparative statics to obtain more primitive conditions for input monotonicity, which typically involve various forms of supermodularity (or increasing-difference) conditions: see \citet*{topkis1998supermodularity} and \citet*{vives2000oligopoly} for general treatments on this topic. Essentially, in settings where input choices are made by a single decision maker, we need the objective function to be supermodular in input variables $x$ and the productivity shock $u$. In settings where the input choices are generated as equilibria of a strategic game between two decision makers, strategic complementarity is typically required (so that the game is supermodular) to establish monotonicity. For games with strategic substitutes, we further need a condition to ensure that the extent of strategic substitutes is not overwhelming: see, for example, \citet*{roy2010monotone}.
We now provide some concrete examples that illustrate how Assumption (ref) can be satisfied: as the optimal choice of a single decision maker, as the Nash equilibrium of a game of strategic complements, as the Nash equilibrium of a game of strategic substitutes.
Suppose that firms optimally choose inputs to maximize profits under the Cobb-Douglas production function with a constant output price. Formally, each firm $i$ solves the problem \[ \max_{X_{i1},X_{i2}}e^{\alpha_{0}+u_{i}}X_{i1}^{\alpha_{1}}X_{i2}^{\alpha_{2}}-Z_{i1}X_{i1}-Z_{i2}X_{i2}, \] where $X_{i1},X_{i2}$ are inputs in its original scale (with $x_{i1},x_{i2}$ denoting the logarithm of $X_{i1},X_{i2}$) and $Z_{i1},Z_{i2}$ are input prices in its original scale (with $z_{i1},z_{i2}$ denoting the logarithm of $Z_{i1},Z_{i2}$). Then the input choice functions $h_{1}$ and $h_{2}$ are characterized by the relevant first-order conditions and have simple closed-form solutions that are linear, increasing in $u_{i}$, and decreasing in $z_{i}$.\footnote{We note that the problem of partially latent inputs is less relevant in that case since the “reduced-form” regression of the observed inputs on the exogenous wages $w_{i}$ will indirectly recover the production function parameters $\alpha$. This corresponds to the “duality approach” to production function estimation as discussed in detail in \citet*{griliches1998}. However, an attractive feature of our approach is also that we can test whether inputs are optimally chosen. If we reject the null hypothesis that inputs are optimal, our estimator is still feasible while duality estimators are not.} In particular, we have: \[ h_{1}\left(u,z\right)=\frac{\alpha_{0}+\left(1-\alpha_{2}\right)\log\alpha_{1}+\alpha_{2}\log\alpha_{2}-\left(1-\alpha_{2}\right)z_{1}-\left(1-\alpha_{2}\right)z_{2}+u}{1-\alpha_{1}-\alpha_{2}}, \] satisfying Assumption (ref). See Appendix (ref) for more details of the derivation.
Assumption (ref) is also satisfied if the inputs are Nash equilibrium choices of two partners, each of whom solves the following optimization problem: given $u,z$ and the other partner's choice $X_{2}$, partner 1 solves \[ \max_{X_{1}}\pi_{1}\left(X,u;Z\right):=\lambda_{1}\left(F\left(X,u\right)-Z_{1}X_{1}-Z_{2}X_{2}\right)+Z_{1}X_{1}-\frac{1}{2}c_{1}X_{1}^{2}, \] where $F\left(X,u\right):=e^{u+\alpha_{0}}X_{1}^{\alpha_{1}}X_{2}^{\alpha_{2}}$ is the Cobb-Douglas production function. The term $F\left(X,u\right)-Z_{1}X_{1}-Z_{2}X_{2}$ is the profit of the firm (as in the single decision maker's problem described above), and $\lambda_{1}\in\left(0,1\right)$ is a positive share of firm profit distributed to partner 1 as “dividends”. Moreover, in addition to the “dividends”, partner 1 receives her wage income $Z_{1}X_{1}$. Finally, $\frac{1}{2}c_{1}X_{1}^{2}$ captures partner 1's quadratic private cost of input $X_{1}$ with $c_{1}>0$.
Similar, partner 2 solves \[ \max_{X_{2}}\pi_{2}\left(X,u;Z\right):=\lambda_{2}\left(F\left(X,u\right)-Z_{1}X_{1}-Z_{2}X_{2}\right)+Z_{2}X_{2}-\frac{1}{2}c_{2}X_{2}^{2}, \] with $\lambda_{2}\in\left(0,1\right)$ and $c_{2}>0$.
This is a (supermodular) game of strategic complementarity since
Furthermore, the payoff functions feature increasing differences between the productivity shock $u$ and the input $x_{j}$:
With these conditions and the theory on monotone comparative statics, say, \citet*{milgrom1990rationalizability}, we can show that the unique Nash equilibrium $X^{*}\left(u,Z\right)$ of this game is strictly increasing in $u$, thus satisfying Assumption (ref). See Appendix (ref) for the detailed proof.
Consider an example with a married couple and two parental households, $j=1,2$, whose wealth levels are respectively $Z_{1}$ and $Z_{2}$, which is based on \citet*{Bergstrom-Blume-Varian-86}. Parents are altruistic toward their married offspring but not toward that offspring's spouse. Parental household $j$ has utility \[ v_{j}\left(X,Z\right)=\log\left(Z_{j}-X_{j}\right)+u\log\left(X_{1}+X_{2}\right) \] where $X_{j}$ is the married couple's gift from parental household $j$ and $u$ is the probability that both parental households think the children's marriage will endure. This leads to a noncooperative game between the two parental households since the incentive for either household to gift the offspring couple diminishes as the other parental household gives more. Formally, parental household $j$'s marginal return from $X_{j}$ \[ \frac{\partial}{\partial X_{j}}v_{j}\left(X,Z\right)=u\cdot\frac{1}{X_{j}+X_{k}}-\frac{1}{Z_{j}-X_{j}} \] is decreasing in the other parental household's return $X_{k}$, and thus the best response of household $j$ \[ BR_{j}\left(X_{k}\right)=\frac{u}{1+u}Z_{j}-\frac{1}{1+u}X_{k} \] is also decreasing in $X_{k}$. Hence, this is a game of strategic substitutes.
There is a unique Nash equilibrium of this game between the two parental households, given by \[ X_{1}^{\ast}=\frac{(1+u)\;Z_{1}-Z_{2}}{2+u},\quad X_{2}^{\ast}=\frac{(1+u)\;Z_{2}-Z_{1}}{2+u} \] for any $u$ and wealth levels $Z_{1},Z_{2}$, provided that the two households that are not “too” different in wealth so that interior solutions in $X_{1}^{*},X_{2}^{*}$ obtain.\footnote{Formally, for the interiority we also need to require that $u$ is strictly positive and bounded away from 0 in this stylized example. One can perturb the example in various ways to ensure interiority without such a restriction, but at the cost of additional complications.} Both $X_{1}^{*}$ and $X_{2}^{*}$ are strictly increasing in the shock $u$, and hence the outcome is strictly increasing in $u$.
Once the partially latent covariates $x_{i1},x_{i2}$ are identified and imputed, researchers may use them to identify and estimate functions or parameters defined based on $\left(x_{i1},x_{i2}\right)$. Note that researchers may use $\left(x_{i1},x_{i2}\right)$ for completely different purpose from the identification and estimation of the outcome function $F$. Hence, our proposed method in Section (ref) can be thought as a monotonicity-based method for data imputation or data combination.
That said, how to identify and estimate the outcome function $F$ is a natural question to ask, given that our method for the identification of partially latent inputs is built upon assumptions on $F$. Hence we focus on the identification and estimation of $F$ in this section.
With the latent inputs identified in Theorem (ref), we are back to equation (ref) \[ y_{i}=F\left(x_{i1},x_{i2},u_{i}\right)+\epsilon_{i}, \] but now we can effectively regard both $x_{i1}$ and $x_{i2}$ as being known, at least for identification purposes. Researchers may proceed to identify the production function $F$ under appropriate application-specific assumptions as in a “standard” setting without the partial latency problem. Hence, the identification of $F$ or other objects of interest is largely “separable” from the partial latency problem, which is the key problem we are solving in this paper.
That said, we note that the estimation of the latent inputs will affect the estimation of (the parameters of) $F$ based on “plugged-in” latent input estimates. This section provides a discussion on how to identify and estimate $F$, and analyzes the impact of the “first-stage” estimation of latent inputs on the final estimator of $F$.
While we cannot cover all relevant specifications of $F$, in this section we will provide both identification and estimation results for the linear case, which is arguably the workhorse model, or at least a natural benchmark, in various empirical applications. We also discuss how our method can be applied under more general settings.
In this subsection we focus on the linear parametric specification of $F$ as in (ref): \[ y_{i}=\alpha_{0}+\alpha_{1}x_{i1}+\alpha_{2}x_{i2}+u_{i}+\epsilon_{i}, \] where our goal is to identify and estimate the unknown parameters $\alpha:=\left(\alpha_{0},\alpha_{1},\alpha_{2}\right)$.
In the presence of the endogeneity problem between $x_{i}:=\left(x_{i1},x_{i2}\right)$ and $u_{i}$, we will need instrumental variables for the identification of $\alpha$. For illustrational simplicity, we impose the following standard IV assumption.
We now turn to the more interesting problem of estimation, propose semiparametric estimators for $\alpha$, and characterize their asymptotic distributions.
We first describe our proposed estimator. Since the identification of latent inputs via equation (ref) is constructive, it suggests a natural estimation procedure:
Step 1 (Nonparametric Regression): obtain an estimator $\hat{\gamma}_{1}$ of $\gamma_{1}$ by nonparametrically regressing $y_{i}$ on $x_{i1}$ and $z_{i}$, among firms with $d_{i}=1$, i.e., those with $x_{i1}$ observed. Similarly, obtain an estimator $\hat{\gamma}_{2}$ of $\gamma_{2}$.
Step 2 (Imputation): impute latent inputs by plugging the nonparametric estimators $\hat{\gamma}_{1},\hat{\gamma}_{2}$ into equation (ref), i.e.,
Step 3 (IV Regression): run either of the following two IV regressions:
We now establish the consistency and the asymptotic normality of $\hat{\alpha}$ and $\hat{\alpha}^{*}$ under the following regularity assumptions.
In view of equation (ref), Assumption (ref) is satisfied if either $\alpha_{1},\alpha_{2}>0$ or $\frac{\partial}{\partial u}h_{1},\frac{\partial}{\partial u}h_{1}$ are uniformly bounded above by a finite constant. Assumption (ref) is needed to ensure that $\hat{\gamma}_{k}^{-1}\left(\cdot,z\right)$ is a good estimator of $\gamma_{k}^{-1}\left(\cdot,z\right)$ provided that the first-stage nonparametric estimator $\hat{\gamma}_{k}$ is consistent for $\gamma_{k}$.
Assumption (ref)(i) is guaranteed if $\gamma_{1},\gamma_{2}$ satisfy certain smoothness condition, e.g. $\gamma_{k}$ possesses uniformly bounded derivatives up to a sufficiently high order. Assumption (ref)(ii) requires that the first-stage estimator converges at a rate faster than $N^{-1/4}$, which is satisfied under various types of nonparametric estimators under certain regularity conditions. This is required so that the final estimator of the production function parameters $\alpha$ can converge at the standard parametric ($\sqrt{N}$) rate despite the slower first-step nonparametric estimation of $\gamma_{1},\gamma_{2}$.
Finally, we state another technical assumption that captures how the first-stage nonparametric estimation of $\gamma_{1},\gamma_{2}$ influences the final semiparametric estimators $\hat{\alpha}$ and $\hat{\alpha}^{*}$ through the functional derivatives of the residual function with respect to $\gamma_{1},\gamma_{2}$. Assumption (ref) below, based on \citet*{newey1994asymptotic}, provides an explicit formula for the asymptotic variances of $\hat{\alpha}$ and $\hat{\alpha}^{*}$ that does not depend on the particular forms of first-stage nonparametric estimators.
Formally, write $w_{i}:=\left(y_{i},x_{i},z_{i},d_{i}\right)$, $\gamma:=\left(\gamma_{1},\gamma_{2}\right)$, and suppress the conditioning variables $z_{i}$ in $\gamma$ for notational simplicity. Define the residual functions
for generic $\tilde{\alpha},\tilde{\gamma}$, and \[ g\left(w_{i},\tilde{\gamma}\right):=g\left(w_{i},\alpha,\tilde{\gamma}\right),\quad g^{\ast}\left(w_{i},\tilde{\gamma}\right):=g^{\ast}\left(w_{i},\alpha,\tilde{\gamma}\right), \] at the true $\alpha$. Define the pathwise functional derivative of $g$ at $\gamma$ along direction $\tau$ by \[ G\left(w_{i},\tau\right):=\lim_{t\rightarrow0}\frac{1}{t}\left[g\left(w_{i},\gamma+t\tau\right)-g\left(w_{i},\gamma\right)\right], \] and similarly define $G^{\ast}\left(z_{i},\tau\right)$ for $g^{\ast}$. Then, following \citet*{newey1994asymptotic}, the influence function can be derived analytically\footnote{See the proof of Theorem (ref) for details on the calculation.} based on $G$ and takes the form of $\varphi\left(w_{i}\right)\overline{z}_{i}\epsilon_{i}$ with
where $\gamma_{k}^{^{\prime}}$ denotes $\frac{\partial}{\partial h_{k}}\gamma_{k}\left(x_{ik};z_{i}\right)$, $\lambda_{1}$ stands for \[ \lambda_{1}\left(x_{i},z_{i}\right):=\mathbb{E}\left[\left.\mathbf{\mathbbm1}\left\{ d_{i}=1\right\} \right|x_{i},z_{i}\right] \] i.e., the conditional probability of observing $x_{i1}$, and $\lambda_{2}:=1-\lambda_{1}$.
The influence function essentially characterizes how the first-stage estimation influences the asymptotic variance of the final estimator. Formally, we present the following assumption, commonly known as an asymptotic linearity condition, which basically requires that the expected error induced by the first-stage estimation is asymptotically equivalent to the sample averages of $\varphi\left(w_{i}\right)\overline{z}_{i}\epsilon_{i}$ and $\varphi^{*}\left(w_{i}\right)\overline{z}_{i}\epsilon_{i}$. In particular, the formula for $\varphi$ and $\varphi^{*}$ given above will be the same regardless of the specific forms of first-step estimators used, provided that some suitable regularity conditions are satisfied.
We emphasize that Assumptions (ref) and (ref) are standard assumptions widely imposed in the semiparametric estimation literature, which can be satisfied by many kernel or sieve first-stage estimators under a variety of conditions. See \citet*{newey1994asymptotic}, \citet*{newey1994large} and \citet*{chen2003estimation} for references. In Assumption (ref) below, we also provide an example of lower-level conditions that replace Assumptions (ref) and (ref) when we use the Nadaraya-Watson kernel estimator in the first-stage nonparametric regression.
The next theorem establishes the asymptotic normality of $\hat{\alpha}$.
We note that, if the latent inputs were observed and the first-step nonparametric regression were not required, the asymptotic variance of standard IV estimator of $\alpha$ would be given by $\Sigma_{zx}^{-1}\text{Var}\left(\overline{z}_{i}\left(u_{i}+\epsilon_{i}\right)\right)\Sigma_{xz}^{-1}$. Hence, the presence of the additional term $\varphi\left(z_{i}\right)$ in $\Omega$ captures the effect of the first-step nonparametric regression on the asymptotic variance of $\hat{\alpha}$.
To obtain consistent variance estimators, define
where \[ \tilde{y}_{i}:=
\] and with
where $\hat{\lambda}_{1}$ is any consistent nonparametric estimator of $\lambda_{1}$. Then the variance estimators can be obtained as \[ \hat{\Sigma}:=S_{x\tilde{z}}^{-1}\hat{\Omega}S_{\tilde{z}x}^{-1} \] with $S_{z\tilde{x}}:=\frac{1}{N}\sum_{i=1}^{N}\overline{z}_{i}\tilde{x}_{i}^{^{\prime}}$.
Similarly, $\hat{\Omega}^{*}$ and $\hat{\Sigma}^{*}$ can be constructed accordingly.
If furthermore $\lambda_{1}\left(x_{i},z_{i}\right)\equiv\lambda_{1}\in\left(0,1\right)$ is assumed, then we may use the sample proportion $\hat{\lambda}_{1}:=\frac{1}{N}\sum_{i}\left\{ d_{i}=1\right\} $.
We also note that, when the first step takes the form of sieve lease squares, the simple procedure in \citet*{ackerberg2012practical}, in which we may “pretend” that the first stage is a parametric model, may be applied to obtain estimates of the asymptotic variance matrix.
Next, we compare the asymptotic variances of $\hat{\alpha}^{*}$ and $\hat{\alpha}$, and show that $\hat{\alpha}^{*}$ is in fact asymptotically more efficient.
The proof is in Appendix (ref). Here we discuss the intuition of Theorem (ref). The error term for the IV regression with the raw outcome $y_{i}$ as the left-hand-side variable is $u_{i}+\epsilon_{i}$, which has a larger variance than the corresponding error term $u_{i}$, if the conditionally expected outcome $\overline{y}_{i}$ is used instead. Even though we do not observe $\overline{y}_{i}$ and must use an estimator $\tilde{y}_{i}=\hat{\gamma}_{1}\left(x_{i1}\right)$ or $\tilde{y}_{i}=\hat{\gamma}_{2}\left(x_{i2}\right)$, the impact of the first-stage estimation error (which can be loosely thought as an average of $\epsilon_{i}$ across $i$) is smaller than the impact of $\epsilon_{i}$ itself.
To see this more clearly, first consider the multiplier “$1+\varphi\left(w_{i}\right)$” in (i): the “1” comes from the one “raw” share of error $\epsilon_{i}$ embedded in each $y_{i}$ that we use as the outcome variable, while “$\varphi\left(w_{i}\right)$” essentially captures the share of influence of the first-step estimation error $\hat{\gamma}-\gamma$ due to $\epsilon_{i}$. Together, we have \[ 1+\varphi=\left(1-\lambda_{1}\frac{\alpha_{2}}{\gamma_{2}^{^{\prime}}}+\lambda_{2}\frac{\alpha_{1}}{\gamma_{1}^{^{\prime}}}\right)\mathbf{\mathbbm1}\left\{ d_{i}=1\right\} +\left(\lambda_{1}\frac{\alpha_{2}}{\gamma_{2}^{^{\prime}}}+1-\lambda_{2}\frac{\alpha_{1}}{\gamma_{1}^{^{\prime}}}\right)\mathbf{\mathbbm1}\left\{ d_{i}=2\right\} , \] while the corresponding multiplier $\varphi^{*}$ on $\epsilon_{i}$ in (ii) is essentially the same except that “$1-\lambda_{1}\frac{\alpha_{2}}{\gamma_{2}^{^{\prime}}}$” becomes “$\lambda_{1}-\lambda_{1}\frac{\alpha_{2}}{\gamma_{2}^{^{\prime}}}$” and “$1-\lambda_{2}\frac{\alpha_{1}}{\gamma_{1}^{^{\prime}}}$” becomes “$\lambda_{2}-\lambda_{2}\frac{\alpha_{1}}{\gamma_{1}^{^{\prime}}}$”. Since $\lambda_{1},\lambda_{2}<1$, the overall multiplier on $\epsilon_{i}$ becomes smaller in magnitude\footnote{Note that $\alpha_{1}/\gamma_{1}^{^{\prime}}\leq1$ and $\alpha_{2}/\gamma_{2}^{^{\prime}}\leq1$ by equation (ref).}. Essentially, by using the estimated conditional expected output $\tilde{y}_{i}$, the raw “1” share of $\epsilon_{i}$ in $y_{i}$ is moved into the first-stage estimation error of $\overline{y}_{i}$, which is then “averaged” and reduced in magnitude to $\lambda_{1}$ or $\lambda_{2}$, thus leading to smaller overall variance.
Lastly, we emphasize that the efficiency comparison in (ref) does not directly relate to the theory of semiparametric efficiency bounds, such as in ackerberg2014asymptotic, which is about asymptotic efficiency of semiparametric estimators under a given criterion function. In fact, by ackerberg2014asymptotic, both estimators based on $y_{i}$ and $\tilde{y}_{i}$ attain their corresponding semiparametric efficiency bounds with respect to their\ different criterion functions $g$ and $g^{*}$. Theorem (ref), however, is a comparison across the two criterion functions $g$ and $g^{*}$: it essentially states that the asymptotically efficient estimator under $g^{*}$ is even more efficient than the efficient estimator under $g$.
Finally, we present a set of lower-level conditions that replace Assumptions (ref) and (ref), when we use the canonical Nadaraya-Watson kernel estimator for the nonparametric regression in Step 1. We emphasize that this subsection simply serves as an illustration of Assumptions (ref)-(ref) and Theorem (ref), as our method does not require the use of a specific form of first-step nonparametric estimators. For sieve (series) first-step estimators, similar results can be derived based on, for example, \citet*{newey1994asymptotic}, \citet*{chen2007large} and \citet*{chen2015sieve}.
Assumption (ref)(i) essentially requires that the proportion of observations with $x_{i1}$ observed and that with $x_{i2}$ observed are both strictly positive, or in other words, the numbers of both types of observations tend to infinity at the same rate of $N$. This guarantees that we can estimate both $\gamma_{1}$ based on observations with $x_{i1}$ and $\gamma_{2}$ based on observations with $x_{i2}$ well enough asymptotically. Assumption (ref)(iv) is the key smoothness condition that will help establish the Donsker property (and a consequent stochastic equicontinuity condition) in Assumption (ref)(i). Assumption (ref)(v)(vi) are concerned with the choice of kernel function $K$ and bandwidth parameter $b$: (v) requires that a “high-order” kernel function (of order $p$) is used, while (vi) requires that the bandwidth is set (with “under-smoothing” ) so that the kernel estimator $\text{\ensuremath{\hat{\gamma}_{k}}}$ converges at a rate faster than $N^{-1/4}$, as required in Assumption (ref)(ii). The requirement of $p\geq4$ in (iii) ensures that (vi) is feasible. Together with the additional regularity conditions in (ii)(ii), these conditions ensure that Assumptions (ref)-(ref) are satisfied. See \citet*[Section 8.3]{newey1994large} for additional details.
If additional instruments are available, it is straightforward to incorporate them in the second-stage regression, which will take the form of a two-stage least square estimator instead of an IV regression. Our results will carry over with suitable changes in notation. For example, the asymptotic variance formula for $\hat{\alpha}$ needs to be adapted as
Consider a potentially nonlinear parametric production function of the form \[ y_{i}=F_{\alpha}\left(x_{i1},x_{i2}\right)+u_{i}+\epsilon_{i} \] After the identification of partially latent inputs via Theorem (ref), the second stage boils down to the estimation of $\alpha$ based on the moment condition $\mathbb{E}\left[z_{i}\left(y_{i}-F_{\alpha}\left(x_{i1},x_{i2}\right)\right)\right]=\mathbf{0}$, which can be obtained via GMM estimation. Technically, since GMM estimators are Z-estimators, the corresponding asymptotic theory in \citet*{newey1994large}, on which the proof of Theorem (ref) is mainly based, still applies with proper changes in notation.
More generally, with any nonparametric production function that is additively separable in $u_{i}$ and $\epsilon_{i}$ of the form \[ y_{i}=F\left(x_{i1},x_{i2}\right)+u_{i}+\epsilon_{i}, \] where $F$ is an unknown function that satisfies Assumption (ref), the only thing that changes is the second-stage nonparametric estimation of $F$ with the imputed inputs $\tilde{x}_{i}$ (or more precisely, with one component known and one component imputed) based on the moment condition $\mathbb{E}\left[z_{i}\left(y_{i}-F\left(x_{i1},x_{i2}\right)\right)\right]=\mathbf{0}.$ The asymptotic theory for this case can be similarly obtained based on theory on nonparametric two-step estimation (e.g. \citealp*{ai2007estimation}, and \citealp*{hahn2018nonparametric}).
In the more general specification (ref): \[ y_{i}=F\left(x_{i1},x_{i2},u_{i}\right)+\epsilon_{i} \] where there is no more additive separability in $u_{i}$, one way to obtain identification and implement IV estimation is by adapting \citet*{chernozhukov2007instrumental} to our current context. Essentially, we would need to impose strict monotonicity of $F$ in $u_{i}$, impose independence of $u_{i}$ from $z_{i}$, normalize the distribution of $u_{i}$ to be uniform, and then exploit a quantile-based residual condition as described in \citet*{chernozhukov2007instrumental}.
Here we report the findings of some Monte Carlo experiments. Table (ref) reports the parameter specifications of the Cobb-Douglas production function that we use in our experiments. We assume that inputs are optimally chosen by a profit maximizing firm as discussed in detail in Appendix (ref). Specification 1 is the baseline specification. These parameters were chosen so that the simulated data are broadly consistent with the descriptive statistics of our application that we discuss in detail in Section (ref). Specification 2 has a larger variance in productivity shocks ($u_{i}$). Specification 3 has a smaller variance in wage distributions ($z_{i}$). Finally, specification 4 has a larger variance in measurement errors ($e_{i}$).
Specifically, we consider $L$ different markets, with each market containing $I$ firms, so that the total number of firms is $N=L\times I$. Firms in the same market $l$ all pay the same local wages, which we use as the instrumental variables. Local wages are drawn from a joint log-normal distribution with mean $\mu_{z}$ and variance $\sigma_{z}$ where wages for the two inputs are positively correlated. The firm-level idiosyncratic productivity shocks and the measurement errors are independently drawn from normal distributions with zero means and standard deviations $\sigma_{u}$ and $\sigma_{e}$, respectively. The productivity shocks and the measurement errors are independently and identically distributed both across firms and markets. We consider different configurations of $L$ and $I$: specifically, $L=$ 50, 100, 500 and $I=$ 1, 50, 100.
For each experiment, we compute the difference between the true parameter value and the sample average of the estimates using $M=1000$ replications, which is a measure of the bias of our estimator. We also estimate the root mean squared error (RMSE) using the sample standard deviation of our estimates.
Note that our data generating process mechanically implies $x_{i1}$ and $x_{i2}$ have a linear relationship with $y_{i}$. We estimate $\gamma_{1}\left(\cdot,z_{i}\right)$ and $\gamma_{2}\left(\cdot,z_{i}\right)$ using second degree polynomials. Not surprisingly, we find that the estimated coefficients on quadratic terms are almost 0. The interpolated functions $\gamma_{1}^{-1}$ and $\gamma_{2}^{-1}$ are also almost linear.
Table (ref) summarizes the performance of two different estimators: the two stage least squares estimator (TSLS) when all inputs are observed as well as our first version of TSLS when inputs are imputed and the observed output is used as the dependent variable.\footnote{Our second version of TSLS, which uses the expected output as the dependent variable, performs almost identically to the first version in this Monte Carlo simulation.} We refer to this version of the TSLS estimator as the “matched” TSLS estimator. As we would expect given our asymptotic results, the matched TSLS performs almost as well as the standard TSLS estimator under these ideal sampling conditions. This finding holds for all four different specifications and several choices for the number of firms within a market and the number of local markets.
Next, we investigate how our estimator performs when we have a relatively small number of observations in each market. Considering an extreme case, we simulate data for $L=500$ and $I=1$. In this case, as we only have a single firm in each market, we cannot impute the missing input variable using within market information. Instead, we pool observations across markets and estimate conditional expectations conditional on $x_{1}$ (or $x_{2}$), $z_{1}$, and $z_{2}$.\footnote{Note that the missing inputs are imputed for each market separately when $I\neq1$.} Table (ref) also summarizes the bias and RMSE where $L=500$ and $I=1$. We find that the matched TSLS estimator performs almost as well as the standard TSLS estimator that assumes that both inputs are observed. The only case where the matched TSLS estimator exhibits relatively large bias and RMSE is when the variance of the measurement errors is large (Spec 4). In Appendix (ref), we present a Monte Carlo experiment where we have partially latent wages.
We have not conducted a Monte Carlo analysis to evaluate how sensitive our estimator is to other misspecification errors. For example, it might be interesting to evaluate the performance of our estimator if input decisions (and hence output) are functions of both demand and supply shocks. We conjecture that the magnitude of the bias that will result from ignoring demand shocks will likely depend on the correlation structure of the demand and supply shocks and the market structure of the industry.
Our main application focuses on the estimation of team production functions. We introduce our data set and discuss our main empirical findings. Finally, we discuss other data sets that have a similar structure than the one we use and potential applications of our methods.
We study team production functions in pharmacies. This industry has undergone a dramatic change over the past decades. An industry that used to be primarily dominated by local independent pharmacies has been transformed by the entry of large chains that operate in multiple markets.
According to Goldin and Katz (2016) one important technological change in the pharmacy sector “is the extensive use of information technology systems and an increase in prescription drug insurance, which have both enhanced the ability of pharmacists to hand off clients. Improvements in information technology have enhanced the ability of pharmacists to leave a coherent and comprehensive record of each client, increasing the substitutability of pharmacists and reducing consumer preferences for particular pharmacists. Because of the increase in insurance coverage, pharmacists can access the prescriptions of clients through Pharmacy Benefit Managers (PBM) even if the scripts were not filled at that pharmacy... [Another] change is the standardization of pharmacy products and services. Medications have been increasingly produced by pharmaceutical companies rather than being compounded in individual pharmacies and hospitals. The greater standardization of medications has meant that the idiosyncratic expertise and talents of a particular pharmacist have become less important ... [these] changes make pharmacists better substitutes for each other and enable an almost costless handoff of clients." (p. 708-709)
An important question is the extent to which this transformation has been driven by technological change that has benefited large chains over smaller independently operated pharmacies. If this is in fact the case, these technological changes may help to explain why this profession has become so popular with females goldin2016most.
The main data set that we use is the National Pharmacist Workforce Survey of 2000 which is collected by Midwestern Pharmacy Research. The data comes from a cross-sectional survey answered by randomly selected individual pharmacists with active licenses. The data set is composed of two types of information: information about pharmacists and information about the pharmacy each pharmacist works at.
Information at the pharmacy level includes the type of pharmacy (Independent or\ Chain), the hours of operation per week, the number of pharmacists employed, and the typical number of prescriptions dispensed at the pharmacies per week. The store-level information is provided by an individual pharmacist who works at the pharmacy, thus the quality of the responses may depend on how knowledgeable the person is about the pharmacy. However, considering that most of the pharmacists in our sample are observed to be full-time pharmacists, the quality of the firm-level data is likely to be high. The number of prescriptions dispensed at the pharmacy is our measure of output. As a consequence, we do not have to use revenue based output measures which could bias our analysis as discussed, for example, in \citet*{epple2010new}.
Table (ref) summarizes the means of key variables that are observed at the firm or pharmacy level. After eliminating cases with missing input/output information, we observe 332 pharmacists. Table (ref) suggests that there are some pronounced differences between chains and independent pharmacies. Chains are more likely to be located in larger urban areas than independent pharmacies. They also operate longer hours per week. Interestingly, hourly production measured by the number of prescriptions per hour is, on average, similar to the independent pharmacies with similar employment size.\footnote{Most pharmacies in our sample have one manager pharmacist and one employee pharmacist, but there are a few pharmacies with a larger employment size.} We explore these issues in more detail below and test whether the different types of pharmacies have access to the same technology.
The survey also collects various information about pharmacists including hours of work, demographics, and household characteristics. Most importantly we observe the position at the pharmacy (Owner/Manager or \ Employee). We treat hours of the manager and hours of the employees as the two input factors in the production function.
Information related to the individual pharmacists is summarized in Table (ref). Employee pharmacists at independent pharmacies work fewer hours than the employee pharmacists at chain pharmacies, and hourly earnings are lower than those of the employees at the chains. Pharmacists in managerial positions at independent pharmacies work more hours than managers at chain pharmacies, but they have lower hourly earnings on average.
We observe only one pharmacy in each local labor market, which is defined as the 5-digit zip code area. Hence, we use the version of our estimator that averages across local markets as discussed in Section (ref). We use imputed hourly wages for the manager and the employer as the additional observed covariates ($z_{1},z_{2}$) for the first-stage estimation.\footnote{In this application, we only observe the wage for the observed type. Thus, wages are imputed using local demand shifters in 5-digit zip code levels and pharmacists' characteristics. We have verified with additional Monte Carlo simulation exercise reported in Appendix C that our method performs as well as the standard TSLS estimator with this variation.} In the second stage estimation, we use the observed wage for the observed position, the imputed wages, and principal components of local demand shifters as additional instruments.\footnote{The local demand shifters include total population size, median household income, and proportion of households with retirement income.}
We implement two versions of our “matched” TSLS estimator: the first estimator uses the observed outputs while the second one uses expected outputs. Since the observed output is subject to a measurement error, the semi-parametric estimator using expected outputs offers the potential of some efficiency gains as discussed in Theorem (ref). Table (ref) summarizes our findings. We report the estimated parameters of the Cobb-Douglas production function as well as the estimated standard errors. In addition, we report standard F-statistics for the first stage of the TSLS estimator to test for weak instruments. Overall, we find that our instruments are sufficiently strong in most cases.\footnote{As a robustness check, we also explored a different matching algorithm which estimates the expectation of output conditional on local demand shifters rather than wages. The results are consistent although the matching algorithm with local demand shifters gives slightly larger point estimates with slightly less precision.}
Table (ref) shows that we estimate most of parameters of the production function with good precision. Correcting for potential measurement error by using the expected output as the dependent variable, we achieve similar, maybe even slightly more plausible estimates.
Our results provide several insights into understanding the difference between independents and chains. First, our results indicate that chains may have a different production function than independent pharmacies. A formal joint hypothesis test reported in Table (ref) rejects the null hypothesis that the coefficients of the production function are the same.
Second, our findings also suggest that managers may be more effective in chains than independents. A formal one-sided t-test reported in Table (ref) rejects the null hypothesis that the two coefficients that characterize the marginal product of inputs are the same. Our findings imply that larger chains are more efficient in many other dimensions, such as task assignments or logistics, so that managers at larger chains may be able to focus more on prescription-related tasks compared to managers at independent pharmacies.
Finally, we find that chains have a significantly lower residual variance than independents. A formal F test reported in Table (ref) rejects the null hypothesis that the residual variance of independents is greater than or equal to the residual variance of chains. Note that all the tests are based on the estimation results with the expected outputs as the dependent variable.
We also test whether the observed labor inputs are indeed the optimal choice of firms. If the inputs are optimally chosen, the coefficients can be directly estimated from equation ((ref)) in Appendix (ref). Under the assumption of Cobb-Douglas production, we can test the optimality by jointly testing the null hypothesis of equality of both coefficients. Table (ref) shows the results. A formal Wald test rejects the null hypothesis of optimality. Thus, the direct inversion of the optimality conditions cannot be applied to estimate the parameters of the production function, whereas our new estimator is feasible. We note, however, that the test of the optimality of inputs here is built upon the premise that the linear model is correctly specified and all the assumptions underlying our method are valid.
Although most pharmacies in our sample have one manager and one pharmacist, there are a few pharmacies with more than one employee pharmacist. For this subset of pharmacies, we compute the total hours worked by employee pharmacists by multiplying the reported hours worked from an employee by the number of employees. Then, the second imputation step is applied based on the total hours worked by all employees. In this process, we implicitly assume the labor hours from two different employees are perfect substitutes. As a robustness check, we also estimate a version of production function which has an elasticity of substitution between the hours worked by different employees equal to one. Table (ref) summarizes this version of the estimation result. The estimated parameters show that employees have slightly lower marginal products at both independents and chains compared to our baseline estimation, but in general our estimation result is robust to how we treat employee inputs from pharmacies with more than one employee.
We thus conclude that chains have different production functions than independent pharmacies which may partially explain the change in the observed market structure of that industry.
One main concern with this application is that the variation of output (the number of prescriptions) in the pharmacy industry is also driven by demand shocks. We offer four observations.
First, our framework works as long as there is only one structural shock. Our model is agnostic about whether $u_i$ is a technology or a demand shock. However, the choice of instruments depends on the interpretation of the structural shock. In particular, wages are less plausible instruments if the structural shock is primarily a demand shock.
Second, our approach can also be applied in a model with two structural shocks as long as the second shock (the demand shock) is unexpected and iid across firms. The main results of the paper are still valid since input decisions are by assumption not affected by these unexpected demand shocks, and we use expected and not realized output in the matching procedure. Hence, we implicitly assume that demand shocks can be treated as we treat measurement error, i.e. as additively separable iid shocks.
Third, it is reasonable to consider a more general model with two structural shocks and measurement error. This generalized model can then account for the possibility that input decisions are functions of both demand and supply shocks. Moreover, these shocks do not necessarily enter into the production function in a simple additively separable way. As we discuss in detail in the conclusions of this paper, this models falls outside the scope of this paper. More research is needed to develop plausible identification strategies for this generalized model.
Finally, there are additional issues that arise when estimating production functions in retail compared to manufacturing. For example, capacity constraints, rationing based on wait time, and service quality may all be important. Unfortunately, we do not observe these variables in our data set. Hence, we cannot evaluate how important these factors are. More research is needed to fully address this important research question.
As we discussed above, the partially latent data structure that we study in this paper arises quite naturally in matched employer-employee data sets which contain information collected from individuals as well as information collected from businesses or establishments. Hence, they contain useful information about the firm such as output and revenues as well as the employees of the firm. However, the sampling design often implies that only a small subset of the employees of a firm are included in the sample. If the survey does not sample all employees in a firm then some important labor inputs are likely to be latent from the perspective of the econometrician.\footnote{For an early survey of employer-employee data sets see \citet*{Abowd-Kramaz-99}.}
Consider the Workers Establishment Characteristics Database (WECD) which matches long form respondents of the 1990 U.S. Census of Population to data on their employers from the Longitudinal Research Database (LRT). Note that the LRT not only has detailed output measures, but also includes measures of the capital stock, which is important in manufacturing. Since the long-form was only given to a random sample of approximately five percent of the population, we can typically only match five percent of the workers for large plants and a much lower percentage for some smaller plants. For example, \citet*{HNT-99} focused on large plants with more than 20 employees and rely on the aggregate data in the LRD to estimate the production function. Hence, they do not differentiate between different types of labor inputs, such blue and white collar workers. The methods developed in our paper allows researchers to use the data from the WECD to estimate the production function of small plants, even if we do not observe white and blue collar workers for each small plant.\footnote{The New Worker-Employer Characteristic Dataset is an extension of the WECD that also contains firms. Another relevant data set is the Longitudinal Employer-Household Dynamics (LEHD) which is drawn from state unemployment insurance administrative files, and hence, contains a larger subset of workers. It suffers from missing data problems, especially for firms with high turnover in the labor force. Hence, we often only observe a subset of workers for these firms. This is particularly problematic for small firms in the service sector.}
Another well-known employer-employee data set that applied this sampling design is the Workplace and Employee Survey (WES), a longitudinal survey of workplaces and their employees administered annually by Statistics Canada between 1999 and 2006. Every year, a representative sample of approximately 6,000 employers was surveyed. A maximum of twenty-four employees were randomly interviewed from each sampled workplace in each odd year and re-interviewed the following year. Thus, we only observe a small subsample of the employees at each firm in the WES.\footnote{See, for example, \citet*{Dostie-Javdani-20}, for more details.}
Other applications of our techniques can be based on surveys conducted by professional associations. Consider, for example, the After the JD , which is a nationally representative longitudinal data set constructed by the American Bar Foundation and the National Association for Law Placement. The data set follows students who graduated in 2000. The respondents were surveyed three times: once each in 2003, 2007, and 2012. The survey thus contains detailed information about associates and partners. The commonly available version of the After the JD does not contain the identity of the law firm. However, this information is available to internal researchers at the American Bar Association. Hence, one should be able to match the employees in the AJD to individual law firms and study the productivity differences among law firms using our technique.
Finally, there are numerous applications outside of industrial organization. In Appendix D show that the partially latent covariate problem can also arise in the study of intergenerational transmission of human capital, wealth or attitudes. It is rather common that we do not sample all relevant household members. Here, we consider, the Child Development Supplement of the PSID to study the transmission of human capital from parents to children. If a child grows up in traditional family, it is likely that we observe inputs from both mother and father. However, traditional family arrangements have become less common during the past decades and thus the focus has shifted on evaluating the impact of growing up in less traditional environments. If a child grows up in a non-traditional family one of the two parents' inputs are often missing. We can also apply our framework to study achievement functions. We find that there are some significant differences between married and divorced parents. In particular, divorced fathers have no significant impact on child quality.
We have developed a new method for identifying econometric models with partially latent covariates. We have shown that a broad class of econometric models that play a large role in industrial organization and labor economics can be non-parametrically identified if the partially latent covariates are monotonic functions of a common shock. Examples that fall into this class of models are production and skill formation functions. The partially latent data structure arises quite naturally in these settings if we employ an “input-based sampling” strategy, i.e. if the sampling unit is one of multiple labor input factors. It is plausible that the sampling unit will only have incomplete information about the other labor inputs that affect output. Our proofs of identification are constructive and imply a sequential, two-step semi-parametric estimation strategy. We have discussed the key problems encountered in estimation, characterized rate of convergence, and the asymptotic distribution of our estimators. Our application focuses on estimating team production functions. Using a national survey of pharmacists, we have found some convincing evidence that chains have different technologies than independently operated pharmacies. In particular, managers appear to be to have higher marginal products in chains.
Finally, our research provides ample scope for future research in econometric methodology. We have restricted ourselves to applications in which our method of identification can be combined with standard IV techniques to estimate the functions of interest. Much of the recent panel data literature has focused on dynamic inputs in the presence of adjustment costs, and more research is needed to extend the idea in this paper to a fully fledged dynamic panel data framework. We have also restricted ourselves to systems of inputs with a single common shock. Another potentially interesting research question is how our methods can be extended to more complicated econometric structures with multiple shocks.
\nocite{npws,psidcds}