Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
116,321 characters · 13 sections · 171 citation commands
Identification of Time-Varying Transformation Models with Fixed Effects, with an Application to Unobserved Heterogeneity in Resource Shares
We provide sufficient conditions for point-identification in a general class of fixed-$T$ time-varying nonlinear panel models. This class has a response variable equal to a time-varying weakly monotonic transformation of a linear index of regressors, fixed effects, and error terms. In contrast, almost all existing results for this class of models require time-invariance of the transformation. Our theorems imply novel identification results for time-varying versions of some commonly used models.
Specifically, we consider models of the type:
where $h_{t}$ is an unknown weakly monotonic transformation, $\alpha_{i}$ are unobserved individual-specific effects, $X_{it}$ is a vector of strictly exogenous observed explanatory variables with coefficients $\beta$, and $U_{it}$ is an error term drawn from a stationary distribution.
Our setting has the following four features:
The panel model literature is very large, with many papers showing identification of $\beta$, and sometimes of $h_{t}$, in models with two or three of the four features above (see, e.g., ArellanoBonhomme2012 for an overview). Ours is the first paper to study identification of models with these four features, all of which are demanded by our microeconomic model and empirical application. We refer to ((ref)) together with the four features above as the fixed-effects linear transformation (FELT) model, following the terminology in Abrevaya1999.
Our contribution is at least three-fold. First, we provide sufficient conditions for the identification of $h_{t}$ and $\beta$ for models in the FELT class. Second, for the case where $h_{t}$ is strictly monotonic, we provide results on the identification of some features of the distribution of fixed effects. Third, we provide a full-commitment intertemporal collective household model whose implied quantity demand equations lie in the FELT class, and estimate the model using Bangladeshi panel data.
We provide sufficient conditions for the point-identification of $h_{t}$ and $\beta$ in ((ref)) for two non-nested cases: one where $U_{it}$ is drawn from an arbitrary but stationary distribution, and one where it is drawn from the logistic distribution. Our approach provides a systematic way of analyzing models nested in the FELT class. An immediate implication of our work is that extensions to time-varying and/or nonparametric counterparts of well-known models can now be shown to be identified. For example, the following models are all nested in the FELT class and are now identified: ordered choice with time-varying thresholds; censored regression with time-varying censoring points; the multiple-spell generalized accelerated failure-time (GAFT) duration model; and the Box-Cox panel model with time-varying parameters.
For the case where $h_{t}$ is strictly monotonic we provide additional identification results for the conditional mean (up to location) and the conditional variance of the distribution of fixed effects. To the best of our knowledge, there are no results in the fixed-$T$, fixed-effects, nonlinear panel literature that cover this aspect of model identification.
Our theoretical work builds on established results from DoksumGasko1990 and Chen2002 who show that cross-sectional transformation models can in general be binarized into a set of related binary choice models. In this paper, we show that we can binarize in a panel setting when the transformation $h_{t}$ varies in an arbitrary way over time.
The key innovation underlying our theoretical work is time-varying binarization. For an arbitrary threshold $y_{t}$, we define the following binary random variable:
where $h_{t}^{-}$ is the generalized inverse of $h_{t}$ and where the equality follows from specification (ref) and weak monotonicity of $h_{t}$. Varying the threshold $y_{t}$ in ((ref)) across time periods converts a FELT model into a collection of binary choice models. This conversion is what we call time-varying binarization.\footnote{Muris uses time-varying binarization in a panel ordered logit model where the transformation is time-invariant and parametric.} However, to the best of our knowledge, previous papers that used binarization in a panel setting, e.g., Chen2010, ChernozhukovWP2018, restrict the thresholds to be equal across time periods, i.e. $y_{t}=y$ for all $t$. Essentially, it is the relaxation of this restriction that enables us to show identification of time-varying transformations $h_{t}$.
Once a FELT model has been converted into a collection of binary choice models via time-varying binarization, we invoke Chamberlain1980 and Manski1987 to show identification of the resulting binary choice models. We then re-assemble the identified models to obtain identification of $h_{t}$ and $\beta$ in the FELT model. Omitting the fact that any FELT model can be transformed into many binary choice models obtains identification of $\beta$ only.
We provide a full-commitment intertemporal collective household model that implies time-varying quantity demand functions of the form ((ref)). The nonlinear quantity demand functions in our model are time-varying because quantity demands depend on prices, and prices are unobserved but vary across the waves of our panel. The fixed effects in our economic model have a clear interpretation: they are logged resource shares, defined as the fractions of total household expenditure consumed by each of its members. Resource shares are not directly observable, but are important because unequal resource shares across household members signal within-household inequality. Our econometric results then imply that the quantity demand equations, the mean (up to location), and the variance of the resource shares are all identified. This is useful because it allows us to characterize the variation of, or inequality in, resource shares.
Our intertemporal collective household model ---along with the identification results above--- permits the use of short panel data to study resource shares within households. Previous cross-sectional methods to measure resource shares have imposed the identifying restriction that resource shares do not vary with household budgets, e.g., dlp13, and have relied on a random-effects model for unobserved heterogeneity in resource shares, e.g., Dunbarlp19. Our framework relaxes both these restrictions, and shows how panel data can enrich the study of resource shares. Ours are the first empirical estimates of a full-commitment collective household model in a short panel, and we demonstrate the importance of accounting for both observed and unobserved heterogeneity in women\textquoteright s resource shares.
Using a two-period Bangladeshi panel dataset on household expenditures, we show that less than half of the variation in women's resource shares can be explained by observed covariates. This means that there is much more inequality within households than would be suggested by variation in observed factors. We also find that women's resource shares are negatively correlated with household budgets. This means that women in poorer households have larger resource shares, and are therefore less poor than their household budgets would suggest. Further, that resource shares are found to covary with household budgets suggests caution in using the cross-sectional identifying restriction of independence suggested by dlp13.
Section (ref) provides a review of the related literature. In Sections (ref) and (ref), we provide our main identification results. Section (ref) introduces our collective household model, which uses data described in Section (ref). We present estimates of women's resource shares in Section (ref). All proofs, descriptive statistics for the data, additional robustness results and estimation details are in the Appendix.
We describe in detail how we connect to, first, the econometrics literature, and, second, to the literature on collective household models.
The literature on panel models is vast. Despite the vastness of this literature, we are not aware of any paper that delivers all four features discussed above. Below, we highlight the key differences between our approach and approaches in the literature that lack one or more of our key features.
Feature 1: We show identification in fixed-$T$ panel models. The incidental parameter problem occurs in fixed-effect panel models with a finite number of time periods, see NeymanScott1948.
Ours is a fixed-$T$ approach, with $n\to\infty$ and works even if $T=2$. A large literature analyzes the behavior of fixed effects procedures under the alternative assumption that the number of time periods goes to infinity, e.g., HahnNewey, ArellanoHahn2007, ArellanoBonhomme2009, FernandezVal2009, FernandezValWeidner2016, and ChernozhukovWP2018. In this setting, it is generally possible to identify each fixed effect, and consequently, the distribution of fixed effects. In our model, we show identification of specific moments of this distribution even though the number of time periods is fixed.
Feature 2: We allow weakly monotonic time-varying transformation $h_{t}$. Abrevaya1999 provides a consistent estimator of $\beta$ (the “leapfrog” estimator) in the FELT model under the restriction that the transformations are strictly monotonic.\footnote{SChen2010b in Remark 6 discusses a version of Abrevaya1999 that allows for some weak monotonicity due to censoring. He focuses on $\beta$ and does not discuss identification of $h_{t}$, although his Remark 1 sketches an approach for estimation of $h_{t}=h$ for all $t$.}
Abrevaya2000 considers a model that allows for weak monotonicity but restricts the transformations to be time-invariant (and allows for nonseparable errors). He provides a consistent estimator for $\beta$ only.\footnote{ChernozhukovWP2018 uses a distribution regression technique that is closely related to our binarization approach, and consequently accommodates weakly monotonic transformations. However, theirs is a large-$T$ setting.} A literature on duration models also considers time-invariant transformations that are weakly monotonic due to censoring, e.g., Lee2008, Khan2007, Chen2010,Chen2010a,Chen2012,Chen2012ET, and Chen2012; we review this below.
A more recent literature has focused on identification issues in a class of panel models with potentially non-monotonic but time-invariant structural functions (or strong assumptions on how those functions vary over time), e.g., HoderleinWhite2012, ChernValHahnNewey2013, ChernValHoderleinHolzNewey2015. These papers focus on (partial) identification of partial effects, and the approaches employed there preclude identification of the structural function(s) or of the distribution of fixed effects.
Feature 3: We allow nonparametric transformations and nonparametric errors. Bonhomme2012 proposes a general-purpose likelihood-based approach to obtain identification for models with parametric $h_{t}$ and parametric $U_{it}$, even allowing for dynamics.
Our model requires strictly exogenous regressors, precluding many dynamic structures. But, our Theorems 1 and 2 apply even when $h_{t}$ is nonparametric, $U_{it}$ is nonparametric, or both are nonparametric.
The setting with parametric transformations and parametric errors covers many models previously shown to be identified, including the time-invariant fixed-effect panel versions of: binary choice (e.g., Rasch1960, Chamberlain1980, Magnac2004, and Chamberlain2010); the linear regression model with normal errors; and the ordered logit model (e.g., das_panel_1999,Baetschmann2015,Muris). Application of our results immediately shows identification of the time-varying versions of these models. This result is novel for the ordered logit model, where our results imply identification of time-varying thresholds.
Parametric transformation models with nonparametric errors are widely studied, starting with Manski1987 for the binary choice fixed effects model. (Aristodemou2020 provides partial identification results for ordered choice with nonparametric errors.) Parametric panel data censored regression models also fit into our framework, and were studied intensively starting with Honore1992, see also, e.g., Charlier2000, HonoreKyriazidou2000Ecta, Chen2012ET. These papers show identification of $\beta$ for the linear model with time-invariant censoring and nonparametric errors. In this context, our results show identification of models that were not previously known to be identified. In particular, the model is identified even if the transformation is nonparametric (as opposed to linear or Box-Cox) and time-varying and/or where the censoring cutoff is time-varying.\footnote{Many papers in the literature on censored regression have focused on endogeneity. For example, HonoreHu2004 allow for endogenous covariates, and KhanPonomarevaTamer2016 study the case of endogenous censoring cutoffs. Our results do not cover the case of endogenous regressors or cutoffs. HorowitzLee2004 and Lee2008 consider dependent censoring, where the censoring cutoff depends on observed covariates and the error term follows a parametric distribution. We do not consider dependent censoring.}
Duration models can be recast as transformation models with nonparametric transformations (see Ridder1990). Consequently, the large literature on identification of duration models is related to our work.
Consider the multiple-spell mixed proportional hazards (MPH) model with spell-specific baseline hazard, analyzed in Honore1993.\footnote{HorowitzLee2004 show identification of this model under the restriction that the baseline hazard is the same for all spells, analogous to time-invariant $h_{t}$. Chen2010a considers the same model, but relaxes the restriction that errors are type 1 EV, but shows identification of only the common parameter vector $\beta$.} This model can be obtained from FELT by letting (i) $h_{t}^{-1}(v)=\log\left\{ \int_{0}^{v}\lambda_{0t}\left(u\right)du\right\} $, where $\lambda_{0t}$ is the baseline hazard for spell $t$, (ii) $\alpha_{i}$ and $U_{it}$ are independent across $t$, and (iii) $U_{it}$ is independent of $X_{i}$ and distributed as EV1. Honore1993 derives sufficient conditions for the identification of this model (Lee2008 provides consistent estimators under other parametric error distributions). Our theorems immediately provide the novel result that this model is identified when the error terms are drawn from a nonparametric distribution.
Consider the single-spell generalized accelerated failure time (GAFT) model introduced by Ridder1990 (see also vandenBerg2001) that has non EV1 errors, and is consistent with a duration model. Just like the MPH model, it can be extended to a multiple-spell setting, see, e.g., Evdokimov2011. Abrevaya1999 shows that $\beta$ in the multiple-spell GAFT model is consistently estimated. However, he does not show identification of the transformation $h_{t}$, which can be seen as dual to identification of the spell-specific baseline hazard function.\footnote{Khan2007 establish consistency of an estimator of the regression coefficient in GAFT under the restriction that the baseline hazard is the same for all spells, analogous to time-invariant $h_{t}$. } Evdokimov2011 considers identification of a related version of the multiple-spell GAFT with spell-specific baseline hazard, but requires continuity of $\alpha_{i}$ and at least 3 spells ($T\geq3)$. Our results show identification of both $\beta$ and $h_{t}$ in the multiple-spell GAFT model, imposing no restrictions on $\alpha_{i}$ and requiring just 2 spells ($T=2$).
Feature 4: We allow for unrestricted fixed effects. A related literature considers restrictions on the joint distribution of $\left(\alpha_{i},X_{i1},...,X_{iT}\right)$. For example, AltonjiMatzkin2005 impose exchangeability on this joint distribution, and BesterHansen2009 restricts the dependence of $\alpha_{i}$ on $\left(X_{i1},...,X_{iT}\right)$ to be finite-dimensional. In our model this joint distribution is unrestricted.
A further group of papers establishes identification of panel models, including the distribution of $\alpha_{i}$, by using techniques from the measurement error literature that: (i) impose various assumptions on $\alpha_{i}$, such as full support and/or continuous distribution; (ii) assume serial independence of $U_{it}$; and (iii) restrict the conditional distribution of $\left(\alpha_{i},U_{i1},\dots,U_{iT}\right)$ conditional on $\left(X_{i1},...,X_{iT}\right)$, see, Evdokimov2010, Evdokimov2011, Wilhelm2015, and Freyberger18. In contrast, our results on the identification of the conditional variance of $\alpha_{i}$ do not require (i). All our other results, including identification of the dependence of $\alpha_{i}$ on observed covariates, are free of assumptions like (i), (ii) and (iii).
Finally, special regressor approaches (see the review in Lewbel2014) have identifying power in transformation models with fixed effects. They require the availability of a continuous variable that is independent of the fixed effects. With such a variable, one can show identification of transformation models in the cross-sectional case (ChiapporiKomunjerKristensen2015) and in the panel data case, e.g., HonoreLewbel2002, AiGan2010, lewbelyang2016, chen_exclusion_2019. Our results do not invoke a special regressor. Further, we are not aware of any special regressor-based papers that identify time-varying transformations or the distribution of fixed effects.\footnote{We conjecture that the existence of a special regressor would be sufficient to identify time-varying nonparametric transformations, and, with strict monotonicity, the distribution of fixed effects. However, we think that a setting with completely unrestricted fixed effects is useful in a variety of empirical applications, including our own.}
We also show identification of some aspects of the distribution of fixed effects. Correlated random effects models identify the distribution of individual effects, but at the cost of restricting their distribution. To our knowledge, we are the first to show identification of moments of this distribution in a nonlinear panel model, when that distribution is unrestricted.
We show the practical importance of these innovations in our empirical work below. Identification of the conditional mean and variance of the distribution of fixed effects in a context with time-varying transformations is essential to our investigation of women's access to household resources in rural Bangladeshi.
Dating back at least to becker62, collective household models are those in which the household is characterized as a collection of individuals, each of whom has a well-defined objective function, and who interact to generate household level decisions such as consumption expenditures. Efficient collective household models are those in which the individuals in the household are assumed to reach the (household) Pareto frontier. Chiappori88,Chiappori92 showed that, like in earlier results in general equilibrium theory, the assumption of Pareto efficiency is very strong. Essentially, it implies that the household can be seen as maximizing a weighted sum of individual utilities, where the weights are called Pareto weights. This in turn implies that the household-level allocation problem is observationally equivalent to a decentralized, person-level, allocation problem.
In this decentralized allocation, each household member is assigned a shadow budget. They then demand a vector of consumption quantities given their preferences and their personal shadow budget, and the household purchases the sum of these demanded quantities (adjusted for shareability/economies of scale and for public goods within the household). For the special case of an assignable good, which is demanded by a single known household member, the household purchases exactly what that person demands given their shadow budget.
Resource shares, defined as the ratio of each person's shadow budget to the overall household budget, are useful measures of individual consumption expenditures. If there is intra-household inequality, these resource shares would be unequal. Consequently, standard per-capita calculations (assigning equal resource shares to all household members) would yield invalid measures of individual consumption and poverty (see, e.g., dlp13). In this paper, we show identification of the conditional mean (up to location) and conditional variance of the distribution of resource shares in a panel data context.
There are many ways to identify resource shares with cross-sectional data. A common identifying assumption (used by, e.g., dlp13) is that resource shares are independent of household budgets in a cross-sectional sense. This identifying restriction has been used to estimate resource shares, within-household inequality and individual-level poverty in many countries (dlp13 and Dunbarlp19 in Malawi; Bargain2014 in Cote d'Ivoire; Calvi in India; DeVreyer2016 in Senegal; Bargain2018 in Bangladesh). In our model, we show identification of the response of the conditional mean of resource shares to observed covariates, even if resource shares are correlated with (lifetime) household budgets. Consequently, we can test this identifying restriction.
dlp13 does not accommodate unobserved heterogeneity in resource shares. Two newer papers, cKim17 and DunbarLewbelPendakur2019 consider identification in cross-sectional data with unobserved household-level heterogeneity in resource shares. Like cKim17 and Theorem 1 in DunbarLewbelPendakur2019, our work investigates identification of the distribution of resource shares up to an unknown normalization. However, the results in those papers are of the random effects type. That is, the authors impose the restriction that the conditional distribution of resource shares is independent of the household budget. In this paper, we consider a panel data setting with household-level unobserved heterogeneity in resource shares, without any restriction on the distribution of resource shares. Further, we show sufficient conditions for identification of the conditional variance of (logged) resource shares.
The literature cited above considered one-period micro-economic models. But, many interesting questions about households, and the distribution of resources within households, are dynamic in nature. For example: how do household members share risk?; how do household investments relate to individual consumption?; how can we use information from multiple time periods to estimate resource shares when there is unobserved household-level heterogeneity?
ChiapporiMazzocco2017 review the literature on collective household models in an intertemporal setting. These models generally come in two flavours--limited commitment or full commitment--depending on whether or not the household can commit to a permanent Pareto weight at the moment of household formation. Full-commitment models answer \textquotedblleft yes\textquotedblright , and limited-commitment model answer \textquotedblleft no\textquotedblright . Limited commitment models have commanded the most theoretical attention. Much effort has gone into testing the full-commitment model against a limited-commitment alternative, e.g., Ligon98,Mazzocco07,MazzocoRuizYamaguchi14,Voena15.
Fewer papers study the identification of Pareto weights or resource shares in an intertemporal context. LiseYamada19 use a long panel of Japanese household consumption data to estimate how Pareto weights (which are dual to resource shares) depend on observed covariates and on unanticipated shocks. Their model does not allow for correlated unobserved heterogeneity, and they require many observations for each household (so as to see when Pareto weights change). They find evidence that Pareto weights do change, so that the full-commitment model does not hold in Japan.
Our model is one of full-commitment, allows for correlated unobserved household-level heterogeneity, and is identified in a short (e.g., 2 period) panel. So we provide a complement to the approach of LiseYamada19 for cases where the data are not rich enough to estimate a limited-commitment model. The key cases here are where the panel is too short to see changes in Pareto weights, or where the observed covariates leave too much room for unobserved heterogeneity. The cost of covering these cases is the assumption of full commitment.
Full commitment models are more restrictive, but may be useful nonetheless. ChiapporiMazzocco2017 write \textquotedblleft In more traditional environments (such as rural societies in many developing countries), renegotiation may be less frequent since the cost of divorce is relatively high, threats of ending a marriage are therefore less credible, and noncooperation is less appealing since households members are bound to spend a lifetime together.\textquotedblright We use a full commitment setting to estimate resource shares for rural Bangladeshi households.
In this paper, we adapt the general full-commitment framework of ChiapporiMazzocco2017 to the scale economy and sharing model of BrowningChiapporiLewbel13. Then, like dlp13 do in their cross-sectional analysis, we identify resource shares on the basis of household-level demand functions for assignable goods
. In our general model, observed household-level quantity demand functions for assignable goods depend on resource shares, and resource shares depend a time-invariant factor (a fixed effect) representing the initial (and permanent) Pareto weights of household members.
We then provide a parametric form for utility functions that results in demand equations for assignable goods that are nonlinear in shadow budgets, and have logged shadow budgets that are linear in logged household budgets and a fixed effect. Further, demand equations are time-varying because prices vary over time. Such demand equations fall into the FELT class, and are therefore identified in our short-panel setting. The parametric model also gives meaning to the fixed effect: it equals a logged resource share, so its distribution is an economically interesting object. So, our micro-economic theory demands an econometric model that allows for time-varying transformations and that can identify moments of the conditional distribution of fixed effects.
Dropping the $i$ subscript, let $Y=\left(Y_{1},...,Y_{T}\right)^{'}$ and $X=\left(X_{1}',...,X_{T}'\right)^{'}$. We then rewrite FELT as a latent variable model:
and denote the supports of $Y_{t},\,Y_{t}^{*},\,X_{t}$ by $\mathcal{Y}\subseteq\mathbb{R}$, $\mathcal{\mathcal{\mathcal{Y}}^{\textnormal{*}}=\mathbb{R}},$ and $\mathcal{X}\subseteq\mathbb{R}^{K}$, respectively.\footnote{The supports may be indexed by $t$.}
We provide sufficient conditions for identification of $\left(\beta,h_{t}\right)$.\footnote{The results in this section were previously circulated in the working paper BotosaruMuris. That paper also introduces four estimators, depending on whether the outcome variable is discrete or continuous, and on whether the stationary distribution of the error term is nonparametric or logistic. In this paper, we use a GMM estimator instead.} We consider two non-nested cases. The first case does not impose parametric restrictions on the distribution of $U_{t}$, requiring only that it is conditionally stationary. In this case, the idiosyncratic errors may be serially dependent and heteroskedastic. The second case assumes that $U_{t},\,t=1,\cdots,T$, are serially independent, standard logistic, and strictly exogenous.
It may appear that the second case is a special case of the first. However, the second case requires weaker assumptions on the distribution of the regressors (c.f. Assumption (ref) below) while imposing stronger assumptions on the error distribution. For both cases, we maintain the assumption below:
This assumption allows us to work with the generalized inverse $h_{t}^{-}:\mathcal{Y}\rightarrow\mathcal{Y}^{*}$, defined as: \[ h_{t}^{-}\left(y\right)\equiv\inf\left\{ y^{*}\in\mathcal{\mathcal{Y}^{\textnormal{*}}}:\mbox{ }y\leq h_{t}\left(y^{*}\right)\right\} , \] with the convention that $\text{inf}\left(\emptyset\right)=\text{inf}\left(\mathcal{Y}\right)$.
It is well-known (see e.g. DoksumGasko1990 and Chen2002) that cross-sectional transformation models can be binarized into a set of binary choice models. Binarization has been used previously in panel settings (see e.g. Chen2010, ChernozhukovWP2018), but those approaches have restricted the threshold to be equal across time periods. An exception is Muris, who uses time-varying thresholds in a panel ordered logit model with a time-invariant and parametric transformation.
We now describe time-varying binarization. For an arbitrary threshold $y_{t}\in\underline{\mathcal{Y}}\equiv\mathcal{Y}\backslash\inf\mathcal{Y}$,\footnote{We use $\underline{\mathcal{Y}}$ instead of $\mathcal{Y}$ because $D_{t}\left(\inf\mathcal{Y}\right)=1$ almost surely for all $t$.} we define the following binary random variable:
where the equality follows from specification (ref) and weak monotonicity of $h_{t}$. Varying the threshold $y_{t}$ in ((ref)) across time periods converts any FELT model into a collection of binary choice models.
Two time periods are sufficient for our identification results, so we let $T=2$ in what follows. Consider the vector of binary variables \[ D\left(y_{1},y_{2}\right)\equiv\left(D_{1}\left(y_{1}\right),D_{2}\left(y_{2}\right)\right), \] for any two points $(y_{1},y_{2})\in\mathcal{\underline{\mathcal{Y}}}^{2}$. Our identification strategy for $\left(\beta,h_{1},h_{2}\right)$ is based on the observation that $D\left(y_{1},y_{2}\right)$ follows a fixed effects binary choice model for any $(y_{1},y_{2})\in\mathcal{\underline{\mathcal{Y}}}^{2}$. This result is summarized in Lemma (ref) below.
The proof of identification proceeds in three steps. First, we show identification of $\beta$ and of $h_{2}^{-}\left(y_{2}\right)-h_{1}^{-}\left(y_{1}\right)$ for arbitrary $\left(y_{1},y_{2}\right)\in\mathcal{\underline{\mathcal{Y}}}^{2}$. In the resulting binary choice model relating $D_{t}\left(y\right)$ to $X$, the difference $h_{2}^{-}\left(y_{2}\right)-h_{1}^{-}\left(y_{1}\right)$ is the coefficient on the differenced time dummy, and $\beta$ is the regression coefficient on $X_{2}-X_{1}$. For a given binary choice model, identification of $\beta$ and of $h_{2}^{-}\left(y_{2}\right)-h_{1}^{-}\left(y_{1}\right)$ follows Manski1987 for the nonparametric version of FELT, and Chamberlain2010 for the logistic version. This result is summarized in Theorem (ref) below.
Second, we show that varying the pair $\left(y_{1},y_{2}\right)$ over $\underline{\mathcal{Y}}^{2}$ obtains identification of \[ \left\{ h_{2}^{-}\left(y_{2}\right)-h_{1}^{-}\left(y_{1}\right),\,\left(y_{1},y_{2}\right)\in\underline{\mathcal{Y}}^{2}\right\} . \]
Third, we show that identification of this set of differences obtains identification of $h_{1}$ and $h_{2}$ under a normalization assumption on $h_{1}^{-}$.
This result is presented in Theorem (ref).
Figures 3.1 and 3.2 illustrate the intuition behind our identification strategy for two arbitrary functions, $h_{1}$ and $h_{2}$, both accommodated by FELT. The line with kinks and a flat part represents an arbitrary function $h_{1}$, while the solid curve represents an arbitrary function $h_{2}$. Consider Figure 3.1. Pick a $y_{1}\in\mathcal{\underline{\mathcal{Y}}}$ on the vertical axis. For all $y\leq y_{1}$, $h_{1}\left(y\right)$ gets mapped to zero, while for all $y>y_{1}$, it gets mapped to one. Now pick a $y_{2}\in\underline{\mathcal{Y}}$. For all $y\leq y_{2}$, $h_{2}\left(y\right)$ gets mapped to zero, while for all $y>y_{2}$, it gets mapped to one. This gives rise to a fixed effects binary choice model for $\left(D_{1}\left(y_{1}\right),D_{2}\left(y_{2}\right)\right)$, also plotted in the figure as the grey solid lines. Our first result in Theorem (ref) identifies the difference $h_{1}^{-}\left(y_{1}\right)-h_{2}^{-}\left(y_{2}\right)$ at arbitrary points $\left(y_{1},y_{2}\right)$, as well as the coefficient $\beta$. It is clear that normalizing $h_{1}^{-}\left(.\right)$ at an arbitrary point identifies the function $h_{2}^{-}\left(y_{2}\right)$ at an arbitrary point $y_{2}$. This is captured in Figure 3.2. There, for an arbitrary $y_{0}\text{, }h_{1}^{-}\left(y_{0}\right)=0.$ Then, as $y_{2}$ is arbitrary, Figure 3.2 shows that moving $y_{2}$ on its support traces out the generalized inverse $h_{2}^{-}$ on its domain. Theorem (ref) wraps up this argument by showing that $h_{1}$ and $h_{2}$ are identified from their generalized inverses.
In this section, we provide nonparametric identification results for $\left(\beta,h_{1},h_{2}\right)$. Parts of our identification proof build on Manski1987, who in turn builds on manski_maximum_1975,Manski1985.
Assumption (ref) places no parametric distributional restrictions on the distribution of $U_{it}$ and allows the stochastic errors $U_{it}$ to be correlated across time. The first part of the assumption, (ref)(i), is a stationarity assumption, requiring time-invariance of the distribution of the error terms conditional on the trajectory of the observed regressors and on the unobserved heterogeneity. This assumption excludes lagged dependent variables as covariates. Additionally, as noted by, e.g., Chamberlain2010, although it allows for heteroskedasticity, it restricts the relationship between the observed regressors and $U_{it}$ by requiring that even when $x_{1}\not=x_{2}$, $U_{1}$ and $U_{2}$ have equal skedasticities. This type of stationarity assumption is common in linear and nonlinear panel models, e.g., ChernValHahnNewey2013 and references therein.
Assumption (ref)(ii) requires full support of the error terms. It guarantees that, for any pair $\left(y_{1},y_{2}\right)\in\underline{\mathcal{Y}}^{2}$, the probability of being a switcher is positive. In our context, being a switcher refers to the event $D_{1}\left(y_{1}\right)+D_{2}\left(y_{2}\right)=1$, so that Assumption (ref) guarantees that $P\left(D_{1}\left(y_{1}\right)+D_{2}\left(y_{2}\right)=1\right)>0$. This assumption is similar to Assumption 1 in Manski1987.
Let $\Delta X\equiv X_{2}-X_{1}\,$ and for an arbitrary pair $(y_{1},y_{2})\in\mathcal{\underline{Y}}^{2},$ define
Let $W\equiv(\Delta X,-1)'$ and $\theta\left(y_{1},y_{2}\right)\equiv\left(\beta,\gamma\left(y_{1},y_{2}\right)\right),$ so that $\left(\ref{eq:med_sgn}\right)$ can be written as \[ \text{med}\left(D_{2}\left(y_{2}\right)-D_{1}\left(y_{1}\right)|X,D_{1}\left(y_{1}\right)+D_{2}\left(y_{2}\right)=1\right)=\text{sgn}\left(W\theta\left(y_{1},y_{2}\right)\right). \] For identification of $\theta\left(y_{1},y_{2}\right)$ we impose the following additional assumptions.
Assumption (ref)(i) requires that the change in one of the regressors be continuously distributed conditional on the other components. Assumption (ref)(ii) is a full rank assumption. These assumptions are standard in the binary choice literature concerned with point identification of the parameters.
Assumption (ref) resembles Assumption 2 in Manski1987, the difference being that our assumption concerns $W$, which includes a constant that captures a time trend. The presence of this constant requires sufficient variation in $X_{t}$ over time. No linear combination of the components of $X_{t}$ can equal the time trend.
Assumption (ref) imposes a normalization on $\beta$, namely that the norm of the regression coefficient equals 1. Scale normalizations are standard in the binary choice literature, and are necessary for point identification when the distribution of the error terms is not parameterized. Normalizing $\beta$ (instead of $\theta$) avoids a normalization that would otherwise depend on the choice of $\left(y_{1},y_{2}\right)$. In this way, the scale of $\beta$ remains constant across different choices of $\left(y_{1},y_{2}\right)$. Alternatively, one can normalize the coefficient on the continuous covariate (cf. Assumption (ref)(i)) to be equal to one. In our economic model in Section (ref) the latter assumption holds automatically.\footnote{There are models with sufficient structure on the transformation $h_{t}$ where identification is possible without a normalization on the regression coefficient. Examples include the linear regression model, the censored linear regression model in Honore1992, and the interval-censored regression model in abrevaya_interval_2020.}
So far, we have identified the regression coefficient $\beta$ and the difference in the generalized inverses at arbitrary pairs $\left(y_{1},\,y_{2}\right)$. We consider now identification of the functions $h_{1}$ and $h_{2}$ on $\underline{\mathcal{Y}}$.
Such a normalization is standard in transformation models, see, e.g., Horowitz1996. Without this normalization, all identification results hold up to $h_{1}^{-}\left(y_{0}\right)$. We normalize the function in the first time period only, imposing no restrictions on the function in the second period beyond that of weak monotonicity (cf. Assumption (ref)). In Section (ref), we show that this normalization assumption is not necessary for our results on the identification of the conditional mean or of the conditional variance of the fixed effects.
In this section, we show identification of $\left(\beta,h_{1},h_{2}\right)$ when the error terms are assumed to follow the standard logistic distribution. The logistic case is not nested in the nonparametric case. In particular, when the errors are logistic, we do not require a continuous regressor. However, we require conditional serial independence of the error terms.\footnote{See Chamberlain2010 and Magnac2004 for more details about identification under nonparametric versus logistic errors in the panel data binary choice context.}
Assumption (ref)(i) strengthens Assumption (ref) by requiring the errors to follow the standard logistic distribution and to be serially independent. Note that one consequence of this assumption, which specifies the variance of the error terms to be equal to 1, is to eliminate the need to normalize $\beta$. On the other hand, Assumption (ref)(ii) imposes weaker restrictions on the observed covariates relative to Assumption (ref), since it does not require the existence of a continuous covariate. Sufficient variation in $\Delta X$ is sufficient to obtain identification of the vector $\beta$ when the error terms follow the standard logistic distribution.
If $(h_{1},h_{2})$ are invertible, we can use the previous identification theorem to identify features of the distribution of the fixed effects conditional on observed regressors. These features are the change in the conditional mean function of $\alpha$ and the conditional variance of $\alpha$ conditional on $X_{1},X_{2}$. These results are relevant since in our collective household model, the fixed effects represent the log of resource shares, and both the standard deviation of these resource shares and the response of their conditional mean to covariates are key parameters of interest in the empirical literature. As this is relevant to our application, we note here that a normalization assumption, such as (ref), on the demand function in the first period is not necessary for these results on the resource shares because, e.g., we only need their deviation with respect to the mean of the fixed effects.
In this section, we provide sufficient conditions for the identification of the change in the conditional mean function of the fixed effects, defined as:
as well as for the conditional variance of the fixed effects. For these results, the normalization assumption (ref) is not necessary. To provide intuition for this, let \[ c_{1}\equiv h_{1}^{-1}\left(y_{0}\right), \] at an arbitrary $y_{0}\in\mathcal{\underline{Y}}$ and $g_{t}\left(y\right)\equiv h_{t}^{-1}\left(y\right)-c_{1}$ for all $y\in\mathcal{\underline{Y}}$. Note that Theorem (ref) recovers \[ \widetilde{U}_{t}\equiv\alpha-U_{t}=h_{t}^{-1}\left(Y_{t}\right)-X_{t}\beta, \] up to $c_{1}$, so that the joint distribution of $\left(\widetilde{U}_{1},\widetilde{U}_{2},X_{1},X_{2}\right)$ is identified up to $c_{1}$. By placing restrictions on the distribution of $\left(\alpha,U_{1},U_{2},X_{1},X_{2}\right)$, we can then recover our features of interest.
Second, define the conditional variance of the fixed effects as
For this second result, we strengthen our assumptions to include, among others, serial independence of the error term. This allows us to pin the persistence in unit $i$'s time series on $\alpha_{i}$ instead of on serial dependence in the errors.
It may be possible to obtain the entire conditional distribution of the fixed effects under the assumption that $\left(\alpha,U_{1},U_{2}\right)$ are mutually independent by using arguments similar to those in AB_distributional_2012.
In this section, we construct a model of an efficient full-commitment intertemporal collective (FIC) household. We provide sufficient conditions implying that observed household-level quantity demand equations are time-varying functions of fixed effects that lie in the FELT class.
Essentially,
we
adapt the intertemporal collective household model of ChiapporiMazzocco2017 to the model of household scale economies given in the static household model of BrowningChiapporiLewbel13. We assume efficiency, i.e. that the household members together reach the Pareto frontier. Consequently, our model does not account for inefficiency due to e.g. consumption externalities or information asymmetries.
We use subscripts $i,j,t$. Let $i=1,...,n$ index households and assume the household has a time-invariant composition, with $N_{ij}$ members of type $j$. Let $j=m,f,c$ for men, women and children. Let $t=1,2$. Let $z$ be a vector of time-varying household-level demographic characteristics, and let the numbers of household members of each type, $N_{im}$, $N_{if}$ and $N_{ic}$, be (time-invariant) elements of $z_{it}$. Like ChiapporiMazzocco2017, this is a model with uncertainty, so we use the superscript $s=1,2$ to index states in the second period only. As in ChiapporiMazzocco2017, the use of 2 time periods and 2 states is for illustration only; the model goes through with any finite number of states or periods. Similarly, inclusion of a risky asset would not change the features of the model that we use.
Indirect utility, $V_{j}(p,x,z)$, is the maximized value of utility given a budget constraint defined by prices $p$ and budget $x$, given characteristics $z$. Let $V_{j}$ be strictly concave in the budget $x$. Indirect utility depends on time only through its dependence on the budget constraint and time-varying demographics $z$. Let $v_{ijt}\equiv V_{j}(p_{t},x_{it},z_{it})$ denote the utility level of a person of type $j$ in household $i$ in period $t$.
Household decisions are made on the basis of individual shadow budget constraints, reflecting the economic environment within the household. These shadow constraints are characterized by a shadow price vector faced by all members and shadow budgets which may differ across household members. Let $A_{it}\equiv A(z_{it})$ be a diagonal matrix that gives the shareability of each good, and let it depend on demographics $z_{it}$ (including the numbers of household members). More shareable goods have lower shadow prices of within-household consumption.
For nonshareable goods, the corresponding element of $A_{it}$ equals $1$; for shareable goods, it is less than $1$, possibly as small as $1/N_{i}$ where $N_{i}$ is the number of household members. Goods may be partly shareable, with an element of $A_{it}$ between $1/N_{i}$ and $1$. With market prices $p_{t}$, within-household shadow prices are given by the linear transformation $A_{it}p_{t}$.
BrowningChiapporiLewbel13 also allow for inequality in shadow budgets. Let $\eta_{ijt}$ be the resource share of type $j$ in household $i$ in time period $t$. It gives the fraction of the household budget consumed by that type. Each person of the $N_{ij}$ people of type $j$ consumes $\eta_{ijt}/N_{ij}$ of the household budget $x_{it}$, so they each have a shadow budget of $\eta_{ijt}x_{it}/N_{ij}$.
Each household member faces a shadow budget of $\eta_{ijt}x_{it}/N_{ij}$ and shadow prices of $A_{it}p$$_{t}$, so that, within the household, indirect utility is given by
Let $V_{xj}(p,w,z)\equiv\partial V_{j}(p,w,z)/\partial w$ be the monotonically decreasing marginal utility of person $j$ with respect to their shadow budget. Then, $V_{xj}\left(A_{it}p_{t},\eta_{ijt}x_{it}/N_{ij},z_{it}\right)$ is the value of their marginal utility evaluated at their shadow budget constraint.
Let $p_{2}^{s},z_{i2}^{s},\overline{x}_{i}^{s}$ for $s=1,2$ be the possible realizations of state-dependent variables that occur with household-specific probabilities $\pi_{i}^{1}$ and $\pi_{i}^{2}$ (which sum to $1$). Here, $\overline{x}_{i}^{s}$ is the state-specific lifetime wealth of household $i$, revealed in period $2$.
Pareto weights $\phi_{ij}$ vary arbitrarily across households, and depend, for example, on the household-specific expectation of lifetime wealth at the moment of household formation. Because ours is a full commitment model, Pareto weights do not have a time subscript because they are fixed at the moment of household formation, when the full-commitment contract is set. This household-level time-invariant variable will form the basis of our fixed-effects variation.
As in ChiapporiMazzocco2017,
the assumption of efficiency implies that we can represent the household's decisions via the Bergson-Samuelson Welfare Function, $W_{i}$, for the household is
The term in square brackets is the expected lifetime utility, discounted by the discount factor $\rho_{i}$, of each member of type $j$ in household $i$. Each member of type $j$ gets the Pareto weight $\phi_{ij}$. The assumption of efficiency implies that the household reaches the Pareto Frontier; the Pareto weights pinpoint which point on the Frontier is chosen by the household.
Next, substitute indirect utility ((ref)) for utility $v_{ij1}$ and $v_{ij2}^{s}$ into ((ref)),\footnote{In contrast, ChiapporiMazzocco2017 substitute direct utility for utility $v_{ij1}$ and $v_{ij2}^{s}$ using a model of pure private and pure public goods. In that model, each individual's utility is given by their direct utility function, which is a function of their (unobserved) consumption of a vector of private goods and their (observed) consumption of a vector of public goods. } and form the Lagrangian using the intertemporal budget constraint with interest rate $\tau$, $x_{i1}+x_{i2}^{s}/(1+\tau)=\overline{x}_{i}^{s}$, and the adding-up constraints on resource shares, $\sum_{j}\eta_{ij1}=\sum_{j}\eta_{ij2}^{s}=1.$ Each household $i$ chooses $x_{i1},$ $\eta_{ij1}$ and $\eta_{ij2}^{s}$ to maximize $W_{i}$.
Solving this optimization problem obtains for any two types, $j$ and $k$, that:
for $s=1,2$ for all $i$. That is, the household chooses resource shares so as to equate ratios of marginal utilities with ratios of Pareto weights.
An assignable good is one where we observe the consumption of that good by a specific person (or type of person). Assuming the existence of a scalar-value demand function $q_{j}(p,x,z)$ for an assignable and non-shareable good (e.g., food or clothing) for a person of type $j$, BrowningChiapporiLewbel13 show that the household's quantity demand, $Q_{ijt}$, for the assignable good for each of the $N_{ij}$ people of type $j$ is given by \[ Q_{ijt}=q_{j}(A_{it}p_{t},\eta_{ijt}x_{it}/N_{ij},z_{it}). \]
Assuming that the assignable good is a normal good implies that $q_{j}$ is strictly increasing in its second argument, and is therefore strictly monotonic.
Since $p_{t}$ is unobserved and varies over time,\footnote{The assumption that prices are unobserved but vary only with time is important for our empirical application below, where we observe Bangladeshi households in 2 time periods. It implies moment conditions for differenced demand functions that do not depend on prices. If prices vary with any observed variables, then these could be worked into the moment conditions, by conditioning on those variables. However, if prices vary with unobserved variables, our estimation strategy would not work.} we can express $Q_{ijt}$ as a time-varying function of observed data. Defining $\widetilde{q}_{jt}(x,z)=q_{j}(p_{t},x,z)$, we have
This is the structural demand equation that we ultimately bring to the data.
The model above determines resource share via the first-order conditions ((ref)), and these resource shares depend on time-invariant Pareto weights $\phi_{ij}$. But these resource shares are a vector of implicit functions, which may be hard to work with. To make the model tractable, we impose sufficient structure on utility functions to find closed forms for resource shares.
In our empirical example below, we work with data that have time-invariant demographic characteristics, so let $z_{it}=z_{i}$ be fixed over time. This implies that the shareability of goods embodied in $A$ is time-invariant: $A_{it}=A_{i}=A(z_{i})$. We will estimate the demand equation for women's food in nuclear households comprised of $1$ man, $1$ woman and $1-4$ children, so $N_{if}=N_{im}=1$.
Let indirect utilities be in the price-independent generalized logarithmic (PIGL) class (Muellbauer75,Muellbauer76) given by
Here, $V_{j}$ is homogeneous of degree $1$ in $p,x$ if $C_{j}$ is homogeneous of degree $0$ in $p$ and $B$ is homogeneous of degree $-1$ in $p$. $V$ is increasing in $x$ if $B(p,z)$ is positive and $V$ is concave in $x$ if $r(z)<1$. In terms of preferences, this class is reasonably wide. It gives quasihomothetic preferences if $r(z)=1$, and PIGLOG preferences as $r(z)\rightarrow0$ (this includes the Almost Ideal Demand System of Deatonm80).
The functions $C_{j}$ vary across types $j$, and so the model allows for preference heterogeneity between types, e.g., between men and women. The restrictions that $B(p,z)$ and $r(z)$ don't vary across $j$ and that $r(z)$ does not depend on prices $p$ are important: as we see below, they imply that resource shares are constant over time.
Substituting into the BCL model, observed demographics and period $t$ budgets, we have that
marginal utilities are given by \[ V_{xj}(A_{it}p_{t},\eta_{ijt}x_{it}/N_{ij},z_{it})=B(A_{i}p_{t},z_{i})^{r(z_{i})}\left(\eta_{ijt}x_{it}/N_{ij}\right)^{r(z_{i})-1}. \] For $r(z_{i})\neq1$, and for any pair of types $j,k\text{, }$we substitute into ((ref)) and
rearrange to get
The household chooses resource shares in each period and each state to satisfy ((ref)). Since the right-hand side has no variation over time or state, this implies that, given PIGL utilities ((ref)), the resource shares in a given household $i$ are independent of period $t$ and state $s$. However, resource shares do vary with both observed and unobserved variables across households $i$.
Let the fixed resource shares that solve the first-order conditions with PIGL demands be denoted $\overline{\eta}_{ij}$, and define
to equal to the logged resource share of the woman in the household.
Let there be a multiplicative berkson50 measurement error denoted $\exp\left(-U\right)$ which multiplies the budget, so that if we observe $x$, the actual budget is $x/\exp\left(U\right)$. The measurement error is i.i.d. across time and households, which implies stationarity of $U$. Here, the measurement error does not affect resource shares, but does affect the distribution of observed quantity demands.
Then household demand for women's food (the assignable good), $Q_{ift}$, is given by \[ Q_{ift}=\widetilde{q}_{ft}(\exp\left(\alpha_{i}\right)x_{it}/\exp\left(U_{it}\right),z_{i}). \] This is a FELT model, conditional on covariates:
where $h_{t}(Y_{it}^{*},z_{i})=\widetilde{q}_{ft}\left(\exp\left(Y_{it}^{*}\right),z_{i}\right)$, $Y_{it}^{*}=\alpha_{i}+X_{it}-U_{it}$, and $X_{it}=\ln x_{it}$ is the logged household budget.
The assumption that the assignable good is normal means that the time-varying functions $h_{t}$ are strictly monotonic in $Y_{it}^{*}$. One could additionally impose that the demand functions $\widetilde{q}_{ft}$ come from the application of Roy's Identity to the indirect utility function ((ref)). These demand functions equal a coefficient times the shadow budget plus a coefficient times the shadow budget raised to a power, where the coefficients are time-varying and depend on $z_{i}$.\footnote{Individual demands are derived by the application of Roy's Identity to ((ref)), and are: \[ q_{jt}(x,z)=c_{jt}\left(z\right)x^{1-r(z)}+b_{t}\left(z\right)x \] where $c_{jt}\left(z\right)=-\frac{\nabla_{p}C_{j}(p_{t},z)}{B(p_{t},z)^{r}}$, $b_{t}\left(z\right)=-\nabla_{p}\ln B(p_{t},z)$. This notation makes clear that we have time-varying demand functions, due to the fact that prices vary over time. In our application, prices in each period are not observed, so we allow the ($z-$dependent) functions $c_{jt}$ and $b_{t}$ to vary over time. We require that the assignable good be normal, meaning that its demand function is globally increasing in $x$. This form for demand functions is globally increasing if $c_{jt}(z)$, $b_{t}(z)$ and $1-r(z)$ are all positive. } We do not impose that additional structure here; instead, we show in Section (ref) that the estimated demand curves given by FELT are close to the PIGL shape restrictions.
Here, the time-dependence of $h_{t}$ is economically important; it is driven by the price-dependence of preferences and by the fact that prices are common to all households $i$ but vary over time $t$. Further, the fixed effects $\alpha_{i}$ are economically meaningful parameters: they are equal to the logged women's resource shares in each household. The standard deviation of the logs is a common inequality measure, and the standard deviation of $\alpha_{i}$ is identified by FELT given strict monotonicity, as we show in Section (ref). Further, the covariation of $\alpha_{i}$ with observed regressors is identified.
We use data from the $2012$ and $2015$ Bangladesh Integrated Household Surveys. This data set is a household survey panel conducted jointly by the International Food Policy Research Institute and the World Bank. In this survey, a detailed questionnaire was administered to a sample of rural Bangladeshi households. This data set has two useful features for our purposes: 1) it includes person-level data on food intakes and household-level data on total household expenditures; and 2) it is a panel, following roughly $6000$ households over two (nonconsecutive) years. The former allows us to use food as the assignable good to identify our collective household model parameters. The latter allows us to model household-level unobserved heterogeneity in women's resource shares.
The questionnaire was initially administered to $6503$ households in $2012$, drawn from a representative sample frame of all rural Bangladeshi households. Of these, $6436$ households remained in the sample in $2015$. In these data, expenditures on food include imputed expenditure from home production. We drop households with a discrepancy between people reported present in the household and the personal food consumption record, and households with no daily food diary data. Of the remaining data, $6205$ households have total expenditures reported for both $2012$ and $2015$.
In this paper, we focus on households that do not change members between periods.\footnote{That is, we exclude households with births, deaths, new members by marriage or adoption, etc. Although a full-commitment model can accommodate such changes in household composition, it is easier to think through the meaning of a person's resource share if the composition is held constant. } There are $1920$ households whose composition is unchanged between $2012$ and $2015$. Roughly half of these households have more than one adult man or more than one adult women. To simplify the interpretation of estimated resource shares we focus on nuclear households. This leaves $871$ nuclear households comprised of one man, one woman and $1$ to $4$ children, where children are defined to be $14$ years old or younger.
Our household-level annual expenditure, $x_{it},$ is the sum of total expenditure on, and imputed home-produced consumption of, the following categories of consumption: rent, food, clothing, footwear, bedding, non-rent housing expense, medical expenses, education, remittances, religious food and other offerings (jakat/ fitra/ daan/ sodka/ kurbani/ milad/ other), entertainment, fines and legal expenses, utensils, furniture, personal items, lights, fuel and lighting energy, personal care, cleaning, transport and telecommunication, use-value from assets, and other miscellaneous items. These spending levels derive from one-month and three-month duration recall data in the questionnaire, and are grossed up to the annual level. Estimation uses $X_{it}=\ln x_{it}$, the natural logarithm of annual consumption.
The assignable good, $Y_{it}$, is annual expenditure on food for the woman. The surveys contains a one-day (24-hour) food diary with data on person-level quantities (measured in kilograms) of food consumption in 7 categories: Cereals, Pulses, Oils; Vegetables; Fruits; Proteins; Drinks and Others. These consumption quantities include home-produced food and purchased food and gifts. They include both food consumed in the home (both cooked at home and prepared ready-to-eat food), as well as food consumed outside the home (at food carts or restaurants). These one-day food quantities are transformed into one-day food expenditures by multiplying by estimated village-level unit-values (following Deaton1997), and from this, the woman's share of one-day household food expenditure is calculated. Finally, the woman's annual expenditure on food is calculated as her share of one-day food expenditure multiplied by the household's total annual food expenditure (from recall data).
Our model is also conditioned on a set of time-invariant demographic variables $z_{i}$. We include several types of observed covariates in $z_{i}$ that may affect both preferences and resource shares: 1) the age in $2012$ of the adult male; 2) the age in $2012$ of the adult female; 3) the average age in $2012$ of the children; 4) the average education in years of the adult male; 5) the average education in years of the adult female; 6) an indicator that the household has 2 children; 7) an indicator that the household has 3 or 4 children; and 8) the fraction of children that are girls.\footnote{Since household membership is fixed for all households in our sample, age, number and gender composition are time-invariant by construction. However, education level of men and women are time-varying in roughly 20% of households. For our time-invariant education variables, we use the average education across the two observed years.}
For the first five of these demographic variables, in order to reduce the support of the regressors, we top- and bottom-code each variable so that values above (below) the $95^{th}$ ($5^{th}$) percentiles equal the $95^{th}$ ($5^{th}$) percentile values. For all seven of these variables, we standardize the location and scale so that their support is $[0,1]$. This support restriction simplifies our monotonicity restrictions when it comes to estimation, as explained in Section (ref).
We do not trim the data for outliers in the budget or food quantity demands. Instead, we trim the support of the estimated nonparametric regression functions to account for fact that these estimators are high-variance near their boundaries.
Table $2$ in the Appendix, Section (ref) gives summary statistics on these data.
Following our identification results, estimation could be based on composite versions of the maximum score estimator or the conditional logit estimator (see BotosaruMuris). Here, we instead follow a sieve GMM approach that facilitates the inclusion of a large vector of demographic conditioning variables $z$ and the imposition of strict monotonicity on the demand functions (aka: normality of the assignable good).
The women's food demand equation ((ref)) is a FELT model, conditional on observed covariates $z$. Denote the inverse demand functions $g_{t}(Y_{it},z_{it})=h_{t}^{-1}(\cdot,z_{it})$. Given ((ref)) a two-period setting with $t=1,2$, and time-invariant demographics $z_{it}=z_{i}$, we have
implying the conditional moment condition
using stationarity of the conditional distribution of the Berkson errors $\exp\left(-U_{it}\right)$.
We provide a detailed description of our GMM estimator in Appendix (ref). Briefly, we approximate the inverse demand functions, $g_{t}$, $t=1,2,\text{ }$ using 8th order Bernstein polynomials to impose monotonicity (estimates for other orders are reported in Appendix (ref)) and use the nonparametric bootstrap for inference. Because the nonparametric bootstrap may not be valid for this case, we take our estimated confidence intervals with caution, see Appendix (ref).
We characterize several interesting features of the distribution of resource shares. Recall from Theorems (ref) and (ref) that identification of features of this distribution does not impose a normalization on assignable good demand functions, and only identifies the distribution of logged resource shares (fixed effects) up to location. Consequently, we only identify features of the resource share distribution up to a scale normalization. This is related to identification results in cEkeland09 which show identification up to location using assignable goods.
Let $\widehat{g}_{it}=\widehat{g}_{t}(Y_{it},z_{i})$ equal the predicted values of the inverse demand functions at the observed data. Recall that $g_{t}\left(Y_{it},z_{i}\right)=\alpha_{i}+X_{it}-U_{it},$ so we can think of $\widehat{g}_{it}-X_{it}$ as a prediction of $\alpha_{i}-U_{it}$. We then compute the following summary statistics of interest, leaving the dependence of $\hat{g}_{it},\,t=1,2,$ on $z_{i}$ implicit:
Of these, the first 2 summary statistics are about the variance of fixed effects, and are computed using data from both years. Their validity requires serial independence of the measurement errors $U_{it}$. In contrast, the second 2 summary statistics are about the correlation of fixed effects with the household budget, and are computed at the year level. They are valid with stationary $U_{it}$, even in the presence of serial correlation.
We also consider the multivariate relationship between resource shares, household budgets and demographics. Recall that the fixed effect $\alpha_{i}$ subject to a location normalization; this means that resource shares are subject to a scale normalization. So, we construct an estimate of the woman's resource share in each household as $\widehat{\eta}_{i}=\exp\left(\frac{1}{2}\left(\hat{g}_{i1}-X_{i1}\right)+\left(\hat{g}_{i2}-X_{i2}\right)\right)$, normalized to have an average value of $0.33$. Then, we regress estimated resource shares $\widehat{\eta}_{i}$ on $\bar{X}_{i}$ and $Z_{i}$, and present the estimated regression coefficients, which may be directly compared with similar estimates in the cross-sectional literature.
The estimated coefficient on $X_{i}$ gives the conditional dependence of resource shares on household budgets, and therefore speaks to the reasonableness of the restriction that resource shares are independent of those budgets (an identifying restriction used in the cross-sectional literature). Finally, using the estimate of the variance of fixed effects, we construct an estimate of $R^{2}$ in the regression of resource shares on observed covariates. This provides an estimate of how much unobserved heterogeneity matters in the overall variation of resource shares.
Figure (ref) shows our estimates of $h_{1}$ and $h_{2}$ (or, equivalently, of $g_{1}$ and $g_{2}$) for $K=8$, for a family with two children with mean values, $\overline{z}$, of the other demographics. The figures have food quantities $q_{t}$ on the vertical axis and $\widehat{g}_{t}\left(q_{t}\right)$ on the horizontal axis, so the horizontal axis is like a predicted logged household budget. Solid lines give the nonparametric estimates, and 95% pointwise confidence bands for the nonparametric estimates are denoted by dotted lines. Additionally, to provide reassurance that the PIGL utility model---which implies the FELT demand curves---fits the data adequately, we display the PIGL demand curve closest to the FELT estimates in each time period with dashed lines.\footnote{We compute these PIGL demand curves by nonlinear least squares estimation of a pooled $q_{it}$ on $\widehat{g}_{it}$, where the demand curves have the form $q_{it}=c_{t}x^{1-r}+b_{t}x$. We estimate the model on a grid of 198 points, one for each interior percentile of $q_{it}$ in each period $t=1,2$. }
Note that since $g_{t}$ are identified only up to location (of $g_{1}$), we normalize the average of $\widehat{g}_{t}$ to half the geometric mean of household budgets at $t=$$1$, $\overline{x}_{1}$. Because estimated nonparametric regression functions can be ill-behaved near their boundaries, we truncate the estimated functions at the 5th and 95th percentiles of the distribution of $q_{t}$ in each $t$. The key message from Figure (ref) is that these estimated demand curves are somewhat nonlinear, estimated reasonably precisely, and not too far from PIGL. The estimated PIGL curvature parameter is $r(\overline{z})=0.06$, which means that food demands are close to PIGLOG (as in Banks1997).
Table (ref) gives our summary statistics (items 1-4 above), with bootstrapped 95% confidence intervals, for our estimates with 8 Bernstein polynomials (see the Appendix for other lengths of the Bernstein sieve). In the lower panel, we provide estimated regression coefficients, also with bootstrapped 95% confidence intervals, where we regress estimated resource shares $\widehat{\eta}_{i}$ on log-budgets $\overline{X}_{i}$ and demographics $z_{i}$.
Starting with the top panel of Table (ref), the standard deviation of $\alpha_{i}$ is a measure of inter-household dispersion in women's resource shares. If this dispersion is very small, then variation in resource shares does not induce much inequality, and we can reasonably use the household-level income distribution as a proxy for person-level inequality. However, if the dispersion is large, then household-level measures of inequality could be very misleading. Note here that we focus on inequality among women, not on gender inequality, precisely because the location of resource shares is not identified.
The estimated value is roughly $0.26$, with a 95% confidence interval covering roughly $0.15$ to $0.37$. To get a sense of the magnitude for the standard deviation of logged resource shares, suppose that women's resource shares were lognormally distributed. Then our estimated standard deviation of 0.26 is consistent with 95% of the distribution of the resource shares lying in the range $[0.25,0.75]$, which represents quite a bit of heterogeneity across households.
The next row of Table (ref) considers how much of the variation in $\alpha_{i}$ we can explain with observed covariates. The standard deviation of $e_{i}$ gives a measure of the unexplained variation, and gives us an idea of whether household-level unobserved heterogeneity is an important feature of the data. If the standard deviation of $e_{i}$ is very small, then fixed effects are not needed---conditioning on observed covariates would be sufficient. Our estimate of the standard deviation of the unexplained variation in $\alpha_{i}$ is about $0.16$. This is large relative to the overall estimated standard deviation of $0.26$, and suggests that accounting for household-level unobserved heterogeneity is quite important.
The next two rows give the covariance of $\alpha_{i}$ and $X_{it}$. Here, we see that log resource shares $\alpha_{i}$ strongly and statistically significantly negatively covary with observed household budgets (the implied correlation coefficients are close to $-0.8$). This means that women in poor households are somewhat less poor than they appear (on the basis of their household budget), and women in richer households are somewhat more poor than they appear. This is consistent with households that are closer to subsistence having a more equal distribution of resources.
The next two rows give the estimated standard deviation of women's log shadow budgets. This is a scale-free parameter: it does not depend on the location normalization of $\alpha_{i}$ (which corresponds to a scale normalization of shadow budgets). The estimated standard deviations are $0.35$ and $0.38$ in the two periods, respectively. We can compare these with the standard deviation of log-budgets, reported in Table 1, of $0.49$ and $0.53$. The point estimates suggest that there is less inequality in women's shadow budgets than in household budgets. Although the confidence intervals are large, the test of the hypothesis that the standard deviation of log-budgets equals the standard deviation of log-shadow budgets rejects in both years.\footnote{For $H_{0}:Var\left(X_{it}\right)-Var\left(\alpha_{i}+X_{it}\right)$, we have the following estimated test statistics and (confidence intervals). Period 1: $0.110\text{}(0.0236,0.165)$; Period 2: $0.137\text{}(0.0527,0.191)$.}
Thus, if we take these results at face value, there is less consumption inequality among women than household-level analysis would suggest. However, another implication of this is that there is more gender inequality than household level data would suggest. The reason is that household-level analysis of gender inequality pins gender inequality on over-representation of one gender in poorer households. In our data, all households have 1 man and 1 woman, so household-level analysis of gender inequality would show zero gender inequality. But, because women in richer households have smaller resource shares, this induces gender inequality even in these data.
Finding correlation between $\alpha_{i}$ and household budgets is not sufficient to invalidate previous identification strategies for cross-sectional settings that rely on independence between resource shares and household budgets. The reason is that the independence required is conditional on other observed covariates. To get a handle on this, the bottom panel of Table (ref) presents estimates of coefficients in a linear regression of normalized (to average $0.33$) estimated resource shares $\widehat{\eta}_{i}$ on log-budgets $\overline{X}_{i}$ and other covariates $z_{i}$.
The figure below shows the scatterplot of predicted resource shares versus the log household budget. Here, we see a lot of variation in resource shares, and it is clearly correlated with household budgets. The overall variation here provides an estimate of the explained sum of squares in an infeasible regression of true resource shares $\eta_{i}$ on $\overline{X}_{i}$ and $z_{i}$.\footnote{This artificial regression is infeasible because we observe (through $g_{it}$) a prediction of $\alpha_{i}+U_{it}$, not of $\alpha_{i}$ itself. However, because we have an estimate of the variance of $\alpha_{i}$, we can construct an estimate of the variance of $\eta_{i}$ (subject to the scale normalization that it has a mean of $0.33$). In our regression, the LHS variable is $\exp\hat{g}_{it}$, which is a prediction of $\eta_{i}u_{it}$. Since $u_{it}$ are uncorrelated with $\overline{X}_{i}$ and $z$ by assumption, the explained sum of squares from this regression applies to $\eta_{i}$, and we can use it to form an estimate of $R^{2}$, which we report, along with a bootstrapped confidence interval. } We may construct an estimate of the total sum of squares of resource shares from our estimate of the standard deviation of $\alpha_{i}$. This yields an estimate of $R^{2}$ in the infeasible regression, which we interpret as the fraction of variation in resource shares explained by observables.
In the first row of the bottom panel, we see that observed variables explain roughly half the variation in resource shares (the estimate of $R^{2}$ is $0.52$). This magnitude of explained variation is very close to that reported in DunbarLewbelPendakur2019 in their cross-sectional estimate based on Malawian data. However, whereas the estimate in DunbarLewbelPendakur2019 is conditional on the assumption that unobserved heterogeneity in resources shares is independent of the household budget, our estimate allows for correlation of resource shares with the household budget. This large magnitude of unexplained variation (roughly half) suggests that accounting for unobserved heterogeneity in resource shares is quite important.
Consider first the coefficient on $\overline{X}_{i}$. The estimated coefficient is $-0.045$ and is statistically significantly different from $0$. This means that, even after conditioning on other covariates (many of which are highly correlated with the budget), we still see a significant relationship between resource shares and household budgets.
However, the magnitude of this effect is small. Conditional on $z_{i}$, the standard deviation of $X_{it}$ is $0.43$ in year 1 and $0.47$ in year 2. Thus, comparing two households with identical $z$ but which are one standard deviation apart in terms the household budget, we would expect the woman in the poorer household to have a resource share $2$ percentage points higher than the woman in the richer household. Thus, the bulk of the variation that makes the standard deviation of women's shadow budgets smaller than that of household budgets is not running through the dependence of resource shares on household budgets, but rather through the dependence of resource shares on other covariates that are correlated with household budgets.
We get a very precise estimate of the conditional dependence of resource shares on household budgets. Overall, then, we see that women's resource shares are statistically significantly correlated with household budgets, even conditional on other observed characteristics. But, the estimated difference in resource shares at different household budgets is quite small. So, we take this as evidence that the identifying restrictions used by dlp13 (and DunbarLewbelPendakur2019) may be false, though perhaps not very false. It does suggest that alternative identifying restrictions---such as those developed here with a panel model---may be useful.
The rows of Table (ref) give several other coefficients that are comparable to other estimates in the literature. Calvi finds that women's resource shares in India decline with the age of the woman. In these Bangladeshi data, we find evidence that women's resource shares are strongly negatively correlated with the age of women and positively correlated with the age of men.
dlp13 find that women's resource shares in Malawi decline with the number of children. Here, we also see that pattern: households with 2 children have women's resource shares $5$ percentage points less than households with 1 child; households with 3 or 4 children have resource shares $12$ percentage points less. dlp13 also find that Malawian women's resource shares are higher in households with girls than households with boys. We do not see evidence of this in rural Bangladesh: the estimated coefficient on the fraction of children that are girls statistically insignificantly different from $0$.
In the Appendix, we also provide estimates analogous to Table 1 using a different assignable good: clothing. Under the model, using different assignable goods should yield the same estimates of resource shares.\footnote{Food is a plausible assignable good (because if one person eats it, nobody else can), but it may not be non-shareable (because there may be scale economies in cooking). In contrast, clothing may be plausibly non-shareable, but it may not be assignable (because, e.g., mothers and daughters might wear each others' clothes). See the Appendix for details on clothing estimates.} This is roughly what we find in our estimates using women's clothing.
Our estimates use 8th order Bernstein polynomials to approximate the inverse demand functions(ref). In Appendix B, we present estimates using Bernstein polynomials of order $K=1,4,8,10$ and show that our finite-dimensional parameter estimates have roughly the same value for $K\geq8$.
In summary, in these rural Bangladeshi households, we find evidence that women's resource shares have substantial dependence on household-level unobserved heterogeneity and are slightly negatively correlated with household budgets. The former suggests that random-effects type approaches to the estimation of resource shares may be inadequate. The latter suggests that consumption inequality faced by women is actually smaller than household-level consumption inequality. It also suggests that cross-sectional identification strategies invoking independence of resource shares from household budgets, such as dlp13, could be complemented by panel-based identification strategies such as ours.
\singlespacing