Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
80,967 characters · 12 sections · 56 citation commands
On a log-symmetric quantile tobit model applied to female labor supply data
{\bf Abstract.} The classic censored regression model (tobit model) has been widely used in the economic literature. This model assumes normality for the error distribution and is not recommended for cases where positive skewness is present. Moreover, in regression analysis, it is well-known that a quantile regression approach allows us to study the influences of the explanatory variables on the dependent variable considering different quantiles. Therefore, we propose in this paper a quantile tobit regression model based on quantile-based log-symmetric distributions. The proposed methodology allows us to model data with positive skewness (which is not suitable for the classic tobit model), and to study the influence of the quantiles of interest, in addition to accommodating heteroscedasticity. The model parameters are estimated using the maximum likelihood method and an elaborate Monte Carlo study is performed to evaluate the performance of the estimates. Finally, the proposed methodology is illustrated using two female labor supply data sets. The results show that the proposed log-symmetric quantile tobit model has a better fit than the classic tobit model.\\ {\bf Keywords:} Log-symmetric distributions; Quantile regression; Monte Carlo simulation; PSID, PNAD.\\
\onehalfspacing
In a sample, there is a possibility that the dependent variable cannot be observed in its entire domain or it is not observed for some of the individuals that incorporate it, namely, censoring is present in the data. This occurs when the information on the dependent variable is not fully available for some units of the sample. However, for these units, data on explanatory variables are known across their domain. Thus, in these cases, it is necessary to work with models that take censoring into account, such as the tobit model; see long:97.
The classic tobit model was introduced by t:58 and has been widely used in economics, in addition to several other areas; see, for example, amemiya:84, helsel:11 and bgls:17. This model is used to left-censored dependent variables and was motivated by a study of the relationship between family spending on durable goods and family income. In this study, a portion of the dependent variable (family expenditure) was zero, that is, censored at a fixed limit value. Other studies in which the dependent variable is censored at zero for some observations include fair:77,fair:78 for modeling the number of extramarital affairs, jarque:87 for modeling family spending in various groups of commodities, and melenbergsoest:96 for modeling holiday expenses, among others.
The classic tobit model can also be used to estimate labor supply equations. Fixing working hours as the dependent variable, there is the possibility that the hours worked take the value zero. This happens when one or more individuals do not work, that is, for those who do not offer hours of work; see details in moffitt:82. According to the labor supply model of heckmanmaCurdy:80, which was discussed in islam:07, the censored model is relevant in cases where the sample consists of individuals chosen at random and with hours reported as zero if the individual does not work. When it occurs, the techniques used for estimating linear models are inappropriate due to the presence of censoring.
One of the limitations of the classic tobit model t:58 is the assumption that the error term is normally distributed. Although the normal distribution is widely used in several areas, it may not be appropriate to model dependent variables where positive skewness and/or light- and heavy-tails are present. In this context, the class of log-symmetric distributions represent important tools to overcome such problems. The log-symmetric distributions are a generalization of the log-normal distribution and have as its special cases distributions that have lighter or heavier tails than those of the log-normal, as well as bimodal distributions; see, for example, j:08, vanegasp:15,vanegasp:16a,vanegasp:16b,vanegaspaula:17 and medeirosferrari:16. In addition to the log-normal distribution, other examples of log-symmetric distributions are the log-Student-$t$, log-power-exponential and extended Birnbaum-Saunders, among others.
Although important, the use of an appropriate distribution to describe the error distribution of a tobit model does not provide a more comprehensive picture of the effect of the explanatory variables on the dependent variable. In this sense, quantile regression plays an important role, being able to model conditional quantiles as a function of explanatory variables; details on quantile regression can be seen in hn:07 and dfv:14. In addition, quantile regression modeling is more efficient in cases where errors are not normally distributed or when the dependent variable has extreme value.
In this context, the primary objective of this paper is to propose a quantile tobit model based on log-symmetric distributions. The strategy for the proposal of the new tobit model uses a reparameterization of the log-symmetric distributions proposed by ssls:20, which has the quantile as one of its parameters. The secondary objectives are: (i) to obtain the maximum likelihood estimates of the model parameters; (ii) to carry out a Monte Carlo simulation to evaluate the performance of the maximum likelihood estimates; and (iii) to apply the proposed methodology to two real female labor supply data sets. The first application uses data from the Panel Study of Income Dynamics (PSID) which is an American longitudinal household survey, whereas the second application uses data extracted from the Brazilian National Household Sample Survey (PNAD) carried out in 2015. The PSID data were studied by bgls:17 using the classical tobit model, and the PNAD data were filtered by the authors from the original data available at the Brazilian Institute of Geography and Statistics (IBGE) website\footnote{The use of the PSID and PNAD databases on female labor supply is relevant since the participation rate of women in the labor market has been a transforming factor in the labor market. In the United States case, for example, jacobsen99 claims that the increase in female participation in the labor market is the most striking economic statistic in the 20th century. As for Brazil, scomenfilho01 emphasize that there was a strong increase in female participation rates, especially for women with 1 to 11 years of study. On the other hand, barrosjatobamendoca95 highlight that the study of the mechanisms and motivations that explain the increase in the rate of female participation in the labor market, in addition to the fact that this rate is a basic socio-economic indicator, have boosted attention to the area. However, even today, women are less likely to participate in the labor market than men, in addition, they are also more likely to be unemployed in most countries of the world, according to a recent study by the International Labour Organization iol2018. This justifies the use of the PSID and PNAD databases on female labor supply, given that the interest is in using censored data. In these data, part of the women do not work and for that reason they declare the income as being zero, which is classified as censored.}. The proposed approach can be seen as a generalization of the works of desousaetal:18 and slnb:20. In general, the results of the applications showed that the proposed log-symmetric quantile tobit model provides better fit as compared to the classic tobit model.
The advantages of the proposed log-symmetric quantile tobit model over the classic normal tobit model are: (a) greater flexibility in terms of distributional assumption, since the log-symmetric class incorporates several special cases, such as the log-normal, log-Student-$t$, log-power-exponential, and extended Birnbaum-Saunders, among others; (b) greater flexibility for data modeling, allowing to consider the effects of explanatory variables on the dependent variable along the spectrum of the dependent variable, due to the quantile approach; and (c) flexibility to accommodate heteroscedasticity, since the proposed tobit model allows the insertion of explanatory variables in the dispersion parameter. On the other hand, the quantile approach proposed in this work differs from the existing models in the literature, since we introduce a quantile tobit model based on a reparametrization of the error distribution. The existing quantile tobit models studied in the literature, for example, powel:86, buchinsky:98, jilinzhang:12, yuehong:12 and alhamzawiali:18, are based on the minimization approach introduced by koenker:78,koenker:05. Thus, the proposed approach is unprecedented in the literature for tobit models.
The rest of this paper proceeds as follows. In Section (ref), we briefly describe the class of log-symmetric distributions, both in its classical representation and in its reparameterization by quantile proposed by ssls:20. In Section (ref), we introduce the log-symmetric quantile tobit model, and then provide details on estimation, interpretation of parameter estimates and residual analysis. In Section (ref), we carry out a Monte Carlo simulation to assess the performance of the maximum likelihood estimates. In Section (ref), two applications to PSID and PNAD data are studied. Finally, in Section (ref), we discuss conclusions and some possible future research on this topic.
This section briefly describes the classical log-symmetric distributions vanegasp:16a and those based on the quantile proposed by ssls:20. The log-symmetric distributions reparameterized by the quantile, that is, those that have the quantile as one of their parameters, will be used to propose the log-symmetric quantile tobit model.
A random variable $T$ follows a log-symmetric distribution with scale parameter $\lambda>0$ and power parameter $\phi>0$, if its probability density function and cumulative distribution function are given by
and
respectively, where $G(\omega)=\eta{\int^{\omega}_{-\infty} g(z^2) \,\textrm{d}z }$, $\omega\in\mathbb{R}$, with $\eta$ being a normalizing constant and $g(\cdot)$ a density generator. In this case, the notation $T\sim\textrm{LS}(\lambda,\phi,g)$ is used. The $100q$-th quantile of $T\sim\textrm{LS}(\lambda,\phi,g)$ is given by
where $G^{-1}$ is the inverse of $G$ given in (ref). Table (ref) presents some density generators $g$ for some log-symmetric distributions; see details in vanegasp:16a. Note that the generator $g$ may involve and extra parameter $\xi$.
Consider a fixed number $q \in (0,1)$ and $Q$ the $100q$-th quantile of $T\sim\textrm{LS}(\lambda,\phi,g)$ given in (ref). Then, considering the one-to-one transformation $(\lambda,\phi) \mapsto (Q,\phi)$, ssls:20 proposed a reparameterization of the classical log-symmetric distribution, where the probability density function and the cumulative distribution function are given respectively by
and
implying the notation $T\sim\textrm{QLS}(Q,\phi,g)$. If $T\sim\textrm{QLS}(Q,\phi,g)$, ssls:20 have shown that the following properties hold: (a) $cT\sim\textrm{QLS}(cQ,\phi,g)$, with $c>0$; (b) $T^c\sim\textrm{QLS}(Q^c,c^2\phi,g)$, with $c>0$. We then readily have the following relation:
Let $T_{i}$ be a positive censored variable to the left at point $\Psi$, namely, it is observable for values greater than $\Psi$ and censored for values less than or equal to $\Psi$. Based on (ref), the log-symmetric quantile tobit model can be formulated as
where $\epsilon_i \sim \textrm{QLS}(1, 1, g)$, $Q_i=\exp(\bm{x}_i^\top\bm \beta)$ and $\phi_i =\exp(\bm{w}^{\top}_{i}\bm{\kappa})$, with $\bm{\beta} =(\beta_0,\ldots,\beta_{k})^\top$ and $\bm{\kappa}=(\kappa_0,\ldots,{\kappa_{l}})^\top$ denoting vectors of regression coefficients, and ${\bm{x}}^{\top}_{i}= (1,x_{i1},\ldots, x_{ik})^\top$ and ${\bm{w}}^{\top}_{i} = (1,w_{i1}, \ldots, w_{il})^\top$ denoting vectors of explanatory variables fixed and known associated with $Q_i$ and $\phi_i$, respectively.
The estimation of the parameters of the quantile log-symmetric tobit model presented in (ref) can be done by the maximum likelihood method. Let ${\bm T}=(T_1,\ldots,T_m,T_{m+1},\ldots,T_n)^{\top}$ be a sample of size $n$ from the quantile log-symmetric tobit model that contains $m$ left-censored data at $\Psi$ and $n-m$ uncensored data. Then, the corresponding log-likelihood function for the parameter vector ${\bm\theta} = ({\bm \beta}^{\top},{\bm \kappa}^{\top})^{\top}$ is given by
where $Q_{i}=\exp(\bm{x}_i^\top\bm \beta)$, $\phi_i=\exp(\bm{w}^{\top}_{i}\bm{\kappa})$, $G$ is as defined in (ref), and $g$ is given in Table (ref). By taking the logarithm of (ref), we obtain the log-likelihood function $\ell({\bm\theta})$, that is,
where $$ \ell_i({\bm\theta}) =
$$ By taking the first derivative of $\ell({\bm\theta})$ with respect to ${\bm \beta}$ and ${\bm \kappa}$, we obtain the score vector, that is,
where $ \dot{\bm \ell}_i({\bm\theta}) = (\dot{\bm \ell}_{i{\bm \beta}}^\top({\bm\theta}),\dot{\bm\ell}_{i{\bm \kappa}}({\bm\theta}))^\top$, with $$ \dot{\bm\ell}_{i{\bm\beta}}({\bm\theta}) =
$$ $$ \dot{\bm\ell}_{i{\bm\kappa}}({\bm\theta}) =
$$ where $\Pi(\xi^c_i)=({{\rm d}G(u)/{\rm d}u|_{u=\xi^c_i}})/{G(\xi^c_i)} $ and $\Delta(\xi_i^2)=({{\rm d}g(u)/{\rm d}u|_{u=\xi_{i}^2}})/{g(\xi_{i}^{2})} $, with \\ $\xi^c_i=({\log(\Psi)-\log(Q_i)+\sqrt{\phi_i}z_{p}})/{\sqrt{\phi_i}}$, $\xi_i=({\log(t_i)-\log(Q_i)+\sqrt{\phi_i}z_{p}})/{\sqrt{\phi_i}}$, $\gamma_i^c=\log(\Psi)-\log(Q_i)$ and $\gamma_i=\log(t_i)-\log(Q_i)$.
The maximum likelihood estimate for $\bm{\theta}$ is obtained by maximizing the log-likelihood function (ref) by equating the score vector $\dot{\bm\ell}(\bm{\theta})$, which contains the vector of first derivatives of $\ell({\bm\theta})$, to zero, providing the likelihood equations. In this case, as there is no analytical solution, they are solved by using the Broyden-Fletcher-Goldfarb-Shanno (BFGS) iterative method for non-linear optimization. Note that in (ref), the extra parameter $\xi$ is assumed to be fixed. The reason for this lies in the works of l:97 and kanoetal93. In the first work it is shown that the robustness of the Student-$t$ distribution to outlying observations holds only when the degree of freedom parameter is fixed, instead of being estimated directly in the log-likelihood function. In the second work, the authors report difficulties in estimating the extra parameter for the power-exponential distribution. Therefore, the parameter $\xi$ is estimated using the profiled log-likelihood. Two basic steps are required:
Under regularity conditions, the asymptotic distribution of $\widehat{\bm{\theta}}$ is a multivariate normal, that is, $$\sqrt{n}(\widehat{{\bm\theta}}-{\bm\theta})\dot{\sim}\textrm{N}_{k+l+2}\left(\bm{0}_{k+l+2}, {\bm{\Sigma}}_{{\bm{\theta}}}\right),$$ where $\,\dot{\sim}\,$ denotes convergence in distribution and ${\bm{\Sigma}}_{{\bm{\theta}}}$ is the asymptotic variance-covariance matrix of $\widehat{\bm{\theta}}$ ch:74, which is the inverse of the expected Fisher information matrix. We can approximate the expected Fisher information matrix by its observed version obtained from the Hessian matrix $\ddot{\bm \ell}({\bm\theta})$, which contains the second derivatives of $\ell({\bm\theta})$. Thus, ${{\bm{\Sigma}}}_{{\bm{\theta}}}\approx [-\ddot{\bm \ell}({\bm\theta})]^{-1}$, where
The elements of $\ddot{\bm \ell}({\bm\theta})$ are given by $$ \ddot{\bm\ell}_{i{\bm\beta}{\bm\beta}}({\bm\theta}) =
$$ $$ \ddot{\bm\ell}_{i{\bm\beta}{\bm\kappa}}({\bm\theta}) = \ddot{\bm\ell}_{i{\bm\kappa}{\bm\beta}}({\bm\theta})=
$$ $$ \dot{\bm\ell}_{i{{\bm\kappa}}{{\bm\kappa}}}({\bm\theta}) =
$$
The corresponding standard errors can then be approximated by the square roots of the diagonal elements in the variance-covariance matrix evaluated at $\widehat{\bm \theta}$.
The regression coefficient of the proposed tobit model is interpreted in terms of the effect on the latent variable $T_i^*$ in the uncensored part. Let $\beta_j$ be the $j$-th regression coefficient and use the subscript $(j)$ to imply excluding the $j$-th element, such that $\bm{x}_{i(j)}$ and ${\bm\beta}_{(j)}$ are, respectively, the vector of explanatory variables excluding $x_{ij}$ and the regression coefficients excluding $\beta_j$. Note that the quantile of $T_i^*$ is given by
If $x_{ij}$ increases by 1 while keeping ${\bm x}_{i(j)}$ fixed, we obtain
Thus, for any $j$ increasing $x_{ij}$ by 1, the quantile of $T_i^*$ will be multiplied by $\exp(\beta_j)$. This is usually expressed as a percentage change, and
is the approximate percentage increase (or decrease if the value of $\beta_j$ is negative) in the quantile of $T_i^*$ when ${x}_{ij}$ is increased by 1. For $x_{ij}$ dichotomous, $(\exp(\beta_j)-1)\times 100\%$ is the percentage increase (or decrease if $\beta_j$ is negative) in the quantile of $T_i^*$ when ${x}_{ij}$ changes from 0 to 1. Note that when $-0.4\leq \beta_j \leq 0.4$, we can use the approximation $(\exp(\beta_j)-1)\approx \beta_j$; see weisberg:14.
Goodness of fit and departures from the assumptions of the model can be assessed through residual analysis. In this work, we work with martingale-type (MT) residual, which is given by $$ {r}_{_{MT_{i}}}=\textrm{sign}(r_{_{\textrm{\tiny M}_{i}}})\sqrt{-2\left( r_{_{\textrm{\tiny M}_{i}}}+\rho_i\log(\rho_i- r_{_{\textrm{\tiny M}_{i}}})\right)},\quad i=1,\dots,n. $$ where $r_{_{\textrm{\tiny M}_{i}}} = \rho_i + \log(\widehat S(t_i))$, with $\widehat S(t_i)$ being the survival function fitted to the data, and $\rho_i=0$ or $1$ indicating that case $i$ is censored or not, respectively; details on the MT residual can be seen in tgf:90. Simulations results indicate that the empirical distribution of the MT residual is in agreement with the standard normal distribution; see SILVA20094482. Then, a normal quantile-quantile (QQ) plot with simulated envelope can be constructed for the MT residual to verify whether the model is correctly specified.
In this section, the performance of the maximum likelihood estimates of the log-symmetric quantile tobit models are evaluated by means of a Monte Carlo simulation study. The following distributions are considered: log-normal (log-NO), log-Student-$t$ (log-$t$), log-power-exponential (log-PE) and extended Birnbaum-Saunders (EBS). For each Monte Carlo replica, a simulated sample of the log-symmetric quantile tobit model is generated for fixed parameter values, then the maximum likelihood estimates are obtained for each simulated sample. Estimates of bias and mean squared error (MSE) are thus computed from the Monte Carlo replicas as $$\widehat{\textrm{Bias}}(\widehat{\theta}) = \frac{1}{\text{NREP}} \sum_{i = 1}^{\text{NREP}} \widehat{\theta}^{(i)} - \theta \quad \text{and}\quad \widehat{\mathrm{MSE}}(\widehat{\theta}) = \frac{1}{\text{NREP}} \sum_{i = 1}^{\text{NREP}} (\widehat{\theta}^{(i)} - \theta)^2,$$ where $\theta$ and $\widehat{\theta}^{(i)}$ are the true parameter value and its respective $i$-th maximum likelihood estimate, and $\text{NREP}$ is the number of Monte Carlo replicas. The R software has been used to do all numerical calculations; see r2020vienna.
Simulated data from the log-symmetric quantile tobit models are generated according to
where $\epsilon_i \sim \textrm{QLS}(1, 1, g)$, $Q_i=\exp(\beta_0 + \beta_1 x_i)$ and $\phi_i=\exp(\kappa_0 + \kappa_1 w_i)$.
The simulation scenario considers: $\beta_0 = 1.0$, $\beta_1 = 0.5$, $\kappa_0 = 1$, $\kappa_1 = 1.5$, $q = 0.10, 0.50, 0.90$ and $\text{NREP}=5,000$. The explanatory variables $x_i$ and $w_i$ are generated from the Bernoulli(0.5) and Uniform(0,1) distributions, respectively. The values of the extra parameters for the distributions considered are: $\xi=4$ (log-$t$), $\xi=0.3$ (log-PE) and $\xi=0.5$ (EBS). The value of $\Psi$ in (ref) is determined so that the censoring proportion is 10% or 40%.
Tables (ref)-(ref) present the results of Monte Carlo simulations based on the log-NO, log-$t$, log-PE and EBS distributions. These tables report the bias and MSE obtained for different combinations of censoring proportion, $q$ and sample size ($n$). A look at the results in Tables (ref)-(ref) allows us to conclude that, in general, as the sample size increases, the bias and MSE both decrease, as expected. This results is expected since the maximum likelihood estimator is consistent Greene2003Econometric. Moreover, we observe that, in general, when the censoring proportion increases, the bias and MSE both increase, that is, the performances of the estimates deteriorates, a result also expected since the likelihood function loses information contained in the sample when the percentage of censoring increases W-1992.
In this section, the proposed log-symmetric quantile tobit models are used to analyze the PSID and PNAD data. The PSID data set has already been analyzed in the literature on tobit models, whereas the PNAD data set was filtered by the authors from the original data available at the IBGE website\footnote{https://www.ibge.gov.br/.}. The distributions considered are the same used in the Monte Carlo simulation.
This application considers a data set corresponding to the PSID of 1976, based on data from the previous year, 1975; see m:87. This set contains $n=753$ observations of white married women between 30 and 60 years of age in 1975 (year of the interview was in 1976). Of these 753 women, 325 have a salary equal to zero, that is, censored at zero. Since the proposed models are for positive data, the dependent variable is considered to be $ T + 1 $, such that $\Psi=1$.
The objective here is to study, using log-symmetric quantile tobit models, the labor supply of white married women. The dependent variable is the hourly wage ($T$) (in 1975 US dollars) and the explanatory variables are: $age$, age in years; $educ$, education in years; $chil6$, number of children under 6 in the household; $chil618$, number of children between 6 and 18 years old in the household; and $exper$, years of previous experience in the labor market. These data were previously studied by bgls:17 using the classic normal tobit model and the Student-$t$ tobit model.
Descriptive statistics of the observed women's hourly wages reveal that the mean is equal to 2.375 and the median is equal to 1.625, that is, the mean is greater than the median. Moreover, the coefficient of variation (dispersion around the mean) is 136.52%, indicating a high dispersion of data around the mean. Finally, the coefficients of skewness and kurtosis are equal to 2.772 and 12.755, respectively. The result of the coefficient of skewness indicates the presence of positive skewness and the coefficient of kurtosis indicates the presence of heavy tails. The asymmetric nature of the data is confirmed by the histogram shown in Figure (ref)(a). Such results make the proposed log-symmetric quantile tobit models good candidates, since they are models for asymmetric data, with or without heavy tails. On the other hand, the classic normal tobit model and the Student- $t$ tobit model studied by bgls:17 are not suitable when positive skewness is present.
The proposed models can accommodate heteroscedasticity, then two versions are considered:
where $$Q_i=(\beta_0 + \beta_1\, age_i + \beta_2\, educ_i + \beta_3\, chil6_i +\beta_4\, chil618_i + \beta_5\, exper_i),$$ $\epsilon_i \sim \textrm{QLS}(1, 1, g)$, and
Note that the difference between the two models lies in the insertion of explanatory variables in the dispersion parameter $\phi$. Thus, Model 2 considers the presence of heteroscedasticity.
Table (ref) reports the AIC and BIC values for the adjusted log-symmetric quantile tobit models with $q=\{0.05,0.25,0.50,0.75,0.95\}$; similar results are obtained considering $ q = \{0.01, 0.02, \ldots, 0.99\}$. From this table, we observe that the models with explanatory variables in the dispersion parameter ($\phi$) provide better adjustments compared to those without explanatory variables, based on the values of AIC and BIC. In general, the results of Table (ref) indicate that the log-PE quantile tobit model presents the best fit, a result that is corroborated by the QQ plot with simulated envelope of the MT residual for this model with $q = 0.50$ (similar adjustments are found for other values of $q$), shown in Figure (ref)(b).
We can compare the proposed models to the classic normal tobit model and the Student-$t$ tobit model, which were fitted by bgls:17 using the same data set. The authors obtained $\text{AIC}=2893.85$ and $\text{BIC}=2926.22$ for the normal case, and $\text{AIC}=2760.99$ e $\text{BIC}=2793.36$ for the Student-$t$ case. Thus, we highlight the superiority of the proposed models when compared to the normal classic and Student-$t$ models adjusted by bgls:17 under three basic aspects: (i) the use of asymmetric distributions (log-symmetric distributions) that is more adequate for the PSID data. The normal classic and Student-$t$ tobit models use symmetric distributions; (ii) the possibility of modeling the dispersion and the consequent accommodation of heteroscedasticity, which improves the fit; and (iii) the modeling in terms of quantiles, which provides a richer characterization of the effects of the explanatory variables on the dependent variable. The normal classic tobit and Student-$t$ models do not consider the quantile approach.
Table (ref) reports the maximum likelihood estimates and standard errors for the log-PE quantile tobit model parameters based on Model 2, considering $ q =\{0.05, 0.25, 0.50, 0.75, 0.95\}$. This table also presents the estimation results based on the optimal quantile, denoted by $q_{otm}$, which was chosen through a profile approach, that is, for a grid of values of $ q=\{0.01, 0.02, \ldots, 0.99 \} $, we estimated the model parameters and computed the corresponding AIC and BIC values. Then, the value of $q_{otm}$ was the one which had the lowest AIC and BIC values. From Table (ref), we note that the maximum likelihood estimates of the model parameters change according to the value of $q$, that is, the magnitude of the effect of the explanatory variables varies with $q$. We can interpret the estimated coefficients in terms of the effect on the latent variable $T_i^*$, that is, in terms of the effect on the observed part of the hourly wage; see Subsection (ref). We observe, for example, that an increase of 1 year in the experience ($exper$), increases in $(\exp(0.0974)-1)*100\%=10.23\%$ the $5^{\circ}$ percentile ($q=0.05$) of the hourly wage, while increasing by $(\exp(0.0383)-1)*100\%=3.90\%$ the $75^{\circ}$ percentile ($q=0.75$) of hourly wages. In other words, the effect of increased experience on the observed part of the hourly wage is greater for women with lower income (lower quantiles). From Table (ref) we also observe that the parameter estimates associated with the explanatory variables $age$, $chil6$, $exper$, which model the dispersion, are significant, indicating the presence of heteroscedasticity in the data, justifying the dispersion modeling.
In this subsection, the log-symmetric quantile tobit models are illustrated using data from the PNAD for the year 2015, from the IBGE, which reports demographic and socioeconomic characteristics of the Brazilian population annually. A sample of the PNAD will be used, composed only of women\footnote{The most recent Continuous PNAD data will not be used, as in its dictionary there is a lack of objectivity in some variables of interest for this work, such as the experience/skill variable, which are the years of work in the main activity, as well as the variable marital status, which is also not clearly defined in the Continuous PNAD. It is of interest that the variables of the two samples, PSID and PNAD, be similar. Thus, the 2015 PNAD data will be used only to illustrate the proposed methodology.}. The PNAD data used in this work consists of a sample composed of women aged between 18 and 65 years old, with information on hourly wages and socio-economic characteristics. In total, the sample contains 26,460 observations, of which 387 are censored with a salary equal to zero. The data covers ten metropolitan regions in Brazil (Belém-PA, Fortaleza-CE, Recife-PE, Salvador-BA, Belo Horizonte- MG, Rio de Janeiro-RJ, Curitiba-PR, Porto Alegre-RS, Brasília-DF and São Paulo-SP). Nominal income values were deflated by the National Consumer Price Index (INPC) provided by the IBGE.
The objective is to study women's labor supply. The dependent variable is women's hourly wages ($T$) and the explanatory variables are: $age$, woman's age; $age^2$, woman's age squared; $color$, dummy for color with value 1 (white) or 0 (non-white); $civil$, dummy for marital status with value 1 (married) or 0 (non-married); $minor$, dummy for children under 10 in the household with a value of 1 (yes) or 0 (no); $educ$, captures educational returns on income and is classified by formal years of study ranging from 0 to 16 years; $exper$, captures the returns of the number of years in the main job on income, which is classified by years of work and can vary between 0 to 56 years; $head$, is a dummy that captures the condition of the woman in the household, presenting a value of 1 if the woman in the household is head of the family and 0 otherwise (non-head). Similarly to the first application, the dependent variable, woman's hourly wage, is added to one ($T+1$), such that $\Psi=1$.
The choice of the above-described explanatory variables is due to their importance in the female labor supply literature, in addition to the interest in similarity with the PSID data. The variable $educ$, for example, directly affects female participation rates in the labor market. On the other hand, the presence of women in the productive world does not depend only on market demand, there are other factors that can limit this participation, such as the presence or absence of minors in the household.
Descriptive statistics for women's hourly wages ($T$) indicate that the mean and median are 22.133 and 7.59, respectively. The coefficient of variation is 493.11%, indicating a high dispersion of the data around the mean, whereas the coefficients of skewness and kurtosis are equal to 20.225 and 573.536, respectively. The result of the coefficient of skewness shows the presence of a high positive skewness and the coefficient of kurtosis indicates the presence of heavy tails, confirming the hypothesis of using log-symmetric distributions is plausible. The asymmetric nature of the data is confirmed by the histogram shown in Figure (ref)(a).
Similarly to the first application, two versions of the log-symmetric quantile tobit model are considered:
where $$Q_i=\exp\left( \beta_0 + \beta_1\, age_i + \beta_2\, age^2_i + \beta_3\, color_i + \beta_4\, civil_i + \beta_5\, minor_i\right.$$ $$\left. + \beta_6\, educ_i + \beta_7\, exper_i + \beta_8\, head_i\right),$$ $\epsilon_i \sim \textrm{QLS}(1, 1, g)$, where
In Model 2, explanatory variables are present in the dispersion parameter $\phi$, that is, the presence of heteroscedasticity.
The AIC and BIC values for the adjusted log-symmetric quantile tobit models are presented in Table (ref). In the case of the normal classic tobit model, the values are $\text{AIC}=313909.3$ and $\text{BIC}=313991.1$, which shows the superiority of the proposed models in terms of adjustment. In Table (ref), the reported values of $ q $ are $ 0.05, 0.25, 0.50, 0.75 $ and $ 0.95 $, however, similar results are obtained considering $ q = \{0.01, 0.02, \ldots, 0.99\}$. It is noted that the models with explanatory variables in the dispersion parameter ($\phi$) present better adjustments in all the distributions considered. In general, the log-$t$ quantile tobit model presents the best fit, a result confirmed by the QQ plot with simulated envelope of the MT residual for this model with $ q = 0.50 $; see Figure (ref)(b). It is emphasized that similar plots are obtained for other values of $q$.
The model parameter estimates for the log-$t$ tobit quantile model based on Model 2 considering $ q =\{ 0.05, 0.25, 0.50, 0.75, 0.95\} $ and $ q_{otm}$, are presented in Table (ref). Note that the maximum likelihood estimates of the model parameters change according to the value of $q$, that is, the magnitude of the effect of the explanatory variables varies with $q$. Again, we can interpret the estimated coefficients in terms of the effect on the latent variable $T_i^*$ (observed part of the hourly wage). It is observed, for example, that for white women there is an increase in the $5^{\circ}$ percentile ($q=0.05$) of the hourly wage of $(\exp(0.0219)-1)*100\%=2.21\%$ when compared to non-whites. However, the increase in the $95^{\circ}$ percentile ($q=0.95$) of the hourly wage is of $(\exp(0.0219)-1)*100\%=26.05\%$. That is, the effect of $color$ on the observed part of the hourly wage is greater for women with higher income (larger quantiles). It is also observed in Table (ref) that the parameters associated with the explanatory variables $age$, $age^2$, $color$, $educ$, $exper$ and $head$ that model the dispersion, are significant, indicating the presence of heteroscedasticity in the data.
In this work, a class of quantile tobit models was proposed based on a reparameterization of the log-symmetric distributions. In such reparameterization, the quantile is one of the distribution parameters. The advantages of the proposed models over the classic tobit model include:
A Monte Carlo simulation study was carried out to evaluate the performance of the maximum likelihood estimates. In general, the results have shown good performances of the maximum likelihood estimates in terms of bias and mean squared error. Two applications to real data from PSID and PNAD were carried out to illustrate the proposed methodology. The applications have favored the use of the proposed log-symmetric quantile tobit models over the classic tobit model, illustrating the advantages (i), (ii) and (iii). As the final product of this paper, the authors are preparing an R package r2020vienna, which can be an important tool for many professionals, researchers in the field of economics and statistics, data scientists, among others.
For future research, the following lines can be explored:
Work on these problems is currently in progress and we hope to report these findings in future papers.