EconBase
← Back to paper

Marijuana on Main Streets? The Story Continues in Colombia: An Endogenous Three-part Model

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

134,656 characters · 17 sections · 62 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Marijuana on Main Streets? The Story Continues in Colombia An Endogenous Three-part Model

abstractCannabis is the most common illicit drug, and understanding its demand is relevant to analyze the potential implications of its legalization. This paper proposes an endogenous three-part model taking into account incidental truncation and access restrictions to study demand for marijuana in Colombia, and analyze the potential effects of its legalization. Our application suggests that modeling simultaneously access, intensive and extensive margin is relevant, and that selection into access is important for the intensive margin. We find that younger men that have consumed alcohol and cigarettes, living in a neighborhood with drug suppliers, and friends that consume marijuana face higher probability of having access and using this drug. In addition, we find that marijuana is an inelastic good (-0.45 elasticity). Our results are robust to different specifications and definitions. If marijuana were legalized, younger individuals with a medium or low risk perception about marijuana use would increase the probability of use in 3.8 percentage points, from 13.6% to 17.4%. Overall, legalization would increase the probability of consumption in 0.7 p.p. (2.3% to 3.0%). Different price settings suggest that annual tax revenues fluctuate between USD 11.0 million and USD 54.2 million, a potential benchmark is USD 32 million.

\Keywords{Marijuana demand, Marijuana legalization, Three-part model, Truncation.}

Introduction

Understanding the demand for illicit drugs is relevant, as it is a necessary analysis to comprehend the potential effects of their legalization. Cannabis remains the most commonly used illicit drug worldwide, with an estimated 209 million consumers in 2020, reflecting a 23% increase over the last decade. This is followed by opioids (61 million), amphetamines (34 million), cocaine (21 million), and ecstasy (20 million) UNODC2022. For decades, several countries have entered the debate on the legalization of marijuana. In this paper, we provide a new methodology for examining the consequences of legalizing marijuana, which helps to answer questions regarding its access, and extensive and intensive margins. Therefore, we can have a better understanding on: how much the prevalence and intensity of marijuana use rise under legalization?, is there heterogeneity in response to legalization among different age groups?, and could government policies based on taxation and campaigns about risk perception be effective in curbing use?

We extend jacobi2016marijuana's proposal, who proposed two independent two-part models to study the extensive and intensive margins of marijuana demand taken into account access restrictions when modeling demand for illicit drugs in Australia. In particular, we propose an endogenous three-part model that simultaneously takes into account access, extensive and intensive margins, as well as incidental truncation. To the best of our knowledge, this is the first paper to estimate demand for an illicit drug that considers simultaneously these features. Taking into account access restrictions is relevant as illicit drugs are not as easy to find, thus non-users have little information about how to get marijuana, which is a necessary condition to becoming a user jacobi2016marijuana. Truncation is also relevant when modeling demand for illicit drugs as individuals tend to withhold information about their use lloyd2013stigmatization. We apply our methodological proposal to model the demand for marijuana in Colombia, and study the potential implications of the legalization of marijuana in this market, where marijuana is decriminalized, making different counterfactual exercises trying to respond relevant inquiries regarding marijuana legalization. Colombian market is significant as this is one of the world's main producers of marijuana UNODC2021, resulting in relatively easy access and low prices.

For decades, countries have faced pressures regarding the decision to decriminalize or legalize the marijuana market. It is important to note that countries can implement a policy of liberalizing the marijuana market through either full legalization or decriminalization. According to nkansah2016gateway, legalization occurs when authorities approve the use of a substance previously prohibited by law, thereby eliminating the risk of arrest or fines. Decriminalization suspends criminal sanctions for using or possessing a particular substance; however, it maintains the substance's illegal status, allowing for potential punishment through civil fines, education, social work, and other measures. Uruguay was the first country to legalize recreational marijuana in 2013; some US states have also done so, as well as Canada in 2018 jacobi2016marijuana.

The literature highlights that the use of illicit drugs entails high costs for society, such as pressure on health systems, productivity loss van2006cannabis, delinquency, violence norstrom2014cannabis, incarceration and costs of the criminal justice system, educational performance, child abuse, and corruption, among others maccoun1996assessing, wen2014effect. There is also evidence that fewer young people consider cannabis consumption dangerous, which leads to the normalization of consumption behavior jarvinen2011normalisation. Additionally, there is the hypothesis that the legalization or decriminalization of marijuana serves as an incentive for consumption for younger populations since the early onset of cannabis use substantially increases subsequent consumption rates pudney2004keeping. This scenario implies that it is not easy for countries to decide to liberalize the market. Studies that show post-legalization effects, such as those of Rubin-Kahana2022 and roffman2016, suggest that some negative repercussions may not manifest until 5 or 10 years later, and the same applies to positive impacts.

Those in favor of marijuana legalization argue for a reduction in violence generated by the illegal market donohue2010rethinking, as well as a decrease in the cost of law enforcement (police and judiciary), a reduction in arrests roffman2016. Mace2020, Irvine2020, jacobi2016marijuana, Caputo1994, also speak about tax revenue that could benefit from legalization for the United States, Canada, and Australia, with funds allocated to education, sports, and other addiction programs. On the issue of taxation, miron2005budgetary also provides insights. Those who oppose legalization refer to the negative aspects of legalization, such as the possibility of a price decrease due to the elimination of transaction costs associated with illegality becker2006market. Another relevant aspect is the potency of cannabis and its psychoactive component, which can have increasingly harmful consequences for health, including mental health problems elsohly2016changes, van2006cannabis.

Thus, there is evidence of increased hospitalizations and intoxications due to excessive marijuana consumption post-legalization roffman2016, Rubin-Kahana2022, as well as psychotic and mental health effects Moran2022. Studies show an increase in consumption due to legalization in the United States. The most recent ones, such as that of Mennis2023, show that after the legalization of recreational marijuana, there was an increase in consumption, especially among teenagers. Additionally, researchers have found an increased frequency of marijuana use and a higher prevalence of symptoms associated with marijuana use disorder Kilmer2022. Barker2021 found that legalization had more significant direct effects on those who were already marijuana consumers, surrounded by an increasingly favorable climate for consumption.

The Canadian case Rotermann2020 shows similar results, where the increase in consumption is mainly associated with men over 25 years old. However, it is essential to note that individuals between 15 and 17 years old are experiencing a decrease. At the same time, the prevalence of consumption remains stable, and there is a decrease in the acquisition of marijuana from illegal sources. Rubin-Kahana2022 found different sources with conflicting data for Canadian teenagers, some studies report an increase, but most do not show a pronounced increase, demonstrating ambiguity in the analysis. In the case of Uruguay, Laqueur2020 sought to study the effect of legalization on high school students in Montevideo and regions within the country after legalization, but found no evidence of an impact on cannabis consumption or perceived consumption risk. However, they did find an increase in student perception of cannabis availability after legalization.

In addition to the above, polydrug use, which involves the combination of several substances, has become more visible. Regarding this, the most common combination in the Americas is that of cannabis with stimulants (such as cocaine and ecstasy), followed by opioids with stimulants, and finally, cannabis with opioids. The growing trend of polydrug use poses significant risks to consumers due to the interaction between substances UNODC2022, constituting an essential topic on the global public agenda. Furthermore, since the 1970s, the question has arisen as to whether access to marijuana leads to an increase in the use of other more problematic substances, with the so-called “gateway hypothesis" DeSimone1998, Kandel1975. Regarding this, Kandel1992 argue that the progression to more toxic drugs depends on prior use of cigarettes or alcohol, then marijuana, and then more harmful substances. Among the new trends in consumption are an expansion in the forms of use and more potent products Hammond2021, Rubin-Kahana2022. Furthermore, researchers Guttmannova2021, Kim2021 discovered a positive association between the frequency of cannabis and alcohol consumption. Additionally, Weinberger2022 established a positive relationship between marijuana consumption and frequency among tobacco users.

Finally, different studies have attempted to estimate the price elasticity of marijuana demand in the US, finding that it is inelastic Davis2016, Kilmer2014, Grossman2005, Nisbet1972, ranging from -0.69 to -0.26. Researchers conducted the same exercise in other countries, such as South Africa, Thailand and Australia, and found that demand in those countries is also in an inelastic range Riley2020, Sukharomana2017, Van2007. Gallet2014 observed that marijuana demand shows less responsiveness to prices than other drugs.

Our results suggests that selection into access is relevant for the intensive margin conditional in the extensive margin, and that modeling simultaneously the three stages is important. We found that marijuana is an inelastic good (the average elasticity is -0.45), there is not statistically significance regarding price heterogeneity among age groups, and the risk perception about its use is also not relevant on the intensive margin, but it is very relevant in the access and extensive margin. Moreover, we did not find heterogeneity regarding age splines in the intensive margin, but there is also relevant heterogeneity regarding the access and extensive margin. In general, we found that demographic and socioeconomic features are important to explain the three stages of marijuana demand in Colombia, all parameters estimates give intuitive results, for instance, women have less probability of having access and using marijuana, and their consumption is 45.5% lower than consumption of men.

Regarding the counterfactual exercises, we found that legalization of marijuana implies an overall increase of the probability of use in 0.7 percentage points (p.p.), from 2.3% pre-legalization to 3.0% under legalization. The population group that faces the higher increase in the probability of consumption is younger individuals (20s age spline) with a low or medium risk perception about using marijuana, this is 3.8 p.p., from 13.6% to 17.4%. In addition, tax revenues from taxation to marijuana may fluctuate between USD 11.0 million to USD 54.2 million, this depends on the tax setting. A potential benchmark is USD 32 million, where marijuana tax is US\cent 37.8, which implies a price equal to US\cent 39.1, half the price of actual marijuana for individuals with access.

After this introduction, we show our econometric framework in Section (ref). Section (ref) shows the results of the demand of marijuana in Colombia using information from the National Survey on the Consumption of Psychoactive Substances in 2019. Section (ref) shows results of tests regarding exclusionary restrictions and some robustness checks. Section (ref) shows results of potential implications of marijuana legalization on the extensive and intensive margins, and potential revenues that the government would have from this public policy. Concluding remarks are shown in Section (ref).

Econometric approach

Drug access

We set $A_{im}$ indicating if individual $i$ in market $m$ has access to marijuana,

align[align omitted — 118 chars of source]

where $U_{im}^a=\bm{w}_i^{\top}\boldsymbol{\alpha}_a+o_i\tau_a+\boldsymbol{\omega}_a+d_{im}\beta_a+V_{im}$ is the access latent variable, $i=1,\dots,N$, $m=1,\dots,M$.

Equation (ref) defines if individual $i$ has access to marijuana in market $m$ ($A_{im} = 1$), or not ($A_{im} = 0$), if she/he has a net positive utility from it. The latter is a function of socioeconomic and demographic characteristics ($\bm{w}_i$) such as age splines (teenager or 20’s, 30’s, 40’s, and 50’s or older), socioeconomic strata (low, medium and high), years of education, gender, mental and physical health status (good or bad), a dummy variable indicating if friends or family members use marijuana, and risk perception about using marijuana (low, medium and high). The latter variable is associated with mental, physical or/and social risks. There is no legal risk for marijuana users in Colombia as it is legally allowed the personal dose of regular cannabis (up to 20 grams). The risk perception is potentially influenced by the public policy, as marketing campaigns may affect risk perception about marijuana use. We also control for a dummy variable indicating previous consumption of alcohol and cigarette ($o_i$), characteristics of the market such regional-fixed effects ($\boldsymbol{\omega}_a$), and presence of drug dealers in the neighborhood ($d_{im}$). The latter is a supply side variable that should affect access to marijuana. The location parameters are $\boldsymbol{\alpha}_a$, $\boldsymbol{\omega}_a$ and $\beta_a$, and we assume $V_{im}\sim N(0,1)$, where the variance is set to 1 due to scale identification issues.

Drug extensive margin

The latent variable $U^c_{im} = \bm{w}^{\top}_i \boldsymbol{\alpha}_c + \omega_c + o_i\tau_c + E_{im}$ defines drug use (extensive margin),

align[align omitted — 154 chars of source]

We observe if individual $i$ uses marijuana in market $m$ ($C_{im} = 1$), or not ($C_{im} = 0$). Individuals use marijuana if they have a net positive utility from it, conditional on having access ($A_{im} = 1$). Otherwise, individuals do not use marijuana if they do not have a net positive utility from it conditional on having access. Observe that we have missing values regarding consumption when individuals do not have access to marijuana. This is due to individuals without access may or may not have net positive utility from marijuana use, therefore, we do not have information about their potential preferences, and consequently, this set of individuals do not contribute to identify these parameters.

The net indirect utility defining drug use depends on demographic and socioeconomic characteristics ($\bm{w}_i$), regional-fixed effects ($\omega_c$), and a dummy indicating previous consumption of alcohol and cigarette ($o_i$). The latter variable is due to the gateway hypothesis which indicates that previous consumption of these legal substances precedes use of illicit drugs Kandel1975,Kandel1992. Observe that we do not use price of marijuana in the extensive margin equation as due to its low price in Colombia (US\cent 83 per joint), this variable should not affect the extensive margin. We assume that $E_{im} \sim N(0, 1)$ due to scale identification issues, and the location parameters are $\boldsymbol{\alpha}_c$, $\omega_c$ and $\tau_c$.

Drug intensive margin

We model the consumption quantity (intensive margin), as a function of socioeconomic and demographic characteristics ($\bm{w}_i$), regional-fixed effects ($\omega_y$), price of marijuana ($p_i$), and the interaction between price and age brackets ($w_i^{\text{Age(j)}}\times p_{i}$, where Age(j) refers to j-th age bracket). The latter due to potential heterogeneity regarding price sensitivity among age splines. Observe that we have prices at individual level weighted by quality (see Appendix (ref) sections (ref) and (ref)), this implies that endogeneity would not be a concern due to there is no demand-supply simultaneity neither endogeneous variation in the stochastic error due to quality of marijuana jacobi2016marijuana.

We consider in the intensive margin equation incidental truncation, that is, there are missing values for consumption quantity when individuals do not have access, or having access reporting not to use marijuana.

align[align omitted — 258 chars of source]

Observe that the truncation setting given in equation (ref), as we are not taking into account zero consumption. This is because some marijuana users reporting not consumption due to social stigma. The location parameters related to consumption level are $\boldsymbol{\alpha}_y$, $\omega_y$, $\gamma_y$ and $\gamma_{yj}$, and we assume $W_{im}\sim N(0, \sigma^2_y)$.

Correlation on unobservable variables

We model simultaneously the three stages (access, extensive and intensive margins) due to there should be unobservable variables that drive these stages. Thus, we assume $\Xi_{im} = \left[V_{im} \ E_{im} \ W_{im}\right]^{\top}\sim N_3(\bm{0}, \boldsymbol{\Sigma})$, where

align[align omitted — 200 chars of source]

Observe that if $\boldsymbol{\Sigma}$ is a diagonal matrix, the three stages are independent, and we can perform inference estimating each equation separately. This maybe a situation where there is no strategic search of drug dealers by users, and that more intense consumers do not make a greater effort to get drug dealers when arriving to a new market. On the other hand, if there is endogenous access, which means $\sigma_{ca}\neq 0$ and/or $\sigma_{ya}\neq 0$, we should model simultaneously these equations to get good sampling properties of our estimators. In addition, $\sigma_{yc}$ takes into account potential unobserved dependence between extensive and intensive margins conditional on selection into access.

Estimation strategy

Observe that modeling the joint distribution of access, extensive and intensive margin implies to integrate over a multivariate space to recover the likelihood function. In addition, this is not a standard likelihood as incidental truncation implies that different sets of individuals contribute to different sets of parameters. Particularly, there are three different groups of individuals when taking the model setting given by equations (ref), (ref) and (ref): all individuals ($G_1$) contribute to estimate the location parameters in the access equation, individuals who report having access ($G_2$) contribute to estimate the location parameters in the extensive margin equation and $\sigma_{ac}$, and individuals who report to use marijuana ($G_3$) contribute to estimate the parameters of the intensive margin equation, and $\sigma_{ay}$, $\sigma_{cy}$ and $\sigma_{y}^2$.

Thus, we use data augmenting Tanner1987 to facilitate inference, we treat latent variables as parameters, such that the augmented model is,

align[align omitted — 574 chars of source]

where $\bm{x}_{im}^{s\top}$ is the vector of regressors associated with individual $i$ in market $m$ in stage $s = \left\{\text{access, consumption, quantity}\right\}$, and $\boldsymbol{\theta}^s$ is the vector of location parameters in stage s.

The likelihood function is

align[align omitted — 562 chars of source]

where $G = 3$, $\phi(\cdot|\tilde{\bm{X}}_{im}\tilde{\boldsymbol{\theta}}_{G_s},\tilde{\boldsymbol{\Sigma}}_{G_s})$ is the density function of a normal distribution with mean $\tilde{\bm{X}}_{im}\tilde{\boldsymbol{\theta}}_{G_s}$ and variance $\tilde{\boldsymbol{\Sigma}}_{G_s}$, and $\tilde{M} = f(G_s,M)$, where $f(G_s,M)$ is a function that takes as inputs a state ($G_s$) and a matrix ($M$), and returns as output the appropriate subset of rows and columns of $M$. For instance,

equation[equation omitted — 275 chars of source]
equation[equation omitted — 330 chars of source]

We use the Bayes’ rule to perform inference in our model. However, we implement our inferential algorithm in the unidentified parameter space as getting draws from the posterior distribution in the identified parameter space has a high computational cost and inferior mixing properties than a more straightforward Gibbs sampler traversing over the unidentified space Rossi2005. Thus, we set

align[align omitted — 218 chars of source]

We implement our Gibbs sampling algorithm using standard conjugate independent priors to obtain standard conditional posterior distributions that facilitate computation. In particular, we assume that $\pi(\boldsymbol{\theta},\boldsymbol{\Omega})=N(\boldsymbol{\theta}|\boldsymbol{\theta}_0,\boldsymbol{\Theta}_0)\times \pi(\boldsymbol{\Omega}|\bm{R}_0,r_0)$, that is, a multivariate normal distribution for the location parameters and an inverse Wishart distribution for the unidentified covariance matrix. We use non-informative hyperparameters in all our exercises, that is, $\boldsymbol{\theta}_0=\bm{0}$, $\boldsymbol{\Theta}_0=\text{diag}\left\{1,000\right\}$, $\bm{R}_0=\bm{I}_{3}$ and $r_0=3+2$.

Reparameterizing the data-augmented likelihood function from equation (ref) in terms of $\boldsymbol{\Omega}$, and using the previous prior density function, the Bayes’ rule implies that the posterior conditional distribution for the location parameters is $N(\boldsymbol{\theta}|\boldsymbol{\theta}_n,\boldsymbol{\Theta}_n)$ where $\boldsymbol{\Theta}_n=\left[\sum_{s=1}^{G}\sum_{i\in G_s}\bm{J}_{G_s}\tilde{\bm{X}}_{im}^{\top}\tilde{\boldsymbol{\Omega}}_{G_s}^{-1}\tilde{\bm{X}}_{im}\bm{J}_{G_s}^{\top}+\boldsymbol{\Theta}_0^{-1}\right]^{-1}$ and $\boldsymbol{\theta}_n=\boldsymbol{\Theta}_n\left[\sum_{s=1}^{G}\sum_{i\in {G_s}}\bm{J}_{G_s}\tilde{\bm{X}}_{im}^{\top}\tilde{\boldsymbol{\Omega}}_{G_s}^{-1}\tilde{\bm{T}}_{im}+\boldsymbol{\Theta}_0^{-1}\boldsymbol{\theta}_0\right]$, and

equation*[equation* omitted — 468 chars of source]

where $\bm{I}_{\left\{d\right\}}$ is a $d\times d$ identity matrix, $H$, $K$ and $L$ are the dimensions of $\boldsymbol{\theta}_s$.

We follow a sequential approach chib2009estimation,li2011estimation to get standard conditional posterior distributions that allow to recover $\boldsymbol{\Omega}$. In particular, all observations contribute to estimate $\omega_a^2$, and given the prior distribution of $\boldsymbol{\Omega}$, which implies that the prior distribution of $\omega_a^2$ is inverse gamma with parameters $r_{11,0}$ and $r_0-2$, where $r_{11,0}$ is the element 1,1-th of $\bm{R}_0$ greenberg2012introduction, then the posterior conditional distribution of $\omega_a^2$ is $IG(r_{11,n}, r_0-2+n)$ where $r_{11,n}=\sum_{i=1}^n(T_{1,im}-X_{1,im}\boldsymbol{\theta})^2 + r_{11,0}$, $T_{1,im}$ is the first element of $\bm{T}_{im} (U^a_{im})$ and $X_{1,im}$ is the first row of $\bm{X}_{im}$.

In the next stage, we set

align*[align* omitted — 145 chars of source]

and

equation[equation omitted — 91 chars of source]

Given the prior distribution of $\boldsymbol{\Omega}$, and a consistent partition of $\bm{R}_0$,

align*[align* omitted — 128 chars of source]

the prior distribution of $\omega_{c.1}^2$ is inverse gamma with parameters $r_{22.1,0}^2$ and $r_0$, where $r_{22.1,0}^2=r_{22,0}^2-r_{12,0}^2/r_{11,0}^2$.

In addition, we set

equation[equation omitted — 77 chars of source]

and given the prior distribution of $\boldsymbol{\Omega}$, the prior distribution of $\omega_{ca.1}|\omega_{c.1}^2$ is normal with mean $r_{21,0}/r_{11,0}^2$ and variance $\omega_{c.1}^2/r_{11,0}^2$.

We calculate

equation[equation omitted — 290 chars of source]

where $\bm{T}_{1:2,im}$ and $\bm{X}_{1:2,im}$ are the first and second rows of $\bm{T}_{im}$ and $\bm{X}_{im}$, respectively.

The posterior distribution of $\omega_{c.1}^2$ is inverse gamma with parameters $r_{22.1,n}^2$ and $r_0+|G_2|$, where $r_{22.1,n}^2=r_{22,n}^2-r_{12,n}^2/r_{11,n}^2$, and $|G_2|$ is the number of individuals in group two.

The posterior distribution of $\omega_{ca.1}$ conditional on $\omega_{c.1}^2$ is normal with mean $r_{21,n}/r_{11,n}^2$ and variance $\omega_{c.1}^2/r_{11,n}^2$.

We can recover $\boldsymbol{\Omega}_{22}$ using equation (ref), such that $\omega_{ca}=\omega_{ca.1}\omega_{a}^2$, and equation (ref), where we have that $\omega_c^2=\omega_{c.1}^2+\omega_{ca}^2/\omega_a^2$.

We set

align*[align* omitted — 179 chars of source]

where $\boldsymbol{\Omega}_{32}=\left[\omega_{ya} \ \omega_{yc}\right]$, and

equation[equation omitted — 137 chars of source]

and given the prior distribution of $\boldsymbol{\Omega}$, and a consistent partition of $\bm{R}_0$,

align*[align* omitted — 135 chars of source]

the prior distribution of $\omega_{y.1}^2$ is inverse gamma with parameters $r_{33.1,0}^2$ and $r_0$, where $r_{33.1,0}^2=r_{33,0}^2-\bm{R}_{32,0}\bm{R}_{22,0}^{-1}\bm{R}_{23,0}$.

Given

equation[equation omitted — 112 chars of source]

where the prior distribution of $\boldsymbol{\Omega}_{32.1}|\omega_{y.1}^2$ is matrix normal with mean $\bm{R}_{32,0}\bm{R}_{22,0}^{-1}$ and scale matrices $\bm{R}^{-1}_{22,0}$ and $\omega^{2}_{y.1}$.

Given

equation[equation omitted — 288 chars of source]

The posterior distribution of $\boldsymbol{\Omega}_{32.1}$ conditional on $\omega^2_{y.1}$ is matrix normal with mean $\bm{R}_{32,n}\bm{R}^{-1}_{22,n}$ and scale matrices $\bm{R}^{-1}_{22,n}$ and $\omega^2_{y.1}$. We can recover $\boldsymbol{\Omega}$ using equation (ref), such that $\boldsymbol{\Omega}_{32}=\boldsymbol{\Omega}_{32.1}\boldsymbol{\Omega}_{22}$, and equation (ref), where we have that $\omega^2_y=\omega^2_{y.1}+\boldsymbol{\Omega}_{32}\boldsymbol{\Omega}^{-1}\boldsymbol{\Omega}_{23}$.

The posterior conditional distributions of $U_{im}^a$ and $U_{im}^c$ are truncated normal in the interval $(-\infty,0]$ if $U_{im}^l=0$, and $(0,\infty)$ if $U_{ij}^l=1$, respectively, $l=\left\{a,c\right\}$. Their conditional means are $m_{l,im}=\tilde{\bm{X}}_{l,im}\tilde{\boldsymbol{\theta}}+\tilde{\boldsymbol{\Omega}}_{l,-l}\tilde{\boldsymbol{\Omega}}_{-l,-l}^{-1}(\tilde{\bm{T}}_{-l,im}-\tilde{\bm{X}}_{-l,im}\tilde{\boldsymbol{\theta}})$, and conditional variances $\tau_{l}^2=\tilde{\omega}_{ll}^2-\tilde{\boldsymbol{\Omega}}_{l,-l}\tilde{\boldsymbol{\Omega}}_{-l,-l}^{-1}\tilde{\boldsymbol{\Omega}}_{-l,l}$, where $\tilde{\bm{T}}_{-l,im}$ is the vector $\tilde{\bm{T}}_{im}$ excluding the $l$-th component, $\tilde{\bm{X}}_{l,im}$ is the $l$-th row of matrix $\tilde{\bm{X}}_{im}$, $\tilde{\bm{X}}_{-l,im}$ is the matrix $\tilde{\bm{X}}_{im}$ without the $l$-th row, $\tilde{\boldsymbol{\Omega}}_{l,-l}$ is the $l$-th row of $\tilde{\boldsymbol{\Omega}}$ excluding the $l$-th element, $\tilde{\boldsymbol{\Omega}}_{-l,-l}$ is equal to $\tilde{\boldsymbol{\Omega}}$ excluding the $l$-th row and $l$-th column, and $\tilde{\omega}^2_{ll}$ is the $ll$ element of $\tilde{\boldsymbol{\Omega}}$.

Our econometric framework differs from other literature regarding modeling marijuana consumption in a few fundamental ways. In particular, we follow a simultaneous three-part modeling approach that takes into account that access is a necessary condition for use, such that the latter is endogenously determined with extensive and intensive margins. We also incorporate that use is a necessary condition for the intensive margin, and allow for unobserved correlation between these stages. In addition, we consider the incidental truncation issue due to missing reports when individuals report no to have access to marijuana, or when reporting access, they report not to use it. Omitting access restrictions, potential correlation on unobservable variables and/or incidental truncation may generate inconsistent estimators. There is also a clear link between our modeling strategy and the policy that we want to analyze, as marijuana legalization implies basically free access for all potential users, such that the “breaking the law" hindrance will disappear, this is formally $P(A_{imM}=1)=1$ in our econometric framework, and will be the basis for our counterfactual exercises.

Marijuana demand in Colombia

Data

We apply our approach to the problem of estimating the demand for marijuana in Colombia. We leveraged individual-level data on the consumption of psychoactive substances representative of the entire country in 2019. In particular, our data come from the National Survey on the Consumption of Psychoactive Substances (ENCSPA 2019), which is a national representative survey, performed by the government statistical department of Colombia (DANE) aiming to measure both legal and illegal substance abuse within the population. The survey randomly sampled households from several municipalities in Colombia. It targeted individuals between 12 and 65 years of age, who were selected randomly from all the household members that met the age criterion. The enumerators privately performed the survey. If the chosen person was absent during the survey, the enumerator should return later but was not allowed to change the individual to be interviewed.

We are especially interested in variables describing marijuana consumption patterns based on the literature review and data availability. Our measure of marijuana consumption takes into account quality based on the tetrahydrocannabinol (THC) content (see Appendix (ref), subsection (ref)). In addition, we use a nearest neighbor algorithm to impute prices for those individuals who report not to consume marijuana, and consequently, who do not report average price (see Appendix (ref), subsection (ref)),\footnote{Observe that posterior estimates of our model composed by equation (ref), (ref) and (ref) does not require these prices.} and a basic counting algorithm to construct risk perception about drug use (see Appendix (ref), subsection (ref)).

Table (ref) presents the summary statistics for the key variables in our sample. We consider three different outcome variables, one for each part of the three-part model: First, whether the individual has access or not to marijuana by reporting that it would be easy for her/him to get it. Second, whether the individual is a consumer, specifically if they had consumed marijuana during the last 12 months, and finally, the quantity of marijuana consumed on average per month. The quantity consumed is measured in the number of cigarettes or joints of marijuana; although it is not specified the number of grams, there is some consistency in the sizes of common joints, which are usually about 1 gram.

table[table omitted — 4,713 chars of source]

Our sample population consists of 49,414 individuals. Less than 60% report having access (row 1, column 1); of those reporting having access, approximately 4% report consuming marijuana during the last 12 months (row 2, column 3). Finally, the consumers take 23 joints of marijuana per month on average, with an average price of USD 0.84 (rows 3 and 6, column 5). By definition, and aiming to achieve consistency in the model, individuals who are consumers should have access; this is internally consistent in the data for more than 98% of the population and manually inputted for the remaining 2%\footnote{We replace access equal to one if the person is a consumer for two main reasons. First, it is not intuitive for people who are current users to report that they cannot get marijuana so it is likely an error in responding on the part of these individuals or misreporting due to stigma. Second, for the model to be internally consistent by construction, all users must have access. To review possible changes to the definition of the access measure, we conducted multiple exercises to validate the robustness of the results.} The representative individual in the survey is a female with complete secondary education in her 20s living in a low socioeconomic stratum. More than half of the people are workers, almost 80% report having good mental health and 77% report having good physical health, 91% have a high-risk perception of the use of marijuana, and 36% report having a peer (family or friend) who is a consumer of marijuana. Approximately, 38% of the sample declares to have a drug dealer in the neighborhood, and 32% report to have consumed alcohol and cigarette.

Model results

We perform inference of the model in system (ref) running 6,000 iterations with a burn-in equal to 1,000 and a thin parameter equal to 5, thus we have 1,000 effective posterior draws. We compute several diagnostics to assess the convergence and stationarity of the posterior chains. In general, the posterior chains look good. Particularly, all location parameters have dependence factors that are less than 5, actually most of them less than 2, using Raftery1992's diagnostic, with a 95% probability of obtaining an estimate in the interval $2.5\% \pm 1.0\%$. Regarding Heidelberger1983's and Geweke1992's tests at 5% significance level, all parameter estimates pass the former test, and 147 out of 163 location parameters passed the latter test. The former uses the Cramer-von-Mises statistic to test the null hypothesis that the sampled values come from a stationary distribution, and the latter tests for equality of the posterior means using the first 10% and the last 50% of the Markov chains. Regarding the scale parameters, we have that all parameters have a dependence factor less than 5, and one out of 4 does not passed the conventional values of the Heidelberger1983 and Geweke1992 diagnostics.

table[table omitted — 5,525 chars of source]
comment\begin{table}[h!] \caption{Posterior results of location parameters: Marijuana demand in Colombia.} \begin{threeparttable} \resizebox{0.75\textwidth}{!}{\begin{minipage}{\textwidth} \begin{tabular}{c c c c c c c c}\hline \multirow{2}{*}{Variable} & \multicolumn{3}{c}{Univariate} & \multicolumn{3}{c}{Multivariate} \\ & 2.5% & 50% & 97.5% & 2.5% & 50% & 97.5% \\ \hline \multicolumn{7}{c}{Access equation} \\ \hline Constant & -0.53 & -0.42 & -0.31 & -0.51 & -0.40 & -0.29 \\ Drug dealer in neighborhood & \textbf{0.41} & \textbf{0.43} & \textbf{0.46} & \textbf{0.40} & \textbf{0.42} & \textbf{0.45} \\ Age 30s & \textbf{-0.10} & \textbf{-0.06} & \textbf{-0.03} & \textbf{-0.09} & \textbf{-0.06} & \textbf{-0.02} \\ Age 40s & \textbf{-0.24} & \textbf{-0.20} & \textbf{-0.16} & \textbf{-0.23} & \textbf{-0.20} & \textbf{-0.16} \\ Age 50s and older & \textbf{-0.37} & \textbf{-0.34} & \textbf{-0.31} & \textbf{-0.38} & \textbf{-0.34} & \textbf{-0.31} \\ Stratum medium & \textbf{-0.07} & \textbf{-0.04} & \textbf{-0.02} & \textbf{-0.07} & \textbf{-0.05} & \textbf{-0.02} \\ Stratum high & \textbf{-0.08} & \textbf{-0.03} & \textbf{-0.01} & -0.08 & -0.04 & 0.01 \\ High risk perception drug use & \textbf{0.21} & \textbf{0.28} & \textbf{0.34} & \textbf{0.19} & \textbf{0.26} & \textbf{0.32} \\ Medium risk perception drug use & \textbf{0.49} & \textbf{0.57} & \textbf{0.66} & \textbf{0.45} & \textbf{0.54} & \textbf{0.62} \\ Years of education & \textbf{0.01} & \textbf{0.02} & \textbf{0.02} & \textbf{0.01} & \textbf{0.02} & \textbf{0.02} \\ Female & \textbf{-0.40} & \textbf{-0.37} & \textbf{-0.34} & \textbf{-0.40} & \textbf{-0.37} & \textbf{-0.35} \\ Good mental health & \textbf{-0.18} & \textbf{-0.15} & \textbf{-0.12} & \textbf{-0.18} & \textbf{-0.15} & \textbf{-0.12} \\ Marijuana consumers in network & \textbf{0.71} & \textbf{0.74} & \textbf{0.76} & \textbf{0.71} & \textbf{0.74} & \textbf{0.76} \\ Working & \textbf{0.12} & \textbf{0.15} & \textbf{0.17} & \textbf{0.12} & \textbf{0.15} & \textbf{0.17} \\ Region fixed effects & $\checkmark$ & $\checkmark$ & $\checkmark$ & $\checkmark$ & $\checkmark$ & $\checkmark$ \\ \hline \multicolumn{7}{c}{Extensive margin equation} \\ \hline Constant & \textbf{-2.08} & \textbf{-1.73} & \textbf{-1.41} & \textbf{-2.07} & \textbf{-1.70} & \textbf{-1.37} \\ Alcohol and cigarette user & \textbf{0.88} & \textbf{0.95} & \textbf{1.03} & \textbf{0.86} & \textbf{0.93} & \textbf{1.01} \\ Age: 30s & \textbf{-0.48} & \textbf{-0.39} & \textbf{-0.29} & \textbf{-0.47} & \textbf{-0.38} & \textbf{-0.29} \\ Age: 40s & \textbf{-0.72} & \textbf{-0.60} & \textbf{-0.49} & \textbf{-0.72} & \textbf{-0.59} & \textbf{-0.48} \\ Age: 50s and older & \textbf{-1.15} & \textbf{-1.04} & \textbf{-0.92} & \textbf{-1.14} & \textbf{-1.02} & \textbf{-0.89} \\ Stratum medium & \textbf{0.00} & \textbf{0.08} & \textbf{0.16} & \textbf{0.01} & \textbf{0.08} & \textbf{0.16} \\ Stratum high & \textbf{0.01} & \textbf{0.14} & \textbf{0.25} & \textbf{0.01} & \textbf{0.13} & \textbf{0.25} \\ High risk perception drug use & \textbf{-0.92} & \textbf{-0.79} & \textbf{-0.66} & \textbf{-0.93} & \textbf{-0.79} & \textbf{-0.65} \\ Medium risk perception drug use & -0.27 & -0.12 & 0.04 & -0.27 & -0.11 & 0.04 \\ Years of education & -0.02 & -0.01 & 0.00 & -0.02 & -0.01 & 0.00 \\ Female & \textbf{-0.53} & \textbf{-0.45} & \textbf{-0.38} & \textbf{-0.53} & \textbf{-0.45} & \textbf{-0.38} \\ Good mental health & \textbf{-0.28} & \textbf{-0.21} & \textbf{-0.13} & \textbf{-0.29} & \textbf{-0.20} & \textbf{-0.12} \\ Marijuana consumers in network & \textbf{0.89} & \textbf{0.98} & \textbf{1.09} & \textbf{0.87} & \textbf{0.97} & \textbf{1.08} \\ Working & \textbf{-0.17} & \textbf{-0.10} & \textbf{-0.03} & \textbf{-0.17} & \textbf{-0.10} & \textbf{-0.03} \\ Region fixed effects & $\checkmark$ & $\checkmark$ & $\checkmark$ & $\checkmark$ & $\checkmark$ & $\checkmark$ \\ \hline \multicolumn{7}{c}{Intensive margin equation} \\ \hline Constant & \textbf{4.54} & \textbf{6.21} & \textbf{7.89} & \textbf{2.17} & \textbf{3.8} & \textbf{5.53} \\ $\log\left\{\text{price of marijuana}\right\}$ & \textbf{-0.67} & \textbf{-0.49} & \textbf{-0.32} & \textbf{-0.59} & \textbf{-0.40} & \textbf{-0.21} \\ Age 30s & -2.16 & 0.26 & 2.78 & -1.81 & 0.86 & 3.24 \\ Age 40s & -3.54 & 0.62 & 4.38 & -3.16 & 0.89 & 4.49 \\ Age 50s and older & -8.19 & -3.68 & 0.92 & -6.17 & -2.12 & 1.91 \\ Stratum medium & -0.23 & -0.04 & 0.15 & -0.20 & -0.04 & 0.12 \\ Stratum high & -0.49 & -0.21 & 0.06 & -0.47 & -0.22 & 0.00 \\ High risk perception drug use & \textbf{-0.94} & \textbf{-0.65} & \textbf{-0.36} & -0.41 & -0.16 & 0.08 \\ Medium risk perception drug use & \textbf{-0.88} & \textbf{-0.56} & \textbf{-0.26} & -0.17 & 0.12 & 0.41 \\ Years of education & \textbf{-0.11} & \textbf{-0.09} & \textbf{-0.06} & \textbf{-0.08} & \textbf{-0.06} & \textbf{-0.04} \\ Female & \textbf{-0.73} & \textbf{-0.51} & \textbf{-0.32} & \textbf{-0.87} & \textbf{-0.68} & \textbf{-0.53} \\ Good mental health & -0.13 & 0.06 & 0.24 & -0.20 & -0.04 & 0.13 \\ Marijuana consumers in network & \textbf{0.05} & \textbf{0.40} & \textbf{0.76} & \textbf{0.76} & \textbf{0.98} & \textbf{1.25} \\ Working & -0.03 & 0.15 & 0.32 & \textbf{0.01} & \textbf{0.16} & \textbf{0.33} \\ Age 30s $\times \log\left\{\text{price of marijuana}\right\}$ & -0.33 & 0.00 & 0.32 & -0.40 & -0.10 & 0.25 \\ Age 40s $\times \log\left\{\text{price of marijuana}\right\}$ & -0.54 & -0.05 & 0.48 & -0.59 & -0.10 & 0.42 \\ Age 50s and older $\times \log\left\{\text{price of marijuana}\right\}$ & -0.14 & 0.45 & 1.01 & -0.32 & 0.21 & 0.71 \\ Region fixed effects & $\checkmark$ & $\checkmark$ & $\checkmark$ & $\checkmark$ & $\checkmark$ & $\checkmark$ \\ \hline \end{tabular} \begin{tablenotes}[para,flushleft] \textit{Notes}: Bold font indicates 95% credible intervals that do not embrace 0. This means that there is a probability higher than 95% that the variable has a statistically significant effect.\\ Upper part shows posterior results of the access equation, middle part shows posterior results of the extensive margin equation, and bottom part shows posterior results of the intensive margin equation. There are 38 region fixed effects in each equation. The columns labeled univariate show posterior results of models assuming exogenous stages (access, extensive and intensive margins). The columns labeled multivariate show posterior results of the endogenous three-part model. There are meaningful differences in the intensive margin equation due to a statistically significant unobserved co-variation between the access and intensive margin equations (see Table (ref)).\\ \textit{Source}: Authors' construction. \end{tablenotes} \end{minipage}} \end{threeparttable} \end{table}

Table (ref) reports the posterior estimates. Columns labeled univariate show the posterior results of estimating univariate models, that is, probit models for the access and extensive margin equations, and a linear model for the logarithm of marijuana consumption (intensive margin). Columns labeled multivariate show the posterior results of modeling simultaneously equations (ref), (ref) and (ref) taking into account truncation.

We observe that univariate and multivariate models give similar results regarding the access and extensive margin equations. This is due to Table (ref) suggesting that these two equations are exogenous. However, we observe in Table (ref) that posterior estimates of the univariate and multivariate models are different regarding the intensive margin (see columns (3) and (6)). This is explained by the fact that there is endogeneity between access and intensive margin. In particular, Table (ref) indicates that the unobserved co-variation between these equations is statistically significant. This suggests that conditional on the extensive margin, more frequent marijuana consumers make higher effort to have access to this drug.

Columns (1) and (4) in Table (ref) show that having a drug dealer in the neighborhood increases the probability of having access to marijuana. In particular, the probability of having access to marijuana for the representative individual based on the descriptive statistics (see the last paragraph of the previous subsection) increases by 16.9 percentage points due to presence of drug dealer in the neighborhood, that is, from 36.5% to 53.4%, given the results in column (4). In addition, the probability of having access is lower for women, individuals who declare to have good mental health status, and are older. On the other hand, individuals who live in a low socioeconomic stratum, have a high or medium risk perception about marijuana use, who have more years of education, friends or family members that consume marijuana, and work, have a higher probability of having access to this drug.

We observe in columns (2) and (5) in Table (ref) that previous consumption of alcohol and cigarette increases the probability of consumption. This is evidence for the gateway drug hypothesis. In addition, the probability of use increases with socioeconomic strata and having friends or family members who also consume. On the other hand, the probability of use decreases with age, risk perception, being female, having good mental health and being a worker. The results in Table (ref) allows to predict the potential significant effects of a public policy that increases the risk perception about using marijuana. For instance, the posterior estimates in column (5) indicates that the probability of using marijuana is equal to 40.5% for a man in his 20s that works, who has 12 years of education, consumes alcohol and cigarettes, lives in a low socioeconomic stratum, whose mental and physical health status is good, has friends who consume marijuana and has a low risk perception about using marijuana. On the other hand, this probability is equal to 15.2% for the same individual, except that his risk perception about marijuana use is high, that is, there is a decrease of 25.3 percentage points due to changing the risk perception.

Columns (3) and (6) in Table (ref) show that the marijuana is an inelastic good. Particularly, column (6) shows that this elasticity is on average equal to -0.45, and is statistically significant; this agrees with previous literature Gallet2014,Davis2016,Sukharomana2017,Riley2020. We also observe in this column that there is not statistically significant heterogeneity regarding price sensitivity between age groups. This is relevant from a public policy perspective as a valid concern regarding marijuana legalization is the implications of price variations on young individuals. The multivariate setting also suggests that the effect of risk perception is through the extensive margin, rather than directly on the intensive margin, as is suggested by the univariate modeling framework. This is, conditional on the extensive margin, the risk perception does not have any effect on the intensive margin. Finally, we see that one additional year of education decreases marijuana consumption by 6.9%, and women consume 45.5% ($\exp(-0.607)-1$) less than men, whereas having a network where some individuals consume marijuana increase marijuana consumption by 85.7% ($\exp(0.619)-1$). All these variables are statistically significant.

table[table omitted — 1,319 chars of source]
comment\begin{table}[h!] \caption{Posterior results of scale parameters: Marijuana demand in Colombia.} \begin{threeparttable} \resizebox{0.35\textwidth}{!}{\begin{minipage}{\textwidth} \begin{tabular}{c c c c}\hline Parameter & 2.5% & 50% & 97.5% \\ \hline & \multicolumn{3}{c}{Multivariate}\\ \hline $\sigma_{ca}$ & -0.02 & 0.00 & 0.02 \\ $\sigma_{ya}$ & 7.28 & 7.93 & 8.60 \\ $\sigma_{yc}$ & -0.19 & 0.00 & 0.17 \\ $\sigma_{y}^2$ & 2.85 & 3.11 & 3.38 \\ \hline & \multicolumn{3}{c}{Univariate}\\ \hline $\sigma_{y}^2$ & 1.93 & 2.10 & 2.28 \\ \hline \end{tabular} \begin{tablenotes}[para,flushleft] Notes: Columns labeled multivariate show posterior estimates of the identified covariance matrix. Columns labeled univariate shows the posterior results of the variance of the intensive margin equation.\\ There is a 95% probability that the unobserved co-variation between the access equation and the intensive margin is between 7.28 and 8.60. The univariate model of the intensive margin shows a lower variance than the multivariate model. \\ Source: Authors' construction. \end{tablenotes} \end{minipage}} \end{threeparttable} \end{table}

Statistical checks

Exclusionary restrictions

It is well known that we can achieve identification of causal effects in nonlinear models without exclusion restrictions mcmanus1992common. However, exclusion restrictions improve inference due to reducing variability of estimates because of data variability munkin2003bayesian. We exclusively use presence of drug dealers in the neighborhood in the access equation as this variable affects drug supply which helps to identify demand parameters. Particularly, we would expect that presence of drug dealers would positively affect the probability of an individual having access to marijuana. We argue that the effect of this variable on the extensive and intensive margins should be just through the access.

However, if this supply side variable affects directly the net utility of using marijuana, then our exclusion restriction would not be valid. We try to test this restriction using the subset of individuals who were offered marijuana, as a consequence, they do not have to search for this drug, which means that they are relatively free of the selection issue. We check the statistical significance of this supply side variable running a probit model on the access equation in this subsample. Column 1 of Table (ref) shows the posterior results using a non-informative normal prior distribution, the number of iterations of the Gibbs sampler is 6,000, a burn-in equal to 1,000, and a thin parameter equal to 5. We observe that the presence of drug dealer is not statistically significant in the extensive margin equation. Hence, this result suggests that “drug dealer in neighborhood" has validity as an exclusionary restriction.

We do not use presence of drug dealer in the neighborhood neither if an individual has consumed alcohol and cigarettes any time in her/his life in the intensive margin equation. The gateway drug hypothesis would support the latter variable as pattern of legal substance use would precede the use of illicit substances Kandel1975,Kandel1992. This means that individuals who have consumed substances like alcohol and nicotine would have a higher probability of accessing and using marijuana. However, past or present consumption of these substances should not directly affect the intensive margin. Column 2 of Table (ref) shows the posterior results of estimating the intensive margin of marijuana, that is, the logarithm of quantity, as function of these two variables as well as all regressors in equations (ref) using the subset of individuals that were offered marijuana, that is, the set of individuals that presumably is exogenous to the access equation. Given that we still have endogeneity due to the extensive margin in this set of individual, this is not a formal test of exclusionary restrictions. However, the fact that presence of drug dealer in the neighborhood, and consumption of alcohol and cigarettes are not statistical significant to explain the intensive margin, suggests that these variables have some validity as exclusionary restrictions.

comment\begin{table}[h!] \caption{Exclusionary restrictions validation: Marijuana demand in Colombia for marijuana offer.} \begin{threeparttable} \resizebox{0.75\textwidth}{!}{\begin{minipage}{\textwidth} \begin{tabular}{c c c c}\hline \multirow{2}{*}{Variable} & \multicolumn{3}{c}{Posterior percentiles} \\ & 2.5% & 50% & 97.5% \\ \hline \multicolumn{4}{c}{Extensive margin equation} \\ \hline Constant & -1.96 & -1.55 & -1.20 \\ Drug dealer in neighborhood & -0.01 & 0.06 & 0.14 \\ Alcohol and cigarette use & 0.75 & 0.84 & 0.93 \\ Age: 30s & \textbf{-0.44} & \textbf{-0.34} & \textbf{-0.24} \\ Age: 40s & \textbf{-0.63} & \textbf{-0.50} & \textbf{-0.37} \\ Age: 50s and older & \textbf{-1.12} & \textbf{-0.97} & \textbf{-0.83} \\ Stratum medium & \textbf{0.02} & \textbf{0.10} & \textbf{0.18} \\ Stratum high & -0.07 & 0.07 & 0.20 \\ High risk perception drug use & \textbf{-0.97} & \textbf{-0.82} & \textbf{-0.66} \\ Medium risk perception drug use & -0.36 & -0.17 & 0.01 \\ Years of education & \textbf{-0.03} & \textbf{-0.02} & \textbf{-0.01} \\ Female & \textbf{-0.50} & \textbf{-0.42} & \textbf{-0.34} \\ Good mental health & \textbf{-0.28} & \textbf{-0.20} & \textbf{-0.11} \\ Marijuana consumers in network & \textbf{0.84} & \textbf{0.95} & \textbf{1.08} \\ Working & \textbf{-0.19} & \textbf{-0.10} & \textbf{-0.02} \\ Region fixed effects & $\checkmark$ & $\checkmark$ & $\checkmark$ \\ \hline \multicolumn{4}{c}{Intensive margin equation} \\ \hline Constant & \textbf{4.58} & \textbf{6.45} & \textbf{8.31} \\ Drug dealer in neighborhood & -0.17 & 0.09 & 0.35 \\ Alcohol and cigarette use & -0.01 & 0.18 & 0.37 \\ $\log\left\{\text{price of marijuana}\right\}$ & \textbf{-0.70} & \textbf{-0.51} & \textbf{-0.33} \\ Age 30s & -2.60 & 0.14 & 2.82 \\ Age 40s & -3.84 & 0.47 & 4.66 \\ Age 50s and older & -8.46 & -3.89 & 1.35 \\ Stratum medium & -0.23 & -0.02 & 0.17 \\ Stratum high & -0.36 & -0.08 & 0.22 \\ High risk perception drug use & \textbf{-1.04} & \textbf{-0.71} & \textbf{-0.38} \\ Medium risk perception drug use & \textbf{-1.00} & \textbf{-0.65} & \textbf{-0.30} \\ Years of education & \textbf{-0.12} & \textbf{-0.09} & \textbf{-0.06} \\ Female & \textbf{-0.71} & \textbf{-0.50} & \textbf{-0.28} \\ Good mental health & -0.06 & 0.14 & 0.35 \\ Marijuana consumers in network & -0.15 & 0.26 & 0.68 \\ Working & -0.01 & 0.17 & 0.37 \\ Age 30s $\times \log\left\{\text{price of marijuana}\right\}$ & -0.34 & 0.01 & 0.37 \\ Age 40s $\times \log\left\{\text{price of marijuana}\right\}$ & -0.57 & -0.02 & 0.53 \\ Age 50s and older $\times \log\left\{\text{price of marijuana}\right\}$ & -0.14 & 0.48 & 1.07 \\ Region fixed effects & $\checkmark$ & $\checkmark$ & $\checkmark$ \\ \hline \end{tabular} \begin{tablenotes}[para,flushleft] \textit{Notes}: This table shows posterior percentiles of parameters of the extensive margin (top part) and intensive margin (bottom part) using the subset of individuals that were offered marijuana. These results suggest plausibility of the exclusionary restrictions. \\ \textit{Source}: Authors' construction. \end{tablenotes} \end{minipage}} \end{threeparttable} \end{table}

Finally, we have included marijuana price in the intensive margin equation, but not in the access and extensive margin equations. Although this is not necessary for identification, we consider that excluding price from these equations make sense. First, the argument to exclude marijuana price from the access equation is that individuals who do not have access are unlikely to know the price of the marijuana they would obtain. Second, exclusion of marijuana price from the extensive margin equation is due to marijuana being very cheap in Colombia (US\cent 83), consequently, price is not a barrier to define the extensive margin.

table[table omitted — 2,972 chars of source]
comment\begin{table}[h!] \caption{Exclusionary restrictions validation: Marijuana demand in Colombia.} \begin{threeparttable} \resizebox{0.65\textwidth}{!}{\begin{minipage}{\textwidth} \begin{tabular}{l c c}\hline \multirow{3}{*}{Variable} & \multicolumn{2}{c}{Margin} \\ \cline{2-3} & Extensive & Intensive \\ & (1) & (2) \\ \hline Exclusionary restrictions & & \\ \multirow{2}{*}{\quad Drug dealer in neighborhood} & 0.064 & 0.091 \\ & (0.041) & (0.133) \\ \multirow{2}{*}{\quad Alcohol and cigarette use} & 0.839 & 0.179 \\ & (0.046) & (0.097) \\ Age & & \\ \multirow{2}{*}{\quad 30s} & -0.338 & 0.118 \\ & (0.049) & (1.404) \\ \multirow{2}{*}{\quad 40s} & -0.498 & 0.437 \\ & (0.064) & (2.129) \\ \multirow{2}{*}{\quad 50s and older} & -0.965 & -3.881 \\ & (0.070) & (2.511) \\ Strata & & \\ \multirow{2}{*}{\quad Medium} & 0.097 & -0.024 \\ & (0.044) & (0.104) \\ \multirow{2}{*}{\quad High} & 0.069 & -0.076 \\ & (0.066) & (0.147) \\ Risk perception marijuana use & & \\ \multirow{2}{*}{\quad Medium} & -0.174 & -0.651 \\ & (0.091) & (0.178) \\ \multirow{2}{*}{\quad High} & \textbf{-0.815} & \textbf{-0.708} \\ & (0.083) & (0.170) \\ \multirow{2}{*}{Years of education} & \textbf{-0.015} & \textbf{-0.090} \\ & (0.006) & (0.014) \\ \multirow{2}{*}{Female} & \textbf{-0.418} & -0.496 \\ & (0.043) & (0.109) \\ \multirow{2}{*}{Good mental health} & \textbf{-0.197} & 0.138 \\ & (0.043) & (0.104) \\ \multirow{2}{*}{Marijuana users in network} & \textbf{0.951} & 0.255 \\ & (0.064) & (0.210) \\ \multirow{2}{*}{Worker} & \textbf{-0.100} & 0.176 \\ & (0.043) & (0.096) \\ Price & & \\ \multirow{2}{*}{\quad $\log\left\{\text{price of marijuana}\right\}$} & & \textbf{-0.512} \\ & & (0.100) \\ \multirow{2}{*}{\quad Age 30s $\times \log\left\{\text{price of marijuana}\right\}$} & & 0.015 \\ & & (0.182) \\ \multirow{2}{*}{\quad Age 40s $\times \log\left\{\text{price of marijuana}\right\}$} & & -0.023 \\ & & (0.276) \\ \multirow{2}{*}{\quad Age 50s and older $\times \log\left\{\text{price of marijuana}\right\}$} & & 0.480 \\ & & (0.317) \\ Constant & \textbf{-1.55} & 6.427 \\ & (0.198) & (0.949) \\ Regional-fixed effects & $\checkmark$ & $\checkmark$ \\ \hline Sample size & 12,994 & 1,043 \\ \hline \end{tabular} \begin{tablenotes}[para,flushleft] \textit{Notes}: Bold font indicates statistically significant variables. This table shows the posterior means and standard deviation (in parenthesis) of parameters of the extensive margin (column 1) and intensive margin (column 2) using the subset of individuals that were offered marijuana. These results suggest plausibility of the exclusionary restrictions. \\ \textit{Source}: Authors' construction. \end{tablenotes} \end{minipage}} \end{threeparttable} \end{table}
commentTo study the potential effects of no achievement of the exclusionary restrictions we perform a simulation exercise. Particularly, we simulate the model given by equations (ref), (ref) and (ref) to check the sampling properties of our Bayesian inferential framework. In particular, we set $\bm{x}_{im}^a=\left[1, w_{i2}\right]^{\top}$, $w_{i2}\stackrel{iid}{\sim}N(0,1)$, $\bm{x}_{im}^c=\left[1, z_{i2},z_{i3}\right]^{\top}$, $z_{i2}\stackrel{iid}{\sim}N(0,1)$, $z_{i3}\stackrel{iid}{\sim}N(0,1)$, and $\bm{x}_{im}^y=\left[1, x_{i2},x_{i3}\right]^{\top}$, $x_{i2}\stackrel{iid}{\sim}N(0,1)$, and $x_{i3}\stackrel{iid}{\sim}N(0,1)$. We replicate our exercises 100 times, using a sample size equal to 2,500. We use non-informative priors, $\boldsymbol{\theta}_0=\bm{0}$ and $\boldsymbol{\Theta}_0=\text{diag}\left\{1,000\right\}$, $\bm{R}_0=\bm{I}$ and $r_0=3$ with 1,100 iterations in our Gibbs sampler a burn-in equal to 105 and a thinning parameter equal to 5. We use the posterior means and the 95% credible intervals to check the sampling properties of our proposal. Both criteria minimize quadratic loss functions in parameters and measures of credibility, respectively ramirez2019focused. Tables (ref) and (ref) report the population values, average of posterior means, measures of errors of the posterior means as estimators of the population parameters, root mean square error (RMSE) and mean absolute percentage error (MAPE), coverage and width of the 95% credible intervals. Tables (ref) and (ref) show that the average of the posterior means are equal to population parameters up to two decimals in location and scale parameters, average errors are lower than 4%, on average around 2%, the coverage is approximately 95% on average, and credible intervals are relatively narrow.

Robustness checks

We perform several exercises to check the robustness of the posterior estimates of our baseline specification. First, as the definition of access is very relevant in our analysis, we check our results with another definition of access. In particular, our baseline definition of access is equal to one if an individual responds that is easy to get marijuana, and zero in case that responds that it is difficult, impossible or does not know how to get it. This implies that 58% of our sample has access to marijuana. In this exercise, we relax the access definition by including within the individuals who have access those who report that it would be easy or difficult to get marijuana. This new definition of access implies that 67% of individuals in the sample have access (see Table (ref) in the Appendix). We reran our baseline model using this new definition of the accessibility, the results can be seen in the second column of tables (ref), (ref) and (ref), where we show the posteriors estimates for the access, extensive and intensive equations, respectively. In general, we obtain qualitatively similar results in the alternative scenario relaxing the access definition (column 2) compared to the baseline estimates (column 1) in the three stages. In most of the posterior estimates there are not statistically significant differences in both exercises, except that individuals in high strata do not have statistically significant differences compared with individuals in low strata regarding access in this new set of estimates (see Table (ref)).

We also estimate our baseline model using the subset of uni-personal households. This is because lying is a valid concern when modeling demand for illicit drugs due to, for instance, social stigma lloyd2013stigmatization. We guess that individuals who live alone have less incentives to lie regarding marijuana use. Column (3) in tables (ref), (ref) and (ref) show the results using this sub-sample. We observe again that there are not statistically significant differences compared to the baseline exercise in the access and extensive margin, except that now there are not statistically significant differences regarding strata or being a worker. In addition, years of education is not statistically relevant in the extensive margin. However, we observe some intriguing results in the intensive margin estimates. In particular, three very robust regressors in all the estimations are not statistically significant in this subset: female, marijuana users in the network and price. Although their coefficients have the expected sign. We suspect that can be a power issue due to being just 233 marijuana consumers in this sub-sample.

We also include interaction effects between age splines and risk perception about marijuana consumption in our main specification. We perform this to identify potential heterogeneous effects in this variable among age groups, thus thinking about marketing campaigns targeting young adults to curve marijuana consumption through risk perception. We observe in column (4) of Table (ref) that these variables are not statistically significant, and in general, we get very similar results in this alternative specification compared to the baseline exercise.

We get our price measure in the baseline estimation calculating expenditure in marijuana over quantity. The advantage of this measure is that takes implicitly quality into account when an individual buys different types of marijuana. However, the survey asks directly individuals about price, there is the question “Do you know how much a marijuana cigarette or joint costs? We estimate our model using this alternative measure of price, which implies calculating again the quantity weighted by THC. Column (5) in tables (ref), (ref) and (ref) show the results. It seems that our results are robust to the price measure, the price elasticity is numerically lower under the alternative price compared to the baseline exercise, but there is not statistically significant differences.

A potential endogeneity issue that we did not take into account in our main specification is self-perception about health status. This variable may be considered endogenous jacobi2016marijuana, thus we estimate the baseline specification without the mental and physical self-perception of health status. The results can be seen in column (6) of tables (ref), (ref) and (ref). The results are very similar comparing this exercise with the baseline specification, except that medium strata do not have statistically significant differences with the low strata in the extensive margin in this new setting.

table[table omitted — 4,389 chars of source]
table[table omitted — 4,145 chars of source]
table[table omitted — 4,984 chars of source]
comment\begin{table}[h!] \caption{Robustness checks: Posterior results of scale parameters of marijuana demand in Colombia.} \begin{threeparttable} \resizebox{0.7\textwidth}{!}{\begin{minipage}{\textwidth} \begin{tabular}{l c c c c}\hline \multirow{2}{*}{Posterior measures} & \multicolumn{4}{c}{Baseline specification} \\ \cline{2-5} & $\sigma_{ca}$ & $\sigma_{ya}$ & $\sigma_{yc}$ & $\sigma_{y}^2$ \\ \hline Mean & 0.0014 & 7.2348 & -0.0919 & 2.8564 \\ Standard deviation & (0.0101) & (0.5490) & (0.1111) & (0.1308) \\ \hline & \multicolumn{4}{c}{Access definition} \\ \hline Mean \\ Standard deviation \\ \hline & \multicolumn{4}{c}{Uni-personal household} \\ \hline Mean & -0.0016 & -2.1369 & 0.1530 & 2.4008 \\ Standard deviation & (0.0306) & (6.5156) & (0.5495) & (0.3275) \\ \hline & \multicolumn{4}{c}{Interaction: Age and risk} \\ \hline Mean & & & & \\ Standard deviation & & & & \\ \hline & \multicolumn{4}{c}{Price definition} \\ \hline Mean & & & & \\ Standard deviation & & & & \\ \hline & \multicolumn{4}{c}{Perception: Health status} \\ \hline Mean & & & & \\ Standard deviation & & & & \\ \hline & \multicolumn{4}{c}{Zero consumption} \\ \hline Mean & & & & \\ Standart deviation & & & & \\ \hline \end{tabular} \begin{tablenotes}[para,flushleft] Notes: Bold font indicates statistically significant variables. Columns labeled baseline specification show posterior estimates of the identified covariance matrix of the baseline model. \\ Source: Authors' construction. \end{tablenotes} \end{minipage}} \end{threeparttable} \end{table} Finally, we replace the intensive margin equation given by equation (ref), by a setting where we keep zeros for individuals who report not to consume marijuana, \begin{align} Y_{im}=\begin{Bmatrix} \bm{w}^{\top}_i \boldsymbol{\alpha}_y + \omega_y + p_i\gamma_y + \sum_{j=1}^J w_i^{Age(j)}\times p_{i}\gamma_{yj} + W_{im}, & \begin{Bmatrix} C_{im}=1\\ C_{im}=0 \end{Bmatrix}\\ - & A_{im}=0 \end{Bmatrix}. \end{align} We should take into account that when using equation (ref), there are just two groups in the estimation process: all individuals ($G_1$), and individuals having access ($G_2$). The former contribute to estimate parameters in equation (ref), and the latter contribute to estimate parameters in equations (ref), (ref) and (ref). We have the issue of logarithm of zeros in this specification. Given that approximately 50% of the papers published in American Economic Review between 2016 and 2020 facing this problem opted to add a positive discretionary value to the dependent variable bellego2022dealing, we follow this approach setting $\log\left\{1+Y_{im}\right\}$ in this specification. Column (7) in tables (ref), (ref) and (ref) show these results. We do not see statistically significant differences in the access equation (see Table (ref)). On the other hand, previous consumption of alcohol and cigarette, medium strata, and being a worker are not statistically significant under this latter specification (see Table (ref)),

In general, it seems that the results of the baseline specification are robust, there are not statistically significant differences in most of the cases compared to the alternative measures of relevant variables or model specifications. We observe that the same variables are statistically significant in all three stages of demand for marijuana, and the numeric values of the posterior estimates are relatively similar. However, the sub-sample of uni-personal households present some intriguing results, potentially due to power issues.

Policy analysis

Marijuana legalization

We perform some counterfactual experiments to estimate the potential effects of the legalization of marijuana in Colombia for different representative individuals. Particularly, we estimate the posterior predictive probability for individual 0,

align*[align* omitted — 511 chars of source]

where $\bm{S}$ is the support of integration of $\boldsymbol{\Sigma}$, $p(A_0=1|\bm{T},\bm{X})=1-\Phi(\bm{x}_0^{a\top}\boldsymbol{\theta}^a,1)$, $p(C_0=1|A_0=1,\bm{T},\bm{X})=1-\Phi(\bm{x}_0^{c\top}\boldsymbol{\theta}^c+\sigma_{ac}(U_0^a-\bm{x}_0^{a\top}\boldsymbol{\theta}^a),1-\sigma_{ac}^2)$ and $p(Y_0|C_0=1,A_0=1,\bm{T},\bm{X})\sim N(\mu_{y|ac},\sigma^2_{y|ac})$, where

align*[align* omitted — 299 chars of source]

and

align*[align* omitted — 211 chars of source]

Observe the relevance of the selection parameters, $\sigma_{ac}$, $\sigma_{ya}$ and $\sigma_{yc}$ in the previous expressions. These account for unobserved dependence between the three stages of the demand for marijuana.

The above integral can be estimated in a straight forward way using the draws from the posterior distribution. Therefore, we use simulation to estimate the effects of the legalization of marijuana on the probability of use, given that under legalization the probability of access is equal to 1, that is, $p(A_0=1|\bm{T},\bm{X})=1$, and the amount of consumption, conditional on use, where we take into account that we model $\log(Y_{it})$, so we get by simulation $Y_{it}$, that is, the amount of joints per month.

We show in tables (ref) and (ref) the results of these exercises for the representative individual who has access to marijuana. In particular, this is an individual with 12 years of education, working, good self-perception of health status, family members and friends who do not consume marijuana, but she/he has consumed alcohol and cigarettes, lives in Medellín in a low socioeconomic stratum, and there is a drug dealer in the neighborhood. We have in these tables seven scenarios, rows one to three in each panel show results under different scenarios about risk perceptions of marijuana use, these experiments are based on the fact that marketing campaigns warning about bad consequences of consuming some products may curve demand. For instance, warning labeling explains why consumers become more health-conscious, and consequently, more risk-averse barahona2023, Berg2023, cannoy2023, Kaai2023,Nguyen2023,Nian2023,brennan2022. Rows four to seven show results under different price scenarios, the ones that we analyze in the next subsection for potential tax revenues, taxes contribute to raise revenues for public policy, and help to curve the demand function allcott2019.

We can see in Table (ref) the results for the representative woman, where each panel shows results by age spline. For instance, the first row in the first panel shows the results under a the baseline price for this representative woman (US\cent 78.2), who has a high risk perception about marijuana use. We observe that the predicted probability of having access to marijuana is 74.5%, and the predicted probabilities of marijuana use are 1.30% and 1.75%, overall women with these features, and those with access, respectively. This means that the probability of use overall these representative women increases 0.44 percentage points given a policy of legalization of marijuana. Conditional on access and use, the predicted consumption for this representative woman is 5.9 joints per month. All these estimates has the standard errors that are calculated by simulation using repeated sampling.

We observe from Table (ref) that under legalization of marijuana, risk perception has a higher effect among women in decreasing the probability of use than price. For instance, the second and third rows show that given access, the probability of use increases 9.37 and 4.58 percentage points for women who have a medium and low risk perceptions about marijuana use, compared with just 0.44 p.p. for women with high risk perception. However, the effect of risk perception decreases with age. We observe the same pattern among men (see Table (ref)). Overall, taking into account these three risk scenarios, age splines and gender, and using the expansion factors of the survey, we find that the probability of use increases from 2.3% pre-legalization to 3.0% under legalization.

Given the effect of risk perception, and the patterns of this by age splines, we can deduce that marketing campaigns targeting young individuals will be an effective way to curve demand for marijuana under a legal setting. First, young have a lower risk perception about marijuana use than other age splines, the sample average for a high risk is 88% for the former, whereas this percentage is equal to 92%, 93% and 94% for 30s, 40s and 50s age splines, this means that there is a higher gap among young individuals. And second, reducing more the demand of this group implies that the cumulative effect on consumption through their life span is higher, with potentially more good externalities. However, we should take into account that the effect of risk perception has a limit, this is, on average 91% of the individuals already have a high risk perception about marijuana use. This fact motivates to perform experiments with different prices (taxes), which also means different scenarios regarding government revenues.

Thus, the second set of experiments consider the effect of legalization of marijuana on price, and as a consequence, on access, extensive and intensive margins. Particularly, marijuana legalization implies that the inherent extra cost due to illegality would disappear. However, we should take a potential tax into account. A first benchmark is the average cost of marijuana in Colombia, velez2021medicinal found that the average production cost is US\cent 1 per gram, and taking into account that the average percentage of distribution cost in Colombia is 15%, the cost of one joint of marijuana is approximately US\cent 1.15. In addition, the average return of capital in Colombia is around 15.25% according to Corficolombiana, a prestigious financial institution in this country, fluctuating between 13.9% y 16.6%,\footnote{See \href{chrome-extension://efaidnbmnnnibpcajpcglclefindmkaj/https://investigaciones.corficolombiana.com/documents}{Rentabilidad esperada del capital propio.}} thus a gram of marijuana has a base price without taxes of approximately US\cent 1.33. The first counterfactual exercise uses as reference the tax on cigarettes, which is US\cent 5.9 in Colombia,\footnote{Real price 2019, see \href{https://actualicese.com/certificacion-04-del-07-12-2021/}{CERTIFICACIÓN 4, 2021 of Dirección General de Apoyo Fiscal del Ministerio de Hacienda.}} then the potential price of one gram of marijuana, tax included, is approximately US\cent 7.3. This price can be considered as a potential lower bound due to being less expensive that the legal price of cigarettes in Colombia, which is on average US\cent 11.5. Although, this price is higher than the price of illegal cigarettes, US\cent 5.3.\footnote{Real price 2019, see \href{https://www.semana.com/salud/articulo/consumo-de-cigarrillos-34-de-cada-100-fueron-de-contrabando-esto-revelo-la-federacion-nacional-de-departamentos/202315/}{Estudio de incidencia del consumo de cigarrillos en colombia 2022.}} The second exercise assumes that the price of marijuana is equal to the legal price of cigarettes (US\cent 11.5). Finally, the potential lower bound and the actual price of marijuana offer a spectrum of possibilities for tax scenarios. Thus, we perform two more experiments, 50% decrease and 25% increase with respect to the actual price of marijuana. The former implies a tax of US\cent 37.8 per joint, which is more than 6 times the tax on cigarettes. The latter is based on jacobi2016marijuana, who propose a 25% tax on the actual price of marijuana in Australia. In the Colombian case, this tax would be more than 16 times the tax of cigarettes. We think about the this scenario as a potential upper bound price in the case of legalization of marijuana due to a relatively high price may imply a huge black market of marijuana in this country. For instance, the size of the black market of cigarettes in Colombia is 34%.\footnote{See \href{https://www.semana.com/salud/articulo/consumo-de-cigarrillos-34-de-cada-100-fueron-de-contrabando-esto-revelo-la-federacion-nacional-de-departamentos/202315/}{Estudio de incidencia del consumo de cigarrillos en colombia 2022.}}

The fourth to seventh rows in each panel of tables (ref) and (ref) show the results. We observe that there are not remarkable differences regarding the probability of access and use under different price settings; however, the intensity of consumption decreases with price, as expected. The shape of the intensity margin as a function of age splines has an “inverted U-shape", that is, this is low for 20s and 50s, increases in 30s, and has a peak in 40s. This latter group has a very high level of consumption, however, we should take with caution this result due to the also high volatility level. Observe that price helps to curve the intensity of consumption under a legalization policy; however, there are not significant changes in the extensive margin due to different price regimes.

Other patterns that we observe from tables (ref) and (ref) are that access is lower for individuals who have a low risk perception regarding marijuana use; however, the probability of marijuana use is higher for this group. In addition, access also decreases with age, and is higher for men, who in turn have a higher probability of use, and given use, have a substantially higher level of consumption, approximately two times the level of women.

We perform ceteris paribus exercises in order to isolate the effects of different control variables, and get a better understanding of the situation. However, manthey2023 demonstrated that warning information about product use that potentially curve consumption must be complemented with taxes (pricing). Both approaches should be integrated as part of a comprehensive strategy aimed at mitigating the burdens on public health, social well-being, and economy, of marijuana use.

sidewaystable\caption{Counterfactual experiments: Effects of legalization of marijuana in Colombia for a representative woman.} \begin{threeparttable} \resizebox{1\textwidth}{!}{\begin{minipage}{\textwidth} \begin{tabular}{l c c c c c c c c c c c c c c c c}\hline & & & & \multicolumn{2}{c}{Pred. Prob. access} & & \multicolumn{7}{c}{Pred. prob use} & & \multicolumn{2}{c}{Pred. consumption} \\ \cline{5-6} \cline{8-14} \cline{16-17} \multicolumn{3}{c}{Scenario} & & \multicolumn{2}{c}{All} & & \multicolumn{2}{c}{All} & & \multicolumn{2}{c}{Access} & & Change & & \multicolumn{2}{c}{Consumption} \\ \cline{1-3} \cline{5-6} \cline{8-9} \cline{11-12} \cline{14-14} \cline{16-17} Price scenario & Price & Risk perception & & Mean & Std. Error & & Mean & Std. Error & & Mean & Std. Error & & Mean & & Mean & Std. Error \\ \hline \multicolumn{16}{c}{Woman in 20s or younger} \\ \hline Baseline & US\cent 78.2 & High & & 74.50% & 0.0137 & & 1.30% & 0.0036 & & 1.74% & 0.0048 & & 0.44 p.p. & & 5.90 & 0.30 \\ Baseline & US\cent 78.2 & Medium & & 77.70% & 0.0131 & & 8.30% & 0.0087 & & 10.68% & 0.0110 & & 9.37 p.p. & & 6.58 & 0.33 \\ Baseline & US\cent 78.2 & Low & & 66.50% & 0.0149 & & 9.10% & 0.0091 & & 13.68% & 0.0130 & & 4.58 p.p. & & 9.06 & 0.50 \\ Lower bound & US\cent 7.3 & High & & 70.80% & 0.0143 & & 2.30% & 0.0047 & & 3.24% & 0.0067 & & 0.94 p.p. & & 17.44 & 0.87 \\ Cigarette & US\cent 11.5 & High & & 72.70% & 0.0142 & & 1.50% & 0.0038 & & 2.08% & 0.0053 & & 0.58 p.p. & & 13.74 & 0.64 \\ 50% decrease & US\cent 39.1 & High & & 72.60% & 0.0141 & & 1.60% & 0.0040 & & 2.20% & 0.0054 & & 0.60 p.p. & & 8.22 & 0.39 \\ 25% increase & US\cent 97.8 & High & & 72.00% & 0.0141 & & 2.60% & 0.0050 & & 3.58% & 0.0069 & & 0.98 p.p. & & 5.18 & 0.31 \\ \hline \multicolumn{16}{c}{Woman in 30s or younger} \\ \hline Baseline & US\cent 78.2 & High & & 69.40% & 0.0146 & & 0.50% & 0.0022 & & 0.72% & 0.0032 & & 0.22 p.p. & & 22.20 & 2.64 \\ Baseline & US\cent 78.2 & Medium & & 76.70% & 0.0133 & & 2.50% & 0.0049 & & 3.25% & 0.0064 & & 0.75 p.p. & & 20.80 & 1.65 \\ Baseline & US\cent 78.2 & Low & & 62.80% & 0.0152 & & 3.80% & 0.0061 & & 6.05% & 0.0095 & & 2.25 p.p. & & 30.12 & 2.80 \\ Lower bound & US\cent 7.3 & High & & 70.50% & 0.0144 & & 0.01% & 0.0030 & & 1.28% & 0.1123 & & 1.27 p.p. & & 48.21 & 3.87 \\ Cigarette & US\cent 11.5 & High & & 67.50% & 0.0148 & & 0.05% & 0.0022 & & 0.74% & 0.0033 & & 0.69 p.p. & & 47.54 & 7.40 \\ 50% decrease & US\cent 39.1 & High & & 71.80% & 0.0142 & & 0.40% & 0.0020 & & 0.55% & 0.0028 & & 0.15 p.p. & & 23.29 & 1.87 \\ 25% increase & US\cent 97.8 & High & & 65.50% & 0.0150 & & 0.40% & 0.0019 & & 0.61% & 0.0030 & & 0.21 p.p. & & 17.72 & 1.62 \\ \hline \multicolumn{16}{c}{Woman in 40s or younger} \\ \hline Baseline & US\cent 78.2 & High & & 67.10% & 0.0150 & & 0.30% & 0.0013 & & 0.44% & 0.0025 & & 0.14 p.p. & & 57.68 & 10.86 \\ Baseline & US\cent 78.2 & Medium & & 73.00% & 0.0140 & & 2.00% & 0.0044 & & 2.74% & 0.0060 & & 0.74 p.p. & & 69.14 & 15.44 \\ Baseline & US\cent 78.2 & Low & & 60.80% & 0.0154 & & 1.90% & 0.0043 & & 3.12% & 0.0070 & & 1.24 p.p. & & 133.25 & 34.22 \\ Lower bound & US\cent 7.3 & High & & 66.30% & 0.0150 & & 0.30% & 0.0017 & & 0.45% & 0.0026 & & 0.15 p.p. & & 313.13 & 110.55 \\ Cigarette & US\cent 11.5 & High & & 62.20% & 0.0153 & & 0.60% & 0.0024 & & 0.96% & 0.0039 & & 0.36 p.p. & & 172.62 & 40.93 \\ 50% decrease & US\cent 39.1 & High & & 64.30% & 0.0151 & & 0.20% & 0.0014 & & 0.31% & 0.0022 & & 0.08 p.p. & & 76.03 & 16.98 \\ 25% increase & US\cent 97.8 & High & & 68.90% & 0.0146 & & 0.20% & 0.0014 & & 0.29% & 0.0021 & & 0.09 p.p. & & 63.72 & 11.05 \\ \hline \multicolumn{16}{c}{Woman in 50s or younger} \\ \hline Baseline & US\cent 78.2 & High & & 64.00% & 0.0152 & & 0.10% & 0.0010 & & 0.16% & 0.0015 & & 0.06 p.p. & & 1.52 & 0.17 \\ Baseline & US\cent 78.2 & Medium & & 68.50% & 0.0147 & & 0.90% & 0.0020 & & 1.31% & 0.0043 & & 0.41 p.p. & & 1.72 & 0.22 \\ Baseline & US\cent 78.2 & Low & & 56.70% & 0.0157 & & 0.40% & 0.0020 & & 0.70% & 0.0035 & & 0.30 p.p. & & 2.47 & 0.32 \\ Lower bound & US\cent 7.3 & High & & 59.90% & 0.0155 & & 0.00% & 0.0000 & & 0.00% & 0.0000 & & 0.00 p.p. & & 5.40$^*$ & 1.27 \\ Cigarette & US\cent 11.5 & High & & 60.30% & 0.0155 & & 0.30% & 0.0017 & & 0.50% & 0.0029 & & 0.20 p.p. & & 6.18 & 1.64 \\ 50% decrease & US\cent 39.1 & High & & 61.90% & 0.0154 & & 0.00% & 0.0000 & & 0.00% & 0.0000 & & 0.00 p.p. & & 2.51$^*$ & 0.50 \\ 25% increase & US\cent 97.8 & High & & 60.70% & 0.0155 & & 0.20% & 0.0014 & & 0.33% & 0.0023 & & 0.13 p.p. & & 1.63 & 0.29 \\ \hline \end{tabular} \begin{tablenotes}[para,flushleft] Notes: $^*$ Up to four digits the probability of marijuana use is zero for this individual, but in the highly unlikely case of consumption, this is the expected amount. Predicted measures for the representative woman with access to marijuana. Her level of education is high school (12 years of education), works, has a good self-perception of mental and physical health, no marijuana users in her network, consumes alcohol and cigarettes, lives in Medellín in a low socioeconomic stratum, and there is a drug dealer in her neighborhood. Pred. stands for prediction, and prob. is probability.\\ Source: Authors' construction. \end{tablenotes} \end{minipage}} \end{threeparttable}
sidewaystable\caption{Counterfactual experiments: Effects of legalization of marijuana in Colombia for a representative man.} \begin{threeparttable} \resizebox{1\textwidth}{!}{\begin{minipage}{\textwidth} \begin{tabular}{l c c c c c c c c c c c c c c c c}\hline & & & & \multicolumn{2}{c}{Pred. Prob. access} & & \multicolumn{7}{c}{Pred. prob use} & & \multicolumn{2}{c}{Pred. consumption} \\ \cline{5-6} \cline{8-14} \cline{16-17} \multicolumn{3}{c}{Scenario} & & \multicolumn{2}{c}{All} & & \multicolumn{2}{c}{All} & & \multicolumn{2}{c}{Access} & & Change & & \multicolumn{2}{c}{Consumption} \\ \cline{1-3} \cline{5-6} \cline{8-9} \cline{11-12} \cline{14-14} \cline{16-17} Price scenario & Price & Risk perception & & Mean & Std. Error & & Mean & Std. Error & & Mean & Std. Error & & Mean & & Mean & Std. Error \\ \hline \multicolumn{16}{c}{Man in 20s or younger} \\ \hline Baseline & US\cent 78.2 & High & & 79.30% & 0.0128 & & 4.90% & 0.0068 & & 6.10% & 0.0085 & & 1.30 p.p. & & 10.45 & 0.55 \\ Baseline & US\cent 78.2 & Medium & & 83.70% & 0.0117 & & 17.80% & 0.0121 & & 21.26% & 0.0142 & & 3.46 p.p. & & 10.44 & 0.55 \\ Baseline & US\cent 78.2 & Low & & 71.40% & 0.0143 & & 15.30% & 0.0114 & & 21.43% & 0.0154 & & 6.13 p.p. & & 15.17 & 0.76 \\ Lower bound & US\cent 7.3 & High & & 80.40% & 0.0126 & & 5.60% & 0.0073 & & 6.96% & 0.0089 & & 1.36 p.p. & & 29.38 & 1.47 \\ Cigarette & US\cent 11.5 & High & & 77.50% & 0.0132 & & 5.20% & 0.007 & & 6.70% & 0.0090 & & 1.50 p.p. & & 25.29 & 1.41 \\ 50% decrease & US\cent 39.1 & High & & 72.20% & 0.0132 & & 4.80% & 0.0067 & & 6.21% & 0.0087 & & 0.60 p.p. & & 13.88 & 0.62 \\ 25% increase & US\cent 97.8 & High & & 79.60% & 0.0127 & & 4.40% & 0.0065 & & 5.53% & 0.0081 & & 1.13 p.p. & & 9.98 & 0.46 \\ \hline \multicolumn{16}{c}{Man in 30s or younger} \\ \hline Baseline & US\cent 78.2 & High & & 76.70% & 0.0133 & & 2.70% & 0.0051 & & 3.52% & 0.0066 & & 0.82 p.p. & & 34.46 & 3.09 \\ Baseline & US\cent 78.2 & Medium & & 82.30% & 0.0120 & & 9.50% & 0.0092 & & 11.54% & 0.0111 & & 2.04 p.p. & & 40.53 & 4.50 \\ Baseline & US\cent 78.2 & Low & & 71.70% & 0.0142 & & 6.80% & 0.0079 & & 9.48% & 0.0109 & & 2.68 p.p. & & 48.52 & 4.64 \\ Lower bound & US\cent 7.3 & High & & 76.40% & 0.0134 & & 2.20% & 0.0046 & & 2.88% & 0.0061 & & 0.68 p.p. & & 94.66 & 8.77 \\ Cigarette & US\cent 11.5 & High & & 77.90% & 0.0131 & & 2.30% & 0.0047 & & 2.95% & 0.0061 & & 0.65 p.p. & & 74.21 & 6.24 \\ 50% decrease & US\cent 39.1 & High & & 77.90% & 0.0131 & & 1.90% & 0.0043 & & 2.44% & 0.0055 & & 0.54 p.p. & & 44.01 & 3.63 \\ 25% increase & US\cent 97.8 & High & & 74.70% & 0.0137 & & 1.00% & 0.0031 & & 1.34% & 0.0042 & & 0.34 p.p. & & 33.54 & 4.44 \\ \hline \multicolumn{16}{c}{Man in 40s or younger} \\ \hline Baseline & US\cent 78.2 & High & & 75.20% & 0.0137 & & 0.70% & 0.0026 & & 0.93% & 0.0035 & & 0.23 p.p. & & 123.18 & 43.60 \\ Baseline & US\cent 78.2 & Medium & & 80.20% & 0.0126 & & 4.20% & 0.0063 & & 5.23% & 0.0078 & & 1.03 p.p. & & 106.25 & 26.12 \\ Baseline & US\cent 78.2 & Low & & 67.00% & 0.0148 & & 6.20% & 0.0076 & & 9.25% & 0.0112 & & 3.05 p.p. & & 197.54 & 65.80 \\ Lower bound & US\cent 7.3 & High & & 74.80% & 0.0137 & & 1.20% & 0.0034 & & 1.60% & 0.0045 & & 0.40 p.p. & & 405.79 & 128.76 \\ Cigarette & US\cent 11.5 & High & & 75.10% & 0.0136 & & 0.90% & 0.0030 & & 1.99% & 0.0039 & & 1.09 p.p. & & 266.60 & 67.18 \\ 50% decrease & US\cent 39.1 & High & & 73.60% & 0.0139 & & 1.10% & 0.0033 & & 1.49% & 0.0044 & & 0.39 p.p. & & 272.48 & 135.76 \\ 25% increase & US\cent 97.8 & High & & 72.20% & 0.0141 & & 1.70% & 0.0040 & & 2.35% & 0.0056 & & 0.65 p.p. & & 109.08 & 19.89 \\ \hline \multicolumn{16}{c}{Man in 50s or younger} \\ \hline Baseline & US\cent 78.2 & High & & 70.30% & 0.0144 & & 0.40% & 0.0019 & & 0.57% & 0.0028 & & 0.17 p.p. & & 3.13 & 0.43 \\ Baseline & US\cent 78.2 & Medium & & 73.80% & 0.0139 & & 1.60% & 0.0039 & & 2.17% & 0.0053 & & 0.57 p.p. & & 4.21 & 0.73 \\ Baseline & US\cent 78.2 & Low & & 61.00% & 0.0154 & & 1.70% & 0.0040 & & 2.79% & 0.0067 & & 1.09 p.p. & & 4.65 & 0.77 \\ Lower bound & US\cent 7.3 & High & & 70.80% & 0.0143 & & 0.50% & 0.0023 & & 0.70% & 0.0031 & & 0.20 p.p. & & 10.51 & 2.22 \\ Cigarette & US\cent 11.5 & High & & 70.10% & 0.0145 & & 0.60% & 0.0024 & & 0.85% & 0.0034 & & 0.15 p.p. & & 6.90 & 0.89 \\ 50% decrease & US\cent 39.1 & High & & 70.00% & 0.0150 & & 0.40% & 0.0020 & & 0.57% & 0.0029 & & 0.17 p.p. & & 3.96 & 0.57 \\ 25% increase & US\cent 97.8 & High & & 68.40% & 0.0147 & & 0.30% & 0.0017 & & 0.44% & 0.0025 & & 0.14 p.p. & & 2.87 & 0.43 \\ \hline \end{tabular} \begin{tablenotes}[para,flushleft] Notes: Predicted measures for the representative man with access to marijuana. His level of education is high school (12 years of education), works, has a good self-perception of mental and physical health, no marijuana users in his network, consumes alcohol and cigarettes, lives in Medellín in a low socioeconomic stratum, and there is a drug dealer in her neighborhood. Pred. stands for prediction, and prob. is probability.\\ Source: Authors' construction. \end{tablenotes} \end{minipage}} \end{threeparttable}

Tax revenues

We perform some simulation exercises regarding the potential tax revenue that the government could collect from a tax on marijuana consumption under legalization. In particular, we use the posterior draws to simulate the model for all the individuals in the survey using the predictive framework of the previous subsection, assuming that the probability of access is equal to one for every one under a legal framework. We estimate the probability of consumption given access for each individual, and use it to sample from a Bernoulli distribution, if the realization of this experiment is 0, then the associated consumption is 0, if the realization is equal to 1, then, we predict their consumption conditional on access and use. Then, we use the expansion factors of the survey to estimate the potential yearly tax revenues. The survey represents 23.6 million individuals between 12 and 65 years-old in 2019, this is approximately 75% of the total Colombian population in this age range. Thus, we should consider these predictions as underestimating the potential revenue. Although, the missing 25% of the population is located in relatively isolated rural areas where potentially would have not legal marijuana suppliers. We also take into account that approximately 34% of the demand of marijuana under a legal framework would be in the black market. This figure is based on the situation in the cigarettes market.\footnote{See \href{https://www.semana.com/salud/articulo/consumo-de-cigarrillos-34-de-cada-100-fueron-de-contrabando-esto-revelo-la-federacion-nacional-de-departamentos/202315/}{Estudio de incidencia del consumo de cigarrillos en colombia 2022.}}

We set fourth potential scenarios where all of them assume that the difference between the average cost per gram (US\cent 1.33) and the price is equal to the tax. The first is a lower bound where the price is equal to the cost (US\cent 1.33) plus a tax that is equal to the cigarette tax (US\cent 5.9). The second scenario assumes that the price of marijuana is equal to the average price of a legal cigarette, the third assumes a price that is the 50% of approximately the actual average price that a representative individual with access pays for a joint, and the fourth scenario uses a price that is 25% more expensive than the latter.

We can see in Table (ref) the results. We observe that the annually average tax revenue under legalization of marijuana in Colombia would be between USD 11.0 million and USD 54.2 million, depending on which taxation scheme is used. This is between 2.9% to 14.4% of the tax revenues of cigarettes, according to the Ministry of Finance and Public Credit, that reports tax revenues of cigarettes are equal to USD 374 million in 2019.\footnote{See \href{https://www.minhacienda.gov.co/webcenter/portal/Estadisticas}{Ministerio de hacienda y crédito público}} We should take into account that in Colombia, the amount of cigarettes per month of a smoker is 208, whereas the predicted average consumption of marijuana under the counterfactual of cigarette price is 57 per month, almost four times less. In addition, the predicted probability of use of marijuana is 2.5% in this price setting under legalization, whereas the probability of cigarette is 9.5% (reporting smoking in the last month), according to the ENCSPA survey, that is, approximately 4 times higher.

table[table omitted — 1,175 chars of source]

Concluding remarks

We present an endogenous three-part model to estimate the demand of marijuana in Colombia that allows to infer the potential effects of its legalization finding heterogeneous effects among age groups and gender. Thus, we extend jacobi2016marijuana's proposal modeling simultaneously the three stages of marijuana demand (access, extensive and intensive margins), taking truncation into account.

The main estimation findings indicate that women have a lower probability of access, use, and quantity of consumption than men. Individuals over 30 also have a lower probability of access and use than younger individuals (20s and below), and individuals with good mental health also have a lower probability of access and use. Overall, we find that the demand for marijuana exhibits inelasticity (-0.45); moreover, there is no statistically significant difference in this elasticity in prices across age groups. These results are robust to different specifications, and access, price and intensive margin definitions.

We also find that a legalization policy would increase the probability of use from 2.3% to 3.0%, particularly affecting young individuals, from 4.3% to 5.5%, where risk perception is a relevant driver to curve marijuana demand. Therefore, under a legalization policy, marketing warning the potential bad consequences of consuming marijuana targeting younger individuals would be mandatory. Therefore, part of the potential tax revenues, which under a realistic setting would be approximately USD 32 million, should be invested in these warning campaigns. In addition, we have that given the relatively low price of marijuana in Colombia, this variable can be use to drive the intensive margin, rather than the extensive margin. In any case, warning campaigns should be complemented with taxes as a comprehensive strategy to mitigate negative effects associated with marijuana consumption.

We consider in this study the short-term effects of a legalization policy on the extensive and intensive margins of marijuana demand, and potential tax revenues from this activity. However, future research should consider long-term effects of this policy. In this point, the “gateway hypothesis" is relevant as a legal framework for marijuana consumption may imply consequences on consumption of more toxic drugs due to the polydrug use. In addition, other aspects of a policy of legalization should be considered, for instance, effects on public health, labor and crime. The latter is particularly relevant in Colombia due to the relevance of marijuana in the micro traffic business.

comment\section*{Tables} \begin{table}[ht] \caption{Sampling properties of location parameters estimators: Three-part incidental truncation model.} \begin{threeparttable} \resizebox{0.95\textwidth}{!}{\begin{minipage}{\textwidth} \begin{tabular}{c c c c c c c}\hline Parameter & Population & $\hat{\hat{\theta}}$ & RMSE & MAPE & 95% CI coverage & 95% CI width \\ & & & & & Coverage & Width \\ \hline $\theta_{1}^a$ & 1.00 & 1.00 & 0.03 & 0.02 & 0.91 & 0.09 \\ $\theta_{2}^a$ & -1.00 & -1.00 & 0.02 & 0.02 & 0.91 & 0.09 \\ $\theta_{1}^c$ & 1.00 & 1.00 & 0.03 & 0.02 & 0.95 & 0.12 \\ $\theta_{2}^c$ & 1.00 & 1.00 & 0.03 & 0.02 & 0.93 & 0.09 \\ $\theta_{3}^c$ & -0.50 & -0.50 & 0.02 & 0.02 & 0.97 & 0.07 \\ $\theta_{1}^y$ & 1.00 & 1.00 & 0.03 & 0.02 & 0.99 & 0.10 \\ $\theta_{2}^y$ & 1.70 & 1.70 & 0.02 & 0.01 & 0.92 & 0.07 \\ $\theta_{3}^y$ & 1.50 & 1.50 & 0.02 & 0.01 & 0.94 & 0.07 \\ \addlinespace[.75ex] \hline \end{tabular} \begin{tablenotes}[para,flushleft] Notes: $\hat{\hat{\theta}}= \frac{1}{100}\frac{1}{200}\sum_{r=1}^{100}\sum_{s=1}^{200}\theta_r^{(s)}$, the root of mean square errors is $\sqrt{\frac{1}{100}\sum_{r=1}^{100}(\theta-\hat{\theta}_r)^2}$, the mean absolute percentage error is $\frac{1}{100}\sum_{r=1}^{100}\left|\frac{\theta-\hat{\theta}_r}{\theta}\right|$, coverage of 95% credible interval is equal to $\frac{1}{100}\sum_{r=1}^{100}\mathbbm{1}{\left[\theta\in(\hat{\theta}_{2.5\%,r},\hat{\theta}_{97.5\%,r})\right]}$, and interval width of 95% credible intervals is $\frac{1}{100}\sum_{r=1}^{100}\left|\hat{\theta}_{97.5\%,r}-\hat{\theta}_{2.5\%,r}\right|$, where $\hat{\theta}_r$, $\hat{\theta}_{2.5\%,r}$ and $\hat{\theta}_{97.5\%,r}$ are the posterior mean, 2.5% and 97.5% percentiles based on the posterior draws of replication $r$-th, and $\mathbbm{1}{\left[\cdot\right]}$ is the indicator function.\\ Source: Authors' construction. \end{tablenotes} \end{minipage}} \end{threeparttable} \end{table} \begin{table}[h!] \caption{Sampling properties of scale parameters estimator: Three-part incidental truncation model.} \begin{threeparttable} \resizebox{0.7\textwidth}{!}{\begin{minipage}{\textwidth} \begin{tabular}{c c c c c c}\hline Parameter & Population & $\hat{\hat{\sigma}}$ & RMSE & \multicolumn{2}{c}{95% credible interval} \\ & & & & Coverage & Width \\ \hline $\sigma_{ca}$ & 0.60 & 0.60 & 0.02 & 0.92 & 0.06 \\ $\sigma_{ya}$ & 0.40 & 0.40 & 0.03 & 0.95 & 0.11 \\ $\sigma_{yc}$ & 0.70 & 0.70 & 0.03 & 0.95 & 0.12 \\ $\sigma^{2}_{y}$ & 1.00 & 1.01 & 0.04 & 0.95 & 0.15 \\ \hline \end{tabular} \begin{tablenotes}[para,flushleft] Notes: $\hat{\hat{\sigma}}=\frac{1}{100}\frac{1}{200}\sum_{r=1}^{100}\sum_{s=1}^{200}\sigma_r^{(s)}$, the root of mean square errors is $\sqrt{\frac{1}{100}\sum_{r=1}^{100}(\sigma-\hat{\sigma}_r)^2}$, the mean absolute percentage error is $\frac{1}{100}\sum_{r=1}^{100}\left|\frac{\sigma-\hat{\sigma}_r}{\sigma}\right|$, coverage of 95% credible interval is equal to $\frac{1}{100}\sum_{r=1}^{100}\mathbbm{1}{\left[\sigma\in(\hat{\sigma}_{2.5\%,r},\hat{\sigma}_{97.5\%,r})\right]}$, and interval width of 95% credible intervals is $\frac{1}{100}\sum_{r=1}^{100}\left|\hat{\sigma}_{97.5\%,r}-\hat{\sigma}_{2.5\%,r}\right|$, where $\hat{\sigma}_r$, $\hat{\sigma}_{2.5\%,r}$ and $\hat{\sigma}_{97.5\%,r}$ are the posterior mean, 2.5% and 97.5% percentiles based on the posterior draws of replication $r$-th, and $\mathbbm{1}{\left[\cdot\right]}$ is the indicator function.\\ Source: Authors' construction. \end{tablenotes} \end{minipage}} \end{threeparttable} \end{table}