EconBase
← Back to paper

EASI Drugs in the Streets of Colombia: Modeling Heterogeneous and Endogenous Drug Preferences

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

113,830 characters · 14 sections · 73 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

EASI Drugs in the Streets of Colombia: Heterogeneous Drug Preferences and Marijuana Legalization

abstractThe response of illicit drug consumers to policy changes like legalization is mediated by demand behavior. Since individual drug use is driven by many unobservable factors, accounting for unobserved heterogeneity becomes crucial for designing targeted policies. This paper introduces a finite Gaussian mixture of EASI demand systems to estimate joint demand for marijuana, cocaine, and basuco (a low-purity cocaine paste) in Colombia, accounting for corner solutions and endogenous prices. Our method classifies users into two groups with distinct preferences over consumption: “soft” and “hard” users. Nationally representative survey estimates find drugs are unit-elastic, with marijuana and cocaine complementary. International marijuana legalization episodes along with Colombia's low marijuana production cost suggest legalization is likely to drive prices down significantly. Legalization counterfactuals under the most likely scenario of a 50% marijuana price decrease reveal \$363/year welfare gains for consumers, \$120M in governement revenue, and \$127M dealer losses.

JEL: D12, C11, C35.

Keywords: Bayesian analysis, Demand systems, Drug consumption, Finite mixture models, Price endogeneity.

Introduction

Drug consumption is a growing global market involving an increasing number of users. According to the 2023 report by the United Nations Office on Drugs and Crime UNODC2023, an estimated 219 million people used marijuana in 2021, representing about 4.3% of the adult population. This was followed by opioids (64 million), amphetamines (36 million), cocaine (22 million), and ecstasy (20 million). In order to cope with the public health outcomes from increasing drug consumption, several nations across the world have enacted laws legalizing recreational drug use. Specifically for cannabis, the most-widely consumed substance, recreational use has been legalized in Canada, Georgia, Germany, Luxembourg, Malta, Mexico, South Africa, Thailand, Uruguay, along with 24 US states, three territories, and the District of Columbia, with many other nations currently considering similar policies. Assessing the impacts of these policies ex-ante requires understanding how consumers might react to variations in the prices of drugs associated with their legalization Becker2006, Hall2016. Additionally, joint policy interventions are usually employed to support the legalization or decriminalization of substances. These interventions are more effective if they target specific segments of the population that face particular necessities. Characterizing population segments that are similar in their preferences for drugs can directly translate to more effective targeting of interventions aimed at these groups.

This paper aims to evaluate the potential effects of marijuana legalization among regular consumers on their welfare, government collection, and drug dealers' revenues, using Colombia as a study case. To study the effect of such a policy ex-ante, we must understand how consumers' drug demand behavior responds to price changes, such as those associated with the possible legalization of marijuana in the country. To this end, we introduce a method to estimate the joint demand for illicit drugs in Colombia, taking preference heterogeneity into account to identify potentially diverse responses to price changes. Given the large rates of violent crime associated with Colombian drug activity, policymakers advocate legalization as a strategy to reduce drug profitability given its causal link to violent crime Fajnzylber2002, Pinotti2015, gavrilova2019legal, queirolo2019uruguay. Evidence supports the claim that marijuana legalization introduces legal competition in the drug market, which decreases profitability and can help reduce crime huber2016cannabis,dragone2019crime,brinkman2019not,burkhardt2019short,wu2020spillover,anderson2023public, particularly violent crimes associated with drug trafficking and resolving disputes burkhardt2019short,chu2019joint,gavrilova2019legal,wu2022effects,anderson2023public. The issue of violent drug crime is of great relevance in Colombia, where homicide rates and violence linked to drug trafficking are pronounced puerta2024spatial.

Our paper also contributes to the literature by studying the demand for illicit drugs using a novel data set in a particularly relevant developing country such as Colombia. The Colombian case is of special interest as it is one of the world's leading producers of marijuana and cocaine UNODC2021, resulting in relatively low prices and widespread access to these substances. The research question is also timely, as the Colombian National Congress recently debated the legalization of recreational marijuana in 2023 but failed to secure sufficient votes in the final round of discussions.\footnote{\url{https://www.reuters.com/world/americas/colombia-senate-votes-down-recreational-marijuana-bill-2023-06-21/}.} According to our main dataset of interest, in a 2019 nationally representative survey of Colombians aged 12 to 65 regarding their consumption of psychoactive substances, marijuana is found to be the most-consumed drug, followed by cocaine, and then basuco (a cheap, highly addictive, and toxic by-product of cocaine production), with other drugs showing only trace amounts of measurable national consumption. Therefore, in this research we focus on studying the joint demand for these three most-consumed drugs in Colombia.\footnote{Currently, possession for personal use is decriminalized in Colombia, allowing individuals to carry up to 20 grams of marijuana and 1 gram of cocaine or basuco.}

To address the technical challenges associated with modeling the demand for illicit drugs and to understand the mechanisms driving the potential effects of marijuana legalization, we propose a new estimation framework based on a microeconomically founded model that takes unobserved heterogeneity, corner (zero) solutions, and endogeneity into account. Specifically, we propose a Bayesian inferential framework that allows for clusters of unobserved preference heterogeneity through a finite mixture of Exact Affine Stone Index (EASI) demand systems. The Bayesian framework allows us some additional advantages that are relevant for our application. First, it allows us to account for corner outcomes where individuals can decide not to consume a given set of goods, which is relevant as many consumers in our sample only report consuming one illicit drug out of the three we consider. Second, it becomes straightforward to include membership to unobserved heterogeneity clusters and associated probabilities as parameters to perform data-driven consumer segmentation and recover heterogeneous drug responses. Third, it allows us to obtain inference on the structural quantities of interest, such as price-demand elasticities or predicted revenue under counterfactual scenarios as a by-product of estimation, all of which are highly non-linear functions of data and model parameters. Finally, we can easily impose and test relevant microeconomic restrictions such as symmetry, strict cost monotonicity, and concavity of the cost function, which ensure the recovered demand functions satisfy standard theoretical conditions RamirezHassan2021, RamirezHassan2024. For further details on Bayesian estimation of demand systems, see Tiffin2010, Kasteridis2011, Kehlbacher2020, Jacobi2021, RamirezHassan2021.

Applying our methodology to the analysis of demand for marijuana, cocaine, and basuco in Colombia, the results suggest that unobserved heterogeneity is crucial to obtaining precise and economically relevant demand behavior when taking into account censoring and price endogeneity. Our method automatically partitions our sample into two sub-populations: “soft” and “hard” drug consumers. The former group represents the largest user segment, spending a majority of their expenditure on marijuana leading to sensible and economically rational demand patterns. In contrast, the latter group has large consumption of the cocaine-based substances and scores highly in the survey questions meant to screen for substance addiction, providing a rationale behind our labels. In addition, we find for the “soft” user group that all three own-price elasticities for each drug are statistically indistinguishable from unity, with a complementary relationship between marijuana and cocaine, and a substitution between higher-quality cocaine and lower-quality residual.

In the most likely scenario that the legalization of marijuana results in a 50% decrease in the price of marijuana (given the high illegal price markup for this drug in Colombia), our estimates suggest that commercial sale of marijuana would imply a considerable increase in marijuana consumption among regular consumers, with relatively much smaller effects for cocaine and basuco. This increase is valued by consumers at approximately \$363 USD of annual utility-equivalent expenditure as measured by the Equivalent Variation of the price change. In addition, we find a high probability that the government will capture a large amount of the legal market resulting in considerable revenue gains even in the face of the reduced post-legalization price. Specifically, assuming that post-legalization legal sales of marijuana accounted for two thirds of all sales (with cocaine and basuco remaining illegal) the government accrues \$120 million annual USD at the expense of dealers who experience a loss of \$127 million annual USD in revenue. These figures imply that drug dealers would need to reach around 130% new domestic drugs users compared with pre-legalization to offset such a revenue decline, which is not likely to be realized based on international legalization experiences. Taken together, these findings suggest that a marijuana legalization policy in Colombia is likely to succeed at disincentivizing drug-related criminal activity; the current largest source of homicides and other violence in the country.

This paper connects with a wealth of previous literature on demand systems, demand for addictive substances, unobserved preference heterogeneity, among others. Demand systems serve as essential tools for learning about consumer behavior and its responses to diverse market conditions Deaton1980. There is extensive previous literature using demand systems to understand consumption patterns for some addictive substances, particularly alcohol and tobacco Thies1993, Deza2023. However, there is a conspicuous gap in the literature regarding the study of joint drug consumption patterns, likely due to issues of data availability and general attitudes towards illicit substances. A key contribution of our analysis is to examine not just a single drug in isolation, but the joint demand for cannabis, cocaine, and basuco. This broader perspective is critical as it enables us to explore whether these substances exhibit substitution or complementary relationships and whether these patterns differ across population groups.

An important advantage of demand systems is their ability to consider joint consumption across categories of goods as well as providing insight into determinants of these behaviors. For instance, duffy2003 uses the Quadratic Almost Ideal Demand system Banks1997 to investigate consumer spending patterns in the United Kingdom, challenging the prevailing assumption that advertising significantly influences preferences, particularly for products like tobacco. aristei2010 conduct an innovative exploration of alcohol and tobacco consumption patterns in Italy, effectively addressing criticisms related to unmeasured preferences and correlated unobserved heterogeneity, shedding some light on the joint determinants of these behaviors.

An additional feature of demand systems is that they allow us to study demand under counterfactual scenarios. That is, we can construct conterfactual price scenarios associated with the legalization of marijuana, to account for mixed effects of legalization on marijuana price (due to shifts away from the black market structure). Evidence highlights varied outcomes depending on the region, market conditions, and regulatory frameworks. For example, the Uruguayan government deliberately kept marijuana prices low to undercut the black market, while in Canada, the price of marijuana increased by 32% after legalization, primarily due to regulatory factors.\footnote{See the official website from Statistics Canada on the price of cannabis post-legalization: \url{https://www150.statcan.gc.ca/n1/daily-quotidien/190710/dq190710c-eng.htm}.} In the specific context of Colombia, the production cost of one gram is approximately US\textcent 1 velez2021medicinal. In contrast, the current illegal market price is around US\textcent 83 (see descriptive statistics from our dataset in Table (ref)). This significant disparity indicates a considerable potential for lowering marijuana prices in a legalized market. Thus, the price elasticities of drug dealers' product offerings are critically important for assessing their potential revenue impacts in a post-legalization scenario.

Other recent research adds further structure to the models in order to dive deeper into the determinants of addictive substance consumption. For example, Lacruz2009 give an in-depth examination of alcohol demand among youth in Spain, highlighting the impact of factors such as income, prices, and the addictive nature of alcohol. Using German survey data, Tauchmann2013 challenge the notion of tobacco and alcohol as substitutes by employing a structural model to directly estimate their interdependence, with profound implications for anti-smoking policies. bokhari2018 considers pharmaceuticals, as they are legal but still potentially addictive substances. This paper analyzes the demand for drugs used to treat attention deficit hyperactivity disorder (ADHD) through the lens of demand systems, focusing on the impact of pharmaceutical mergers. Their study emphasizes the importance of selecting the demand model when evaluating the effects of large-scale market changes, particularly to capture drug substitution patterns. Collectively, these findings highlight the value of demand system analysis in understanding drug consumption patterns and informing effective policy strategies across a wide range of substances.

Previous studies additionally and heavily emphasize the relevance of heterogeneity, in both observable and unobservable factors, to understanding complex patterns found in demand data duffy2003, Lacruz2009, aristei2010, Tauchmann2013, bokhari2018. In estimation of demand systems, researchers often rely on observable features to approximate unobserved preference heterogeneity. This has been implemented in the literature through the incorporation of interaction terms Blundell1993, Moro2000 or by categorizing individuals ex-ante using observable features Bonnet2013. Other studies address unobserved heterogeneity using random parameter models Haefen2004, Lewbel2009 or segmentation algorithms Bertail2008, Kehlbacher2020. In the former, heterogeneity emerges from parameter values that vary across units according to some distribution function, while the latter assumes that units are clustered into sub-populations, with similar preferences within clusters and distinct preferences across clusters.

Unobserved heterogeneity in the demand for illicit drugs is particularly relevant when examining the gateway hypothesis Kandel1975, Lynskey2018, which posits that individuals often transition from using softer drugs, such as marijuana, to harder ones like cocaine or basuco. This, in turn, may lead to complementary consumption patterns between drugs. DeSimone1998 found evidence that marijuana consumption can potentially lead to the consumption of other drugs, where they allow for structural estimates of unobserved heterogeneity to affect both marijuana and cocaine. Ours2003, Ours2006 provide evidence of a causal relationship between cannabis and cocaine use, where the use of multiple drugs is generally driven by unobserved individual characteristics. BrettevilleJensen2008 argues that the empirical association between cannabis and heroin use could be spurious. The authors interpret unobserved heterogeneity as “antisocial behavior" and explain that this could simultaneously affect both cannabis and heroin use.

BrettevilleJensen2008 agree that the gateway effect diminishes greatly when unobserved factors are considered, and Melberg2010 examine the gateway hypothesis, explaining that factors like traumatic childhood experiences could be associated with both cannabis and heroin use. If this is the case, the causal impact of cannabis use on the use of harder drugs would be confounded by the presence of the unobserved trauma factor, mistakenly leading to the conclusion that cannabis serves as a gateway to other drugs. Deza2015 explores the dynamic patterns of drug use, distinguishing between the effects of prior drug consumption and unobserved heterogeneity. The former implies that past experiences influence future choices, while the latter suggests inherent tendencies toward drug consumption in certain individuals. This study finds evidence supporting the hypothesis that consumption of hard drugs complements the consumption of alcohol and marijuana. Jorgensen2022 found that marijuana is not a reliable gateway cause of illicit drug use when taking unobserved effects into account, meaning prohibition policies are unlikely to reduce illicit drug use. In summary, the literature finds that heterogeneous effects are relevant when analyzing demand of illicit drugs, and that there is not conclusive evidence about the relationship between marijuana use and the use of “harder” drugs like cocaine.

After this introduction, we present in Section (ref) our data set, descriptive statistics, and some preliminary descriptive patterns. Section (ref) shows the microeconomic and econometric framework, and Section (ref) shows the results applied to our consumption dataset. We present in Section (ref) counterfactual exercises from implied price changes after a simulated legalization policy on consumer welfare, government revenue, and drug dealers' illegal market size. Section (ref) presents our policy recommendations based on results and concluding remarks.

Descriptive Statistics of Colombian Drug Market

We use the National Survey of the Consumption of Psychoactive Substances performed in 2019 (ENCSPA from its name in Spanish) by the Colombian National Administrative Department of Statistics (DANE). This is a nationally representative survey aiming to measure the consumption of legal and illegal psychoactive substances. Individuals between 12 and 65 years old from several municipalities were randomly selected, and the enumerators privately performed the survey. If the chosen individual was absent during the survey, the enumerator should return later. The survey resulted in a total sample of 49,439 individuals, representative of approximately 23.6 million individuals, which is equivalent to roughly the total urban Colombian population in 2019 within the considered age range.

Our primary objective is to examine the potential effects of significant drug price variations in Colombia on the consumption of marijuana and harder drugs (specifically cocaine and basuco); for example, as a result of a marijuana legalization policy. These three drugs were identified as the most relevant illicit drugs in Colombia in terms of consumption and expenditure. Therefore, we restrict our sample to current regular consumers, defined in the survey as individuals who report using at least one of the three illicit substances at least once per month during the interview year. Limiting our sample to regular consumers results in a total of 1,236 users that are representative of 633,490 users nationwide.

A natural concern with self-reported use of psychoactive substances collected via survey is the potential for under-reporting and reliability of consumption amounts disclosed. While this issue is recognized in the survey design by implementing repeated visits, we separately assess the reliability of the survey in Appendix (ref). By conducting similar descriptive statistic analysis for consumers in single-dweller households, we provide evidence that those with the lowest incentive to misrepresent their information behave similarly to the full sample. This assuages potential concerns of data reliability, and is consistent with previous literature that finds generally low instances of under-reporting with larger cases occurring mostly in younger individuals or specialized groups Needle1983, Harrison1993, Darke1998.

Drug market

According to the ENCSPA survey, the total national expenditure on marijuana, cocaine, and basuco in 2019 was USD 226.3 million among regular users.\footnote{We use the average exchange rate in 2019 (COP/USD 3,282.39) to convert all values in Colombian pesos to US dollars throughout the paper banrep_trm_2019.} Figure (ref) shows the share of total expenditure by age groups and drug of choice (defined as the most-consumed drug). We see that individuals in their twenties spend the most, approximately 51.9%, followed by individuals in their thirties (20.1%), teenagers (14.7%), forties (6.9%), and fifties (6.4%). It is concerning that such a large share is spent by younger individuals. We also see in this figure that marijuana represents by far the most relevant expenditure. Conditional on age group, marijuana shares range between 70% (fifties) and 92.7% (twenties). We see that basuco represents the lowest share, ranging between 0.0% (teenagers) and 3.7% (fifties). Cocaine remains as in-between these two substances, with the largest expenditure share coming from individuals in their fifties spending 26.4% of their drug budget on this substance, with a concerningly large share remaining for both teenagers and young adults (14.7% and 6.3%, respectively).

figure[figure omitted — 1,143 chars of source]

Expenditure is composed of three elements: quantities, prices, and total size of the market. Turning our attention first to market size, we classify the total number of regular users according to their drug of choice, resulting in 550,786 total marijuana users, with 73,578 for cocaine and 9,126 for basuco. Figure (ref) shows the distribution of these users by age group. We observe again that individuals in their twenties represent the largest share (53.0%), followed by individuals in their thirties (18.9%), teenagers (15.1%), forties (7.4%), and fifties (5.5%). We also see again that marijuana comprises the largest share, ranging between 88.4% (twenties) and 77.2% (forties). The second largest is associated with cocaine, where its share ranges between 21.0% (forties) and 9.4% (teenagers). Finally, there is basuco, whose share ranges between 5.2% (fifties) and 0.2% (teenagers). It is concerning that most of drug users are individuals less than 30 years old, and that older individuals have higher concentrations of hard drug use.

Table (ref) shows descriptive statistics for individual-level expenditure shares, quantities consumed and prices faced in the market for drugs. Marijuana's share is the largest (86.9%), followed by cocaine (10.7%). We also report the proportion of zeros in the sample, where marijuana has the lowest figure (6.5%), while basuco has the highest (95.5%). This highlights the importance of taking into account the censoring issue in the econometric framework, as standard models cannot account for the exceeding share of consumers that do not demand some of the drugs.

table[table omitted — 1,750 chars of source]

Conditional on the consumption of each drug, the average monthly consumption of marijuana, cocaine, and basuco is 43.1, 6.6, and 21.1 grams, respectively. However, there is substantial variability. For instance, one individual reports consuming 20 grams (joints) of marijuana per day. The unconditional average monthly consumption of marijuana, cocaine, and basuco is 39.8, 1.3, and 0.95 grams, respectively. This is because most individuals consume only marijuana, as illustrated in Figure (ref). The average marijuana price is 83.3 USD \textcent/gr. with a range from 15.3 USD \textcent/gr. to 305.4 USD \textcent/gr. This high variability in prices is also present in the other drugs: cocaine prices range from 61.1 USD \textcent/gr. to 1,221.8 USD \textcent/gr., and basuco from 30.5 USD \textcent/gr. to 305.4 USD \textcent/gr. This heterogeneity is largely attributed to differences in drug quality and suppliers' location.

figure[figure omitted — 1,911 chars of source]

Due to the clear heterogeneity exhibited in both users and expenditure due to age, we consider the distribution of quantities and price similarly disaggregated by age groups for each drug in Figure (ref). In these figures we focus on those individuals who present only positive consumption to rule out the effects of the large outlier at zero shown in Table (ref). We observe that there is some degree of heterogeneity regarding consumption of marijuana and basuco according to Figure (ref). The average consumption of marijuana ranges between 25.1 (fifties) and 51.2 (thirties) joints per month, and the average consumption of basuco ranges between 2.0 (teenagers) and 29.9 (thirties) grams per month. On the other hand, the average consumption of cocaine is fairly homogeneous, around 6.7 grams per month, except for individuals in their twenties (4.5 grams per month). We note that individuals in their thirties have the largest consumption of all the three drugs, and individuals in their twenties have the second largest consumption of marijuana (43.2 joints) and basuco (20.1 grams).

Figure (ref) shows the distribution of average prices by age group among regular users. There is still some small heterogeneity remaining for the prices per dose among age groups. The most expensive drug is cocaine, where its average price ranges between USD/gram 2.1 (fifties) and USD/gram 3.4 (thirties), with marijuana and basuco having similar prices per doses. Marijuana prices faced by consumers ranged between USD/gram 0.61 (teenagers) and USD/gram 1.02 (fifties), and basuco prices ranged between USD/gram 0.77 (teenagers) and USD/gram 0.94 (forties). It is concerning that teenagers get the lowest average prices of marijuana and basuco. Price variation remains at the individual level due to both observed and unobserved factors, which we exploit in our modeling strategy to construct accurate demand estimates.

Consumer characteristics

The previous analysis aggregated individuals according to their drug of choice to obtain market-level statistics. However, our analysis accounts for the fact that a user may consume multiple drugs. Figure (ref) shows the distribution of users across the consumption bundles implied by our considered drug choices (including irregular and non-consumers for reference). This figure reveals that 97.5% of individuals report not regularly consuming any of the three drugs. Among the remaining 2.5%, 2.34% are regular marijuana users, 0.55% consume cocaine, and 0.11% consume basuco. The majority consume only marijuana (1.95%), followed by those who consume both marijuana and cocaine regularly (0.33%), those who exclusively consume cocaine (0.11%), and those who consume all three drugs (0.06%). Notably, the second-largest group consists of individuals who jointly consume marijuana and cocaine, suggesting potential complementarity between these substances on the extensive margin. In contrast, basuco users represent the smallest share, which aligns with expectations given its high addictiveness and severe adverse health effects.

figure[figure omitted — 1,789 chars of source]

Table (ref) presents descriptive statistics for demographic characteristics available in the survey. We see from this Table that the representative (modal) regular drug user is a male that has drug dealers in his neighborhood, lives in a low-socioeconomic stratum, his friends also use drugs, consumes alcohol and cigarettes jointly, is 29 years old, studied for 12 years, and is currently employed. This person spends on average USD 30.7 per month on drugs, has access to marijuana, cocaine, and basuco, has used marijuana per 11 years, reports having good physical and mental health, and has a high-risk perception about drugs.

table[table omitted — 4,223 chars of source]

The bottom set of variables in Table (ref) includes instruments derived from consumers’ geo-referenced locations and detailed geo-referenced data on drug dealer arrests from the previous year (2018). Using Gaussian kernels with a 1,000-meter bandwidth centered on each consumer’s coordinates, we compute a distance-weighted average of nearby arrests. We also construct average marijuana and cocaine prices using geo-referenced price data. On average, we find approximately 482 drug dealer captures, with a large standard deviation of 660 captures. This large number reflects the prevalence of drug dealing in Colombia and how concentrated it is in specific urban areas puerta2024spatial. The marijuana and cocaine prices elicited from drug dens near consumers are very similar to those reported in Table (ref), but they are more precisely measured and are less prone to measurement error Zhen2014. The role of these instruments is expanded upon in sections (ref) and (ref).

To sum up, individuals in their twenties represent the largest market share, where marijuana is the most relevant expenditure (see Figure (ref)). Although there is some degree of heterogeneity in average consumption, where individuals in their thirties have the largest average consumption (see Figure (ref)), it seems that the market composition is explained by the largest number of users being in their twenties (see Figure (ref)).

Finally, we note that most parametric demand systems, such as the AID or QAID, place a restriction on the rank of the Engel curves (relationship between expenditure and consumption) such that only quadratic relationships can be modeled using these demand systems. This is not the case for the EASI demand system as it can deal with an arbitrarily large rank of the Engel curves. While this is a less restrictive assumption when studying the demand for more standard goods, the demand for illicit drugs presents non-linearities beyond the simple quadratic patterns in Engel curves recoverable by other demand systems. As shown in Figure (ref), a non-parametric estimate of these Engel curves and their slopes using our full sample of consumers showcases highly non-linear patterns for drug consumption, which require flexible demand systems such as the EASI to be captured accurately.

figure[figure omitted — 1,375 chars of source]

Econometric Framework

The EASI demand system Lewbel2009 is constructed to satisfy the standard microeconomic restrictions on demand functions. Specifically, this system satisfies the axioms of choice, such that additivity, homogeneity, and symmetry restrictions are easily imposed to perform estimation. Moreover, the rank in the function space spanned by the Engel curves can be more than three, such that the Engel curves may take flexible shapes as those found in granular demand data. This feature is absent in two of the most widely used demand systems in the literature: the almost ideal demand system Deaton1980 or quadratic AID Banks1997. These properties are particularly relevant in our case given the rank of Engel curves for illicit drugs as shown in Figure (ref) and the fact we work with micro-level data where variation in expenditure is not smoothed out by aggregation Blundell2007, Zhen2014.

Review: the EASI model

Details of the EASI demand system can be found in Lewbel2009. We outline the basic framework here for convenience of exposition. In our application, we simultaneously model $S = 3$ implicit Marshallian budget shares, given by $\bm{\omega}(\widetilde{\bm{p}}_i, y_i, \bm{h}_i, \widetilde{\bm{\varepsilon}}_i) \coloneqq \widetilde{\bm{w}}_i = [\text{marijuana}_i \, , \text{cocaine}_i \, , \text{basuco}_i ]^{\top}$ for individuals $i=1,2,\dots, n$, where

align[align omitted — 656 chars of source]

Implicit utility $y_i$ is an exact affine transformation of the (log) Stone index given by $\widetilde{\bm{p}}_i^{\top} \widetilde{\bm{w}}_i$. Additionally, $e_i$ denotes log-nominal expenditure, and $\widetilde{\bm{p}}_i$ is the log-price vector faced by individual $i$.\footnote{We emphasize that our data allow us to obtain variability in individual price levels, which is not common as in most demand applications consumers face an aggregate measure of price instead.} The system of equations in (ref) can involve polynomials of arbitrary degree $R$ in $y$, providing flexibility for the Engel curves.

The EASI model naturally allows for sources of observable heterogeneity through the inclusion of socioeconomic controls $\bm{h}_i$ and variables that interact with prices ($\bm{h}_i^p$) and implicit utility ($\bm{h}_i^y$). Each of these vectors of controls, with dimensions $M$, $M_p$, and $M_y$ respectively, can contain distinct variables from one another or could be empty (where $h_{i0} = h_{i0}^p = h_{i0}^y \coloneqq 1$ and these first elements are not included in the vector definitions). The stochastic error $\widetilde{\bm{\varepsilon}}_i$, on the other hand, can directly be interpreted as a source of unobserved preference heterogeneity. However, this heterogeneity does not affect the elasticities or Engel curves, as these objects depend on parameters and observable characteristics. Table (ref) presents a complete depiction of all price and expenditure effects that can be derived from the EASI demand system.

To satisfy standard microeconomic regularity conditions in the EASI model, one can impose the following restrictions on the coefficients in system (ref): $\widetilde{\bm{A}}_m\bm{1}_S = \widetilde{\bm{B}} \bm{1}_S = \bm{0}_S$ for cost function homogeneity; the unit-sum constraint of shares requires $\bm{1}_S^{\top} \widetilde{\bm{b}}_0 = 1$, $\bm{1}_S^{\top} \widetilde{\bm{b}}_r = 0$ for $r = 1, \ldots, R$, $\bm{1}_S^{\top} \widetilde{\bm{A}}_m = \bm{0}_S^{\top}$ for $m = 0, \ldots, M_p$, $\bm{1}_S^{\top} \widetilde{\bm{B}} = \bm{0}_S^{\top}$, $\bm{1}_S^{\top} \widetilde{\bm{C}} = \bm{0}_{M}^{\top}$, $\bm{1}_S^{\top} \widetilde{\bm{D}} = \bm{0}_{M_y}^{\top}$ and $\bm{1}_S^{\top} \widetilde{\bm{\varepsilon}}_i = 0$ for $i = 1, \ldots, n$; Slutsky symmetry requires all $\widetilde{\bm{A}}_m$ matrices for $m = 0, \ldots, M_p$ and $\widetilde{\bm{B}}$ to be symmetric; strict cost monotonicity requires $\widetilde{\bm{p}}^{\top} \left[\sum_{r=0}^R \widetilde{\bm{b}}_r r y^{r-1}+\widetilde{\bm{D}} \bm{h}^y + \widetilde{\bm{B}} \bm{p} / 2\right] + 1 > 0$; finally, a sufficient and necessary condition for concavity of the cost function is negative semi-definiteness of the normalized Slutsky matrix $\sum_{m=0}^{M_p} \widetilde{\bm{A}}_m h_m^p+\widetilde{\bm{B}} y + \widetilde{\bm{w}} \widetilde{\bm{w}}^{\top} - \bm{W}$, where $\bm{W}$ is a diagonal matrix whose diagonal equals $\widetilde{\bm{w}}$, and $\bm{0}_S$ and $\bm{1}_S$ are $S$-dimensional vector of zeros and ones, respectively.

Model specfication

We begin by specifying the system of equations (ref) in latent form using the last share ($S$) as the base category or numeraire good (basuco in our case given its generally low expenditure share). These latent shares, denoted by $\widetilde{\bm{w}}_i^* = \left[w_{i1}^*, \ldots, w_{iS}^*\right]^{\top}$, capture the un-normalized marginal utility of individual $i$ of consuming each of the goods considered (illicit drugs in our case). We can then recover observable shares $\bm{w}_i$ from latent ones Kasteridis2011, RamirezHassan2021:

equation[equation omitted — 212 chars of source]

where $L_i = \{l: w_{il}^* > 0\} = \{l: w_{il} > 0\}$ is the set of drugs with positive consumption for individual $i$. Imposing the previously discussed microeconomic restrictions onto (ref) yields an EASI demand system specified directly on latent shares:

equation[equation omitted — 217 chars of source]

where these variables and coefficients are defined from the original quantities as

align*[align* omitted — 1,289 chars of source]

Here, $\bm{p}_i$ represents the vector of relative log prices with respect to the base price $p_{iS}$, and Slutsky symmetry also implies $\bm{A}_m = \bm{A}_m^{\top}$ for $m = 0, \ldots, M_p$ and $\bm{B} = \bm{B}^{\top}$. Note that the unit-sum, cost function homogeneity, and Slutsky symmetry restrictions allow us to recover the share and coefficients from the base category in terms of the information from remaining goods. That is, once we impose these restrictions, we only need to model $s \coloneqq S - 1$ of the shares (marijuana and cocaine).

Our estimation framework also takes endogeneity issues into account. A first source of endogeneity arises mechanically from (ref), with budget shares $\bm{w}_i$ used to construct indirect utility $y_i$, creating simultaneous causality. However, simple valid instruments can be constructed and it has been documented in the literature that this endogeneity is numerically negligible Lewbel2009, Zhen2014. Reverse causality between prices and quantities is also not of general concern when working with micro-level data, as one can argue individual purchase decisions should not affect aggregate market prices.

A final source of endogeneity comes through the relative prices $\bm{p}_i$ due to omitted variables and measurement errors; likely to be more relevant in a micro setting as these are not averaged out when aggregated. For instance, omitted variables can arise through the strategic search of consumers when looking for drug providers, particularly in markets where individuals declare to have easy access to drugs, which translates to relatively good price information and potentially many providers Galenianos2017. Additionally, the prices provided in our data are self-reported by individuals, meaning they can be subject to non-classical measurement errors created by recall, socially desirable responses, or other self-reporting biases Embree1993, Johnson2005, Fadnes2009, Steenkamp2010, Rosenman2011.

Let $\bm{x}_i \coloneqq [1, y_i, \ldots, y_i^R, \bm{h}_i^{\top}, \bm{h}_i^{y\top} y_i]^{\top}$ collect all exogenous variables in a vector of dimension $1+R+M+M_y$ and $\bm{p}_i^* \coloneqq [\bm{p}_i^{\top} h_{i0}^p, \ldots, \bm{p}_i^{\top} h_{iM_p}^p, \bm{p}_i^{\top} y_i]^{\top}$ collect all endogenous variables in a vector of dimension $d^* \coloneqq s (M_p + 2)$. To counteract these sources of endogeneity, we assume we have access to an $\ell$-dimensional vector $\bm{z}_i$ of excluded instruments that are uncorrelated with $\bm{\varepsilon}_i$, are relevant to predict $\bm{p}_i^*$, and we have enough instruments to satisfy $\ell \geq d^*$ as a necessary condition for identification.\footnote{The endogenous $\bm{p}_i^*$ is composed of prices $\bm{p}_i$ and cross-products of $\bm{p}_i$ with exogenous variables $\bm{h}_i^{p}$ and $y_i$. If we have $l$ credible instruments for prices (denoted by $\widetilde{\bm{z}}_i$) with $l \geq s$, then $\widetilde{\bm{z}}_i h_{01}^p, \ldots, \widetilde{\bm{z}}_i h_{iM_p}^p$, and $\widetilde{\bm{z}}_iy_i$ are valid instruments as well as long as they contain sufficient variation across individuals.} We group all structural information in equation (ref) into an $s \times d_{\beta}$ matrix $\bm{F}_i \coloneqq [\bm{I}_s \otimes x_i^{\top} , \, (\bm{I}_s \otimes \bm{p}_i ^{\top})\bm{D}_s h_{i0}^p, \, \cdots , \, (\bm{I}_s \otimes \bm{p}_i^{\top}) \bm{D}_s h_{iM_p}^p, \, (\bm{I}_s \otimes \bm{p}_i^{\top}) \bm{D}_s y_i]$ and all first-stage information into an $d^* \times d_{\gamma}$ matrix $\bm{G}_i \coloneqq [\bm{I}_{d^*} \otimes x_i^{\top}, \, \bm{I}_{d^*} \otimes z_i^{\top}]$.\footnote{$\bm{D}_s$ is defined as the $s^2 \times s(s+1)/2$ duplication matrix such that $\bm{D}_s\operatorname{vec}(\bm{A}) = \operatorname{vech}(\bm{A})$ for any symmetric $s \times s$ matrix $\bm{A}$. Post-multiplication by this matrix comes from the Slutsky symmetry restriction. Additional details for deriving these stacked expressions are provided in Appendix (ref) and (ref).} This results in the following EASI system with endogeneity:

align[align omitted — 199 chars of source]

where $\bm{\beta}$ is the structural coefficient of interest with dimension $d_{\beta} \coloneqq s(1+R+M+M_y) + d^*(s+1)/2$, and $\bm{\gamma}$ are the first-stage coefficients with dimension $d_{\gamma} \coloneqq d^*(1+R+M+M_y+\ell)$. To reproduce the endogeneity in the system, we place a multivariate normal distribution with non-diagonal covariance matrix $\bm{\Sigma}$ between $\bm{\varepsilon}_i$ and $\bm{u}_i$.

To acknowledge further unobserved heterogeneity, we allow for individuals to differ according to unobserved types, where drug demand responses vary across types and are similar for all individuals of the same type. That is, we are assuming that our sample is representative of a population composed by sub-populations, with homogeneous drug preferences within each group and heterogeneity across them. Introduce an individual cluster indicator $\psi_i \in \{1, \ldots, J\}$ such that $\psi_i = j$ means the observation belongs to cluster $C_j$, for $j = 1, \ldots, J$. We will assume $J$, the number of clusters, to be known and experiment with its value in our applications.

As the structural coefficients capture the drug preferences of individuals, we allow for all elements in the structural equation ($\bm{\beta}$ and the relevant components of $\bm{\Sigma}$) to vary at the cluster-level. However, we do not allow for the first-stage parameters ($\bm{\gamma}$ and $\bm{\Sigma}_{u u}$) to vary with the clusters as it is unlikely that the same population segments driving heterogeneity in drug preferences also drive heterogeneity in the reduced form. Additionally, using cluster-specific first-stage regressions do not allow the model to exploit the full variability in instruments to recover these coefficients, leading to artificial issues of weak instruments and larger uncertainty. We incorporate this into our model by assuming

equation[equation omitted — 412 chars of source]

independently across individuals, where $\mathcal{N}(\cdot\mid\mu, \bm{\Sigma})$ represents the density function of a normally distributed variable with mean $\mu$ and variance $\bm{\Sigma}$. Finally, we note that by also including priors for both the cluster indicators ($\psi_1, \ldots, \psi_n$) and assignment probabilities ($\psi_1, \ldots, \psi_J$), their updated posterior produce a data-driven probabilistic assignment of individuals to clusters. This is an important policy tool, as we will see that our model meaningfully classifies individuals according to their drug-use behavior and risk. Additionally, as we correlate the characteristics of these individuals to the model's assignment, it allows us to preemptively identify individuals who are highly at-risk due to their drug use.

Bayesian estimation

We implement a Bayesian inferential framework that allows us to simultaneously handle all key elements of this framework: heterogeneity in structural preferences, censoring at corner solutions, endogeneity, and imposition of microeconomic restrictions. Collect all model parameters into $\bm{\theta} \coloneqq (\bm{\beta}_1, \ldots, \bm{\beta}_J, \bm{\gamma}, \bm{\Sigma}_1, \ldots, \bm{\Sigma}_J, \psi_1, \ldots, \psi_n, \phi_1, \ldots, \phi_J, \bm{w}^*_1, \ldots, \bm{w}^*_n)$, augmented with latent shares. The point of departure for Bayesian analysis is a prior probabilistic belief about unknown parameters, which is updated using sample information $\mathcal{D} = \{\bm{w}_i, \bm{p}^*_i, \bm{F}_i, \bm{G}_i\}_{i=1}^{n}$. Letting $p(\mathcal{D} \mid \bm{\theta})$ represent the likelihood function and $\pi(\bm{\theta})$ be the prior distribution, we can use Bayes' rule to obtain the posterior distribution as

equation*[equation* omitted — 118 chars of source]

Under the distribution for the error terms provided in (ref), the contribution to the likelihood by individual $i$ assigned to an arbitrary cluster $\psi_i = j$ is a function of cluster-specific parameters $\bm{\beta}_j, \bm{\Sigma}_j$ and $\bm{\gamma}$ given by the joint distribution over latent shares $\bm{w}_i^*$ and endogenous variables $\bm{p}_i^*$:

equation[equation omitted — 573 chars of source]

The likelihood function $p(\mathcal{D} \mid \bm{\theta})$ is then the product over all contributions (ref) across individuals $i = 1, \ldots, n$. For a full Bayesian specification, we are simply left with providing priors for estimation. Motivated by the structure of our problem, we assume conditionally conjugate priors with an added twist to the specification: we allow for homogeneous first-stage coefficients while maintaining heterogeneous structural demand preferences (see Eq. (ref) and surrounding discussion for details).\footnote{We also implement heterogeneous first-stage regressions at the same level as the structural equations and provide results below as a comparison.} Combining the likelihood and priors results in the following conditional posterior distributions (where the notation $\bm{\theta}_{-\delta}$ represents the vector of parameters $\bm{\theta}$ with component $\bm{\delta}$ removed):

align[align omitted — 1,099 chars of source]

The posterior hyperparameters in each of these expressions are given by:

align[align omitted — 3,607 chars of source]

Based on these expressions, we provide a Gibbs sampler that can obtain draws from the joint posterior of all parameters. First, conditional on a given value of the latent shares, we obtain a new draw of the model parameters using ((ref))-((ref)). We then draw the latent shares with zero consumption from conditional on those with positive consumption using (ref) and re-compute all latent shares. Repeating this process $S$ times leaves us with a chain of posterior draws $(\bm{\theta}^{(1)}, \ldots, \bm{\theta}^{(S)})$ that we can use to summarize model estimates. Additional details on the computational implementation of our algorithm can be found in Appendix (ref).

Results

We now explore the results of applying our previously described methodology to the analysis of demand for illicit drugs in Colombia. We use the full sample of 1,236 consumers contained in the nationally representative 2019 ENCSPA survey to provide Bayesian inference of the EASI demand system specified by equations (ref), (ref) and (ref) to (ref). In the specification, we include as exogenous information all the variables whose descriptive statistics are presented in Table (ref) except for the geographically distance-weighted variables. To save on degrees of freedom, and as recommended by Lewbel2009, we include these demographics directly in the equation and only explore interactions between price and indirect utility, rather than including additional observed heterogeneity that can greatly increase the size of the estimated demand system.

We tackle potential issues of endogeneity in our application by using instrumental variables based on two sources of information. First, the number of drug-related captures in the neighborhood of the individual allows us to consider supply-side effects exploiting random variation in availability of drug dealers. Second, we use a secondary source of prices to deal with potential search effects (interactions between consumers and dealers) and misreporting by individuals, where these instruments are prices imputed as geographically-weighted averages of the prices in drug dens. As the prices in these locations is both standardized and less prone to measurement error, we can control for additional demand-side variation in consumer prices Zhen2014.

In our implementation, we set the prior hyperparameters to standard non-informative values: for $j = 1, \ldots, J$, $\underline{\bm{\beta}}_j = \bm{0}_{d_{\beta}}$, $\underline{\bm{B}}_j = 1000 \bm{I}_{d_{\beta}}$, $\underline{\bm{\gamma}} = \bm{0}_{d_{\gamma}}$, $\underline{\bm{\Gamma}} = 1000 \bm{I}_{d_{\gamma}}$, $\underline{\alpha} = (1/J) \bm{1}_J$, $\underline{\nu}_j = s(M_p+3)$, $\underline{\nu}_{uu} = s(M_p+3)$, and $\underline{\bm{R}} = \bm{I}_{s(M_p+3)}$. We initially set three clusters in our application (\( J=3 \)). However, we found that one cluster disappeared after discarding the burn-in iterations. Thus, we fixed the number of components at \( J = 2 \), as supported by the variability in the data and instruments. We do not implement any random permutation of the cluster identifiers Fruhwirth2006, as this was shown to hinder convergence of our Gibbs sampler in simulation exercises. We also verify label-switching is not an issue in our application by considering the consistency of individual segmentation of posterior chains (results available upon request).

Using our coefficient draws $(\bm{\theta}^{(1)}, \ldots, \bm{\theta}^{(S)})$ from the EASI specification (allowing or not for unobserved heterogeneity clusters), we can use the posterior draws to compute all relevant microeconomic summaries previously derived in Table (ref). We are then left with posterior draws of these summaries, such that inference on these highly non-linear quantities is a simple by-product of the estimation algorithm; a key feature of Bayesian inference. The coefficients of the EASI model itself are usually not of direct interest, so we provide the full estimates in the Appendix tables (ref), (ref) and (ref). Nonetheless, we highlight that these estimates provide evidence for: (i) the importance of including demographic variables to deal with observed heterogeneity in consumer preferences; (ii) the chosen instruments being relevant and jointly significant in explaining additional variation in drug prices aside from the demographics; and (iii) prices being endogenous due to the significant correlation to consumed shares in their latent disturbances.

We then turn to studying the unobserved heterogeneity clusters recovered by our method, and provide evidence that these identify two clear consumer segments: “soft” and “hard” users. This classification is completely data-driven and automatically identified by the estimation strategy, leading to a valuable policy targeting mechanism in addition to demand system estimates. We provide balance tests across survey variables to further understand the differences between these population segments. Our results confirm that access to and use of harder drugs (among other key demographic variables) are drivers of the classification into one cluster or another, providing further rationale behind our cluster labels. Additionally, we show how the classification into these population segments is correlated to addiction indicators that can be obtained from the survey and that one of the identified segments has similar drug consumption patterns to the population of homeless individuals whom showcase large levels of drug consumption. Heterogeneous Engel curves and income effects are explored in Appendix (ref).

An important aspect of our Bayesian inferential framework is that it allows us to test the microeconomic restrictions imposed for estimation of the EASI demand system. As discussed in Section (ref), while the unit-sum restrictions are imposed due to a mechanical property of expenditure shares, the constraints from Slutsky symmetry, strict cost monotonicity and cost concavity are not innocuous as a way to regularize demand behavior towards microeconomic theory predictions. As Slutsky symmetry implies an equality (or point) restriction, we can use the Savage--Dickey density ratio to calculate the Bayes factor in favor of the restriction.\footnote{Let $M_1$ represent the EASI model imposing Slutsky symmetry and $M_2$ the unrestricted model. Recall that Slutsky symmetry imposes $\widetilde{\bm{A}}_m = \widetilde{\bm{A}}_m^{\top}$ for $m = 0, \ldots, M_p$ and $\widetilde{\bm{B}} = \widetilde{\bm{B}}^{\top}$, meaning $M_1$ imposes an equality restriction on the model parameters of the form $\bm{\theta} = \bm{\theta}_0$. The Bayes factor comparing models $M_1$ and $M_2$ can then be computed using only the unrestricted model according to the Savage--Dickey density ratio kass1995bayes: $BF_{1, 2} = p(\bm{\theta} = \bm{\theta}_0 \mid \mathcal{D}, M_2) / p(\bm{\theta} = \bm{\theta}_0 \mid M_2).$} For cost monotonicity and concavity, which imply inequality (or set) restrictions, we can directly calculate the Bayes factor according to the posterior probability that the constraints are satisfied in our Markov Chain Monte Carlo (MCMC) algorithm draws. Computing these measures for our specification of interest results in very strong evidence in favor of Slutsky symmetry ($2 \log BF_{1,2} = 17.58$), strong evidence in favor of strict cost monotonicity ($2 \log BF_{1,2} = 8.56$), and only slight evidence in favor of cost concavity ($2 \log BF_{1,2} = 1.84$). Evidence generally points towards microeconomic restrictions being satisfied in-sample using the Bayes factor scale provided in kass1995bayes.

Demand-price elasticities

Our first set of results considers the price effects on demand of illicit drugs. As we model the demand of drugs in a fully structural system, we can obtain not just the own-price elasticities of each drug (as usually done in the literature), but also the cross-price elasticities of one drug onto another. This is of particular importance in the face of potential marijuana legalization policies that are likely to have large impacts through price, as the price of one drug will face clear changes due to policy intervention whereas the price paths of other drugs will likely remain fixed in the short-term. In this way, we will be able to study how the effect of a price change due to legalization translates into re-balancing of a consumer's drug bundle.

The main story behind the results is presented in Table (ref), which provides estimates from all our Bayesian specifications. The panels of the Table are arranged in the way we were led by our exploration of the different specifications and understanding of the data. We begin by considering estimation of the EASI model accounting for corner solutions (censoring) due to zero consumption of drugs by many consumers, before augmenting the model with price endogeneity and using external instruments to recover identification (two top panels in Table (ref)). Both models do not account for unobserved heterogeneity. In the endogenous specification, the relevance of endogeneity is evident when comparing the two sets of results, as the inclusion of instruments highlights the high levels of price sensitivity in cocaine and basuco. However, this comes at the cost of a remarkably reduced precision. These results align with previous literature that has found marijuana to be an inelastic product in countries such as Australia jacobi2016marijuana, Colombia ramirez2023marijuana, South Africa Riley2020, Thailand Sukharomana2017, and the United States Davis2016. On the other hand, the demand for more harmful drugs, like cocaine and basuco, has shown greater elasticity jofre2008trading,Gallet2014. Regarding cross-price elasticities, limited evidence finds some complementarity between marijuana and cocaine jofre2008trading.\footnote{Tables (ref) (full sample) and (ref) (“soft” cluster) in the Appendix show similar patterns using least squares (LS) and two-stage LS. However, these estimators do not account for censoring nor unobserved heterogeneity.}

Our next two specifications introduce unobserved heterogeneity by considering two clusters. The results presented are for the subsample classified as “soft” users, whose parameter estimates are the most stable (those for the “hard” cluster are highly unstable or not updated through the data due to its small sample size). Again, we observe the relevance of endogeneity when comparing these two sets of posterior estimates. We improve the precision of the posterior estimates by fixing the parameters of the first stage to be homogeneous across clusters and, consequently, fully exploiting the variability in the instruments, leaving only unobserved heterogeneity in the structural demand equation.\footnote{The two bottom panels in Table (ref) in the Appendix show the results assuming heterogeneity in the price equations. We observe greater variability in the results.}

We focus the remaining analysis of the results based on the last specification with a homogeneous first-stage. This specification identifies a total of 1,069 “soft” consumers in our sample, representative of 551,507 drug consumers nationwide (the remaining 167 in-sample “hard” consumers represent 81,983 national consumers using the expansion factors from the survey). The full marginal posterior distributions and 95% highest posterior density intervals of the demand price elasticities from this specification can be visualized in Figure (ref). Additional visualizations are provided in Figures (ref) (trace plots) and (ref) (autocorrelation plots), with draws shown to satisfy standard posterior convergence diagnostics such that posterior expectations should be accurately computed (results available upon request).

figure[figure omitted — 694 chars of source]

The demand-price elasticity estimates from this specification suggest that “soft” users are unit-elastic to the price of illicit drugs (a 1% increase in price implies a 1% decrease in consumption), such that they are highly responsive and their consumption follows the rational law of demand. These elasticities are statistically significant, as the 95% credible intervals do not include zero. The similarity of these own-price elasticities is novel in comparison to previous literature, which clearly identifies greater price sensitivity for harmful drugs than for marijuana. See the panel Bayesian Censored with Endogeneity but without Mixtures in Table (ref). The difference arises because prior research overlooks unobserved heterogeneity, combining individuals with different unobserved preferences. In particular, we show below that the “soft” cluster spends the most on marijuana, with a small percentage spent on cocaine and basuco. As a result, price variations in these two drugs do not have a high impact on their spending, and consequently the price elasticity exhibits similar values to those found for marijuana. Unobserved preference heterogeneity is therefore crucial for identifying meaningful price effects.

Additionally, as we have the cross-price elasticities of demand, we can determine that both harder drugs, cocaine and basuco, are complementary to marijuana in the sense that it follows the same direction as the own-price effect of marijuana (a 1% increase in the price of marijuana results in a decrease of 0.03% in the quantity of cocaine and 0.11% for basuco). Similarly, marijuana responds in the same direction to changes in the price of harder drugs, suggesting further complementarity, though the magnitude of this effect is one or two orders of magnitude smaller, making the complementarity asymmetric. Finally, we note that cocaine and basuco are instead substitutes, as can be expected from basuco, being a lower-quality product and considered to be inferior to cocaine. Our estimates suggest a 1% increase in the price of cocaine results in an increase in the consumption of basuco of 0.14%, whereas a 1% increase in the price of basuco results in an increased cocaine consumption of only 0.03%. This is in line with the much larger average price per gram of cocaine, such that a 1% increase in the price of cocaine has larger level effects on expenditure than the corresponding increase for basuco. However, the uncertainty associated with the cross-price elasticities varies and the 95% credible intervals include zero, with a high probability of a complementary effect between marijuana and cocaine remaining as the key effect, which is consistent with previous literature. Further results are provided in Appendix (ref).

table[table omitted — 4,421 chars of source]

Consumer segments by drug preferences

The finite mixture framework used in this paper allows us to provide posterior estimates of both cluster indicators and probability of belonging to each cluster for all individuals in our sample. We first provide evidence that the clusters identified by our algorithm are stable, in the sense that the classification of individuals is sharp and consistent across posterior draws. Figure (ref) showcases the posterior inclusion probability to $C_1$, normalized to be the “soft” cluster of consumers, denoted as $\Bar{\phi}_{i, 1}$ in (ref).

Observe that there is a sharp edge that distinguishes the probability of inclusion of individuals to the “soft” and “hard” clusters, with a few individuals falling on the frontier between these two states. This provides a way for the model to identify individuals that can be assigned to either cluster, which can be interpreted in our application as those at risk of moving from being a soft consumer towards harder use. That is, these users are identified as the ones with the highest potential of fulfilling the gateway hypothesis, where their marijuana consumption is large but skewing towards cocaine and basuco.

To add further evidence to the rationale behind our cluster labels, we evaluate the average differences between all relevant covariates collected in the survey. Figure (ref) showcases the sorted T-statistics obtained from a mean-difference test that takes into account the different cluster sizes (see numerical p-values and intervals in Table (ref)). The largest differences between users in the “hard” and “soft” clusters arise due to the former having higher access to and consumption of both cocaine and basuco, including addiction-related questions contained in the original survey. Other key differences in terms of demographic characteristics include larger probabilities that individuals are consumers of substances in their personal networks, have access to a drug dealer, and use tobacco and alcohol jointly for the “hard” cluster. Additionally, individuals in this cluster are older, report feeling mentally healthy less often, have less education, and are more likely to be male.

Based on this classification, we can probe the differences between the characteristics of individuals that are sorted into either cluster. Figure (ref) presents the consumption shares of individuals classified into the cluster of either “soft” or “hard” users. One of the main reason behind this labeling is the fact that consumers classified into the second cluster have larger shares of both cocaine and basuco compared with the share of marijuana. Given the more addictive nature of these substances, when individuals include them in their consumption bundle, it is more likely that this will be correlated to negative outcomes in terms of their covariates.

figure[figure omitted — 1,002 chars of source]

Particular attention must be given to individuals belonging to the “hard" cluster, as their drug consumption patterns are similar to those of homeless individuals, who report basuco as the most-consumed drug according to the Census of Homeless Individuals in Colombia (2019). We also examined the answers provided by the “hard" cluster to questions related to drug-related risks, such as issues with friends, family, or work due to drug consumption. We find that the hard-drug user group, on average, responded positively to 5 out of 8 questions, whereas this rate is measured to less than 1 in 8 for those in the “soft” cluster. Using this procedure, we identify a subset of “soft” consumers classified as “at-risk" in our model (representing 5,716 individuals) who show indication of moving towards harder drug use (green triangles in Figure (ref)).

Marijuana legalization counterfactuals

In this section, we use our estimated drug demand behavior to conduct counterfactual exercises of the potential effects of a marijuana legalization policy. Following the large political backing that such a policy has already amassed in Colombia ---as well as the legalization experience of Uruguay, a country both geographically and economically similar--- it becomes crucial to understand its potential effects. This is emphasized by the idiosyncrasies of the Colombian drug market with its large production amounts, low costs, and salient demographics, as well as the large sources of heterogeneity in drug demand by consumers found in our estimation exercise.

Legalization policies will directly affect consumers' access to the drug and prices faced for its legal purchase. The final effect on marijuana price in Colombia is uncertain as legalization could entail conflicting forces in the market. Accounting for direct production costs for marijuana of US\textcent 1 per potency-adjusted gram velez2021medicinal, as well as average logistical costs and the rate of return, the lower bound on the tax-free price of a gram of marijuana should be US\textcent 2. This matches the average tax-free price of a cigarette in Colombia ramirez2023marijuana. The survey we use shows the average market price of marijuana in Colombia is US\textcent 83 (see Table (ref)), with the difference between this and the tax-free rate likely attributable to the illegality margin. By reducing such large illegality margins, legalization in Colombia is potentially likely to reduce the price of marijuana. On the other hand, marijuana legalization would introduce additional taxes, operating and regulatory costs, which could potentially dominate the final price caulkins2015options.\footnote{Optimal taxation for “sin goods” such as marijuana requires additional information or models on dynamic formation of addictive behaviour Gruber2001, substitutability within and between sin and non-sin goods Arnabal2021, Fleissig2021, demand for potency Hansen2020, or preferences over inequality Lockwood2017, among others. As such information is not available in the survey, this exercise is outside the scope of the current paper and left for future research.}

We therefore consider several scenarios based on the potential price changes of marijuana that such a legalization policy might entail (leaving all other drug prices fixed), considering a representative agent that did not have access to a dealer pre-implementation but now does have access due to legalization.\footnote{We fix the values of the covariates for the representative agent to their means for continuously-distributed covariates (i.e., “average” agent) and to the mode for discretely-distributed covariates (i.e., “modal” agent). These values are fixed throughout the counterfactual exercises to make all results comparable.} Figure (ref) showcases the average post-legalization price of marijuana implied by each scenario. All prices occur naturally within the distribution faced by individuals in the sample, showing potential for external validity of the provided legalization scenarios.

enumerate• Our most likely scenario is one of a 50% decrease in prices of marijuana, given the low production and distribution costs of marijuana in Colombia. This scenario mirrors Uruguay’s recent experience following marijuana legalization, where the price of marijuana was reduced to incentivize demand in the newly legal market pluasbeyond. • The potential for legalization to decrease prices even further, setting the final price at 10% of the initial value. This would be comparable to the current tax-included price of a cigarette in Colombia. • The overhead costs of legalization might outweigh its cost benefits leading to a 25% net increase in the price of marijuana, similar to the counterfactual exercise proposed by jacobi2016marijuana.

We present results showcasing the counterfactual effects of this path of legalization policy on drug consumption, consumer welfare (as measured by the absolute value of Equivalent Variation), government revenue, and illegal drug market size. Estimating the effects of such price changes and implied taxes on aggregate outcomes requires assumptions on both demand and supply forces in the market Angrist1996, Imbens2014. Given the structure of the drug market in Colombia and low marginal cost of production for marijuana, it is reasonable to approximate the supply curve for marijuana as perfectly elastic at traded prices. Under such assumption, price effects can be directly identified from our demand curve estimates that rely on instrumental variables. A limitation of our approach is the inability to include dynamic or post-legalization effects, given that the survey has not been conducted since its initial wave and legalization itself has not been implemented. Therefore, our results can provide short-run effects without an increase on the user base, which can be interpreted as lower bounds for government revenue and upper bounds for dealer losses. Providing results ex-ante to the legalization is crucial as it allows policy makers to consider desirable scenarios prior to implementation.

As most demand system models are based on expenditure shares, they cannot identify changes to total expenditure nor changes in total quantities. Therefore, we approximate percentage changes to the consumed quantity of drug $l$ ($q_l$) given a change to the price of drug $j$ ($p_j$) as $\%\Delta q_l \approx \exp \{\widetilde{\varepsilon}^{M}_{lj} \cdot \Delta \log(p_j) \} - 1$, where $\widetilde{\varepsilon}^{M}_{lj}$ are the estimated demand-price elasticities (see Table (ref)). Given our Bayesian framework, we are able to provide full posterior distributions over all policy quantities involved in the analysis, meaning inference also comes as a by-product of estimation. The uncertainty associated to our Bayesian results for the “soft” cluster of users implies highly accurate estimates of all quantities considered here.

We provide a full description of the results for the 50% decrease in the body of the paper, exploring alternative scenarios in Appendix (ref). Figure (ref) shows the posterior distribution of the equivalent variation (EV) in annual terms following the price change, as deduced from the EASI model RamirezHassan2024:

align*[align* omitted — 325 chars of source]

where $e$ is the total monetary expenditure in the three drugs, $\widetilde{w}_s^0$ and $\widetilde{w}_s^1$ are the pre- and post-legalization expenditure shares, $\widetilde{p}_s^0$ and $\widetilde{p}_s^1$ are pre and post ($\log$) prices, and $\widetilde{a}_{sj}$ is the $sj$-th element of matrix $\widetilde{\bm A}_0$, respectively. The result suggests that the representative “soft” consumer perceives a change in their utility that is equivalent to a median change in total expenditure of approximately \$363 USD, with a 95% credible interval of $(338, 374)$. This change is approximately equivalent to the annual average drug expenditure, meaning consumer welfare effects are considerable and largely due to an increased marijuana consumption that offsets the price decrease, with the remaining drugs experiencing negligible changes.

figure[figure omitted — 1,911 chars of source]

After the legalization is implemented, previous studies suggest only a fraction of users will switch to the legal marijuana market, leaving approximately 34% of the illegal market active ramirez2023marijuana, pluasbeyond,\footnote{For instance, the size of the black market for cigarettes in Colombia is 34% (see the study \href{https://www.semana.com/salud/articulo/consumo-de-cigarrillos-34-de-cada-100-fueron-de-contrabando-esto-revelo-la-federacion-nacional-de-departamentos/202315/}{“Consumo de cigarrillos ilegales Colombia” 2022}). Similarly, 1 out 3 marijuana consumers in Uruguay reports obtaining marijuana from an illicit provider after legalization.} whereas dealers maintain sole control of the cocaine and basuco markets due to their continuing illegal status. Figure (ref) shows the posterior distribution of the estimated government tax revenue from the “soft” consumers, with a median value of approximately \$121.5 million USD and a 95% credible interval of approximately (120, 123) million USD. The posterior distribution is derived from a tax of approximately US\textcent 39 per marijuana joint (calculated as the average price of US\textcent 41 minus the tax-free price of US\textcent 2) and the predicted monthly number of marijuana joints consumed by individuals in this cluster, again considering expansion factors.

Additionally, Figure (ref) shows that dealers would face a median revenue loss of \$127 million USD from the “soft” consumers, the 95% credible interval is $(-129, -125)$ million USD (similar figures for marijuana profits that subtract marginal cost are provided in Figure (ref)). In order to make up for such losses after legalization, drug dealers will need the total number of drug users to increase. The annual weighted average expenditure in drugs post-legalization is estimated to be approximately \$157 USD (using the fraction of expansion factor over the sum of total users as weights). Based on this value, we can calculate the required number of average users required to offset a revenue change equal to the one found in Figure (ref). This provides a full posterior distribution over the number of required users given in each price scenario. Figure (ref) showcases that drug dealers would need an approximately 130% increase in the number of users spending at the annual average rate of \$157 USD to offset the \$127 million USD total revenue decrease. This corresponds to approximately 825,000 new users, which would imply a jump in the regular use rate from 2.5% to 5.8% of Colombians aged between 12 and 65 years. No state or country in the world has exhibited such an increase following legalization anderson2023public. The estimated increase in Colombia would be approximately closer to 30% after legalization ramirez2023marijuana, which is similar to the figures found in the United States after legalization in different states anderson2023public.

table[table omitted — 1,885 chars of source]

Table (ref) summarizes the point estimates and credibility intervals of the policy-relevant quantities across all legalization price scenarios considered. We see that the magnitude and sign of the effect on consumer welfare as measured by the EV depends on the price change, with price decreases having a larger compensating effect on consumers compared to the unlikely price increase. On the other hand, the estimated revenue collected by the government and losses experienced by the dealers remains consistent across the scenarios, reflecting that the largest portion of the policy effect can be explained by the large loss to marijuana market shares experienced by dealers. In summary, a legalization policy applied to the Colombian context is estimated to increase marijuana consumption, increase welfare for consumers through standardized pricing and access, and reallocate considerable revenue sources to local governments from illegal drug dealers. Legalization thus has the potential to decrease drug profitability and disincentivize the use of violence to control the domestic black market, which is the current a large source of violent crime in Colombia.

Policy recommendations and conclusion

This paper models the joint demand for illicit drugs in Colombia, introducing a Gaussian mixture of endogenous EASI demand systems estimated via Bayesian methods. We provide evidence that unobserved heterogeneity is a key driver of drug demand, and using our data-driven cluster assignment, we simultaneously classify individuals into either a “soft” or “hard” drug consumer segment and estimate the drug demand behavior of each segment. We control for both demand- and supply-based endogenous sources of consumer price variation using as instruments geo-referenced, distance-weighted averages of drug-related captures and elicited prices. Tailoring the priors to suit the setting of structural demand modeling and to take full advantage of the variability in instruments, our Bayesian results are able to provide accurate estimates of EASI coefficients and their demand implications. In particular, the results from the “soft” cluster of consumers provide accurate descriptions of drug demand patterns that represent over 550,000 consumers across Colombia.

The framework emphasizes the importance of modeling drug demand jointly, which is not usually available given data limitations in other sources. This importance is reflected in our estimates, where complementarity and substitution effects between drugs arise. Specifically, results for our preferred specification suggest that marijuana and cocaine are complementary in an asymmetric way, such that increases to marijuana prices decrease both marijuana and cocaine consumption, but increases in cocaine price only have sizable effects on the consumption of cocaine, leaving marijuana consumption statistically unchanged. Basuco is instead found to be an inferior substitute for cocaine, though largely unrelated to marijuana. Additionally, our estimates provide a good fit to the implied Engel curves in the data, showcasing the importance of accounting unobserved heterogeneity and price endogeneity to model drug demand in the Colombian setting.

Finally, we used our estimates to evaluate potential effects of a marijuana legalization policies based on different price scenarios the policy could entail. Given Colombia's drug production structure and previous experiences in legalization in developing countries, the most likely scenario of a 50% decrease in the price of marijuana creates sizable gains for consumers and the government as taxing legal providers of marijuana, with heavy losses incurred by actual illegal drug suppliers. These losses account for the fact that an illegal market of marijuana will continue existing alongside the legal purchasing system, but at a greatly reduced size compared to pre-legalization. Specifically, estimates suggest that revenue gains to the government will be close to \$120 million USD, with suppliers losing upwards of \$127 million USD in revenue. These results suggest that the profitability of the illicit drug market would decrease, and consequently, reduced incentives to control the domestic black market would lead to a decrease in violent crimes associated with local drug trafficking. In addition, the estimated unit-elasticity of marijuana and cocaine in Colombia implies that a legalization policy resulting in a 50% decrease in marijuana prices would keep the total size of the drug market largely unchanged, shifting from \$226.3 million USD to an estimated \$228.0 million USD post-legalization. This would come at the cost of higher consumption levels among existing regular users and a higher likelihood of attracting new consumers, as suggested by the literature on the effects of marijuana legalization. Nevertheless, we find that these new consumers would not considerably compensate for the lost income of drug dealers.

Given the largely mixed effects of marijuana consumption on both consumer health and job market outcomes found in the literature, our results imply that while policy makers would clearly benefit from a legalization policy and dealers would be clearly negatively impacted, the effects on consumers are not clear cut. On one hand, while consumption of marijuana is expected to almost double for the actual representative agent (given the 50% price decrease at unit-elasticity), this is valued by consumers at approximately \$363 annual USD of utility-equivalent expenditure, which represents 100% of current weighted average total drug expenditure, and around 12% of the yearly minimum wage in 2019. On the other hand, the increased consumption would have additional heterogeneous indirect effects on public health, productivity, educational attainment, etc. Therefore, it is key that a legalization policy is accompanied by additional targeted policy efforts for each population segment, particularly for individuals in the “hard" and “at-risk" user groups.

Appendices