Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
95,846 characters · 14 sections · 18 citation commands
Measuring the Demand Effects of Formal and Informal Communication : Evidence from Online Markets for Illicit Drugs
Product information is essential for markets to function well. Without accurate knowledge of what they are buying, consumers cannot be expected to maximize personal welfare over the set of product choices. Of course, product characteristics that affect a consumer's ability to optimize in a marketplace go beyond knowledge of its individual components. In the absence of strong institutions, consumer's might greatly value a seller's reputation, or be concerned that a seller might cheat them, particularly if they are unable to observe product quality at time of purchase. Of critical importance in this setting would be the avenues through which consumers can collect information about the product space they seek to purchase from.
In this paper, I examine a market with exactly these characteristics, and attempt to measure the relative importance of differential modes of information transmission in relation to consumer demand. Specifically, I study the demand for illegal products on online “Darknet” markets, which lack credible enforcement of laws and so must rely on incentives from vendor reputation and consumers being appropriately informed to function well,. Because of this and the concentrated set of locations where consumers can gather information about products, I am able to plausibly estimate the impact of product sentiment on future market demand, by controlling for all feasible supply-side decisions that may impact consumer demand. By “sentiment”, I mean whether a vendor is being talked about in a positive or negative way, as measured by the word structure in the message. I find that both informal and formal communication (as measured by forum post and product review sentiment, respectively) have a large effect on consumer demand, and that these effects are comparable in magnitude. I find that these effects scale with sample size of the information set , but little other evidence for heterogeneity in the effects of product sentiment from varying the information sources.
My marketplace is the illegal online market for illicit products, such as guns, narcotics, and fraud services, commonly known as the “Darknet” of the Internet. Due to the illegal nature of all transactions being performed\footnote{even if the product itself is legal in a given country, no government taxes are ever paid by the vendors, so they are always illegal.}, this marketplace is highly anonymous. Vendors and sellers alike are potentially culpable in the distribution of these goods, which in developed countries such as the United States, can result in up to 15 years of imprisonment csa. In addition to the clear personal incentives one might have to remain anonymous, anonymity is enforced by requiring users to connect to these online marketplaces via the TOR protocol. TOR is a specific type of Internet routing service that scrambles ones personal data packets by interchanging them with the data packets of other current active TOR users dingledine2004tor. The net result is that an observer, whether a government or a private organization, is unable to tell what IP address a traced data request is coming from, thereby offering an additional layer of security for both customers and vendors. On top of this, all transactions on the marketplace are done in Bitcoins, a popular cryptocurrency notorious for its ability to anonymize transactions.
Besides the strong protections of anonymity used by the market, Darknet marketplaces operates like virtually any other e-commerce website. Listings are organized into relevant product categories (LSD, Opiods, Firearms, etc.), and customers can use a search engine to locate any specific product they wish to consume or vendor they wish to buy from. On an actual product page, buyers can view a provided description and image of the product from the seller, in addition to the full history of customer reviews from previous buyers, which each have a 0 to 5 star rating, a brief text description provided by previous buyer about their purchase experience and/or the quality of the product, and how long ago the review was made. In addition, the product page has summary statistics on the vendor, such as average rating across all products, total volume of transactions, where the buyer ships to/from in the world, along with a link to a vendor's profile page where one can view past reviews of all previously bought products. Figure (ref) provides an example of a typical listing on the Darknet marketplace I use as my universe for this paper, Agora.
One rather unique feature of these Darknet markets is that, in order for customers to finalize the transaction, buyers are required to leave a review once they have received the product. In order to ensure the market functions while guaranteeing agent anonymity, Darknet markets use an escrow system, where the marketplace acts as clearinghouse and transfers the Bitcoins from the buyer to the vendor only after the consumer verifies that they have received the product they ordered and leaves their review. The reasons for this are threefold:
However, a somewhat unfortunate consequence of this well meaning rule is that many vendors explicitly require customers to “Finalize Early” when receiving products. This practice involves customers finalizing their transactions with the market clearinghouse in advance of the vendor actually shipping the product, so that the vendor does not have to wait for delivery to receive the payment, a process that can sometimes take weeks. A customer may later edit their review appropriately to match their actual experience, but it is common for a reviewer to mandate that those who finalize early also give a 5-star review; otherwise they will not fulfill the transaction. Part of what originally motivated this paper's investigation into text sentiment was the common practice in Darknet markets of mandating 5-star reviews on finalize early orders, but little done to regulate the text accompanying the review (possibly because it is not an input into a vendor's overall rating, or possibly because it is much more difficult to filter text). Perhaps because of the anarcho-libertarian ideology that surrounded the creation of these Darknet marketplaces, this practice was not explicitly banned by moderators, since advertisers are transparent about this requirement if you wish to purchase from them. Instead, moderators built an explicit “No Finalize Early” flag placed on all listings in search results so that consumers could easily filter out product offerings that may require finalize early. Misclassifying one's product listing as no finalize early can result in fines to the vendor and in some cases banning the vendor altogether from the market.
One can imagine that all of this taken together implies a huge premium on a vendor's reputation. Despite the protections put in place, this market should still gravitate towards a set of well-established vendors who are well known to serve customers honestly. For this reason, changes in attributes of a vendor reputation (such as community attitudes towards them) should have more pronounced effects in this market, leading to more precise estimations of their effects.
There is a limited literature on Darknet markets due to both their recent inception and illicit nature. Most of the work is done from a criminology perspective, explaining the difference in structures between Darknet markets and traditional illegal drug rings or cartels. One of the first studies of Darknet markets by Christin christin2013traveling provides a comprehensive measurement analysis of the Darknet market during its initial stages when it consisted of one website. Demant et al. demant2016personal is one of the first papers to examine the Darknet from a social science perspective. They investigate whether consumers on the Darknet are redistributers or direct consumers, and find suggestive evidence that consumers on the Darknet resemble direct-to-consumer sellers, a step below in the drug supply chain.
There is some relevant work that has been done on the importance of communication in online settings. Lewis lewis2006asymmetric examines the effects of differential vendor communication or advertisement on sales in eBay's online platform for buying and selling used cars. Luca & Georgios luca2016fake study the incentives of vendors listed on Yelp to procure review fraud in order to boost their own ratings on the website.
The market I study is one in which consumers choose to purchase different types of good (mostly different types of drugs). I assume that demand for product $j$ in time interval $t$ is given by a Poisson distribution: \[Pr(y_{jt} \textrm{ units purchased} | X = \frac{e^{y\beta X} e^{-e^{\beta X}}}{y_{jt}!} \] More importantly, the result of this functional form is that our expectation has a straightforward exponential form: \[E[y_{jt} | X_{j,t},\mu_j] =e^{X_{j,t}\beta + \mu_j }\] Where $X_{j,t}$ is a vector of time and product varying characteristics, and $\mu_j$ is a product-specific effect on average product demand. More plainly, I model my covariates of interest, $X_{j,t}$, as having a multiplicative $e^\beta$ effect on (mean) product demand for every 1-unit change.
While it is well known that product demand is unobservable, since prices and quantities are results of both demand and supply schedules clearing, I argue in this paper that, due to the unique properties of the market I study, a researcher can observe virtually everything suppliers choose here that are observable to customers. And thus, a researcher can plausibly control for all supply-side decisions a seller takes that impact a consumer's choice. Thus, assuming all supply decisions are perfectly controlled for, remaining variation in sales can be explained as a combination of demand-side variation and noise. From here, I take relevant covariates not determined by the supplier, and determine their effect on consumer demand.
My demand-side covariates will mostly examine the effects of peer communication on a vendor's reputation and hence the demand for their products. I examine the impact of two channels of reputation formation: sentiment or satisfaction displayed by past customers on a product's review page, and sentiment concerning a vendor on the designated market forum, which was specifically established so customers could communicate to each other about events or products within the market. I designate these two channels as “formal” and “informal” communication, since one exists in the formal marketplace setting, where consumers are asked to explicitly review a product, while the other exists in a social setting where users are free to spontaneously discuss whatever they wish. This may sometimes include experiences on buying from certain vendors. I expect that both of these are crucially important to the demand in this market. Because there are virtually no other places on the Internet for consumers to communicate with each other (largely due to criminality concerns), vendor ratings, past reviews, and the forums are the only ways for potential customers to collect information about a product or seller.
It is clear, especially since reviews are required of every past customer, that this will be one of the primary channels for customers to collect information on whether a vendor is selling high or low quality products, and whether or not they are committing fraud. Ex ante, though, it is not obvious that customers would highly value positive or negative reviews of vendors on forums. Since forum posts can be written by anyone, not just past customers, a consumer might view a vendor review on a forum as “cheap talk” since it is costless for a single user to send an arbitrary message on the forums, and thus they may ignore this information entirely. And so, especially since the author's utility is unlikely to be determined by the potential customer's purchase decision, the outcome may be a babbling equilibrium where forum messages are entirely noninformative farrell1996cheap. At the same time, one might imagine there are sufficient incentives for forum members to accumulate social capital among their peers and obtain a reputation to be a member in good-standing with the community (See Wasko & Faraj wasko2005should for an investigation of incentives for knowledge sharing in an Internet forum setting). Given the limited means for consumers to collect information, I hypothesize that in this marketplace, sentiment of vendors displayed on forums will have an influential role in customer demand. In addition, I propose the following hypotheses for how consumer demand will depend on differing sources of vendor sentiment, which I test later in the paper:
For my analysis, I use three datasets concerning the anonymous market for drugs on the Internet. All of these datasets were generously uploaded for public usage by Gwern Branwen gwern.
The first dataset is comprised of weekly, complete, sets of listings available on all Darknet marketplaces from a central Darknet search engine named GRAMS. GRAMS interacts with the API of major marketplaces on the Darknet to obtain a complete set of listings on each marketplace. However, the information GRAMS provides is more basic than that contained in some of the datasets I describe below, a trade-off that must be weighed against its relative completeness. GRAMS has for each item the description, price, and seller name, along with a field for which Darknet market (Silk Road 2, Evolution, Agora, etc.) it is being sold on. Namely, it does not contain any sales data or review data for each listing. I use the GRAMS dataset as a sort of validation set to compare with my incomplete HTML marketplace data, as well as for testing vendor responses to changes in sentiment on an extensive margin. \\ I use some of the evidence in this GRAMS dataset to determine which marketplace websites would be best to study. Figure (ref) shows the percent of listings (from the complete GRAMS data) on the Darknet from three of the largest marketplaces on the Darknet from 2014 to mid 2015. The sample period can be broken up into three phases; before The Silk Road 2 went offline (red line), before Evolution went offline (blue line), and after. On November 6th, 2014, The Silk Road 2 server (started as a successor to the infamous original Darknet market, The Silk Road) was seized by U.S. customs authorities and brought offline sr2bust. It did not enjoy as large of a market share as its predecessor, partly due to a loss of a first mover advantage, and partly due to its buggy interface. Evolution, on the other hand, enjoys the largest market share during the sample period I later focus on (the year 2014). Known for its high degree of security (unlike Silk Road 2), Evolution suddenly, without notice, went offline on March 14th, 2015, to the bewilderment of it's users. It was confirmed by moderators on the website that the owners of the site “cashed out” and stole the asset holdings of vendors currently in escrow on the website, estimated to be worth \$12 million USD evoamount.Because of the abrupt and peculiar circumstances surrounding the closing of the two other markets, I chose to not include these as the main object of study for this paper. Compare this to the shutdown of Agora, the other major market during this period, whose owners publicly announced in August 2015 that they were going to indefinitely take the website offline due to the increasing risk associated with running a Darknet website and security concerns. Everyone's asset holdings on the market were returned free of charge. Considering the sincere nature of it's exit from the marketplace , along with its relatively high market share, Agora seemed to be the best candidate market to study.
The second dataset I use consists of approximately biweekly HTML scrapes of one of the Darknet's most prominent markets for drugs, Agora, from January 2014 to July 2015. The sample period I use for this research is data from the 2014 calendar year . These scrapes contain a rich amount of unstructured data on product listings on the market. Namely, they contain the complete product description provided by the vendor, an (optional) image of the product, the overall rating of the seller alongside the number of transactions the vendor has previously engaged in, shipping location, price, and partitions of the products into categories. Within categories, listings (theoretically) vary only on quantity and quality. In addition, each product page contains a history of reviews for the specific listing from previous customers, much in the style of online marketplaces such as Amazon and eBay. The web pages contain everything a consumer considering a purchase might see. Due to the instability of the TOR network that is required to connect to these online marketplaces, and semi-frequent DDOS attacks on the servers hosting the marketplace, many of these website scrapes are incomplete and often only cover a fraction of the listings on the marketplace on any given day. As a result, the observed time series of any individual product listing may be randomly right or left censored. For example, I might observe a listing in November of 2014 and never see it again, possibly because the vendor took the listing off the market soon after that scrape, or because the vendor left the listing on the site for an additional 2 months, but subsequent scrapes failed to arrive at the URL associated with that listing when crawling the website. The crawling procedure is described in gwern and more or less follows a recursively defined random walk. Because of this, I can treat the listings I do observe as i.i.d. observations from the marketplace, since the random nature of the webcrawl ensures missingness will be unrelated to unobservable characteristics of the product.
The third and final dataset used in this paper is weekly scrapes of the Agora forums. On the Darknet, each marketplace typically has an accompanying forum where buyers and sellers can discuss offerings on the markets. Importantly, this is the main place where customers can discuss whether certain vendors are selling high or low quality items, or whether they are “scams”. Due to the highly illegal and anonymous nature of the marketplace, it is very difficult to externally validate the quality of a listing. Importantly, these forums are virtually the only venue prospective buyers can go to for more information on the product. Consumers cannot go to typical online social networks for fear of being identified as a buyer or seller of illegal products. Since the Agora forums also require a Tor connection to anonymize users, it is a relatively safe area to discuss legitimate questions on products and sellers. The only other known place on the Internet where some discussions of Darknet markets occurs is Reddit, but the discussion is relatively limited for confidentiality concerns. The number of threads from 2014 in the Agora subreddit was 1,551; as a comparison, in this same period, the number of forumn threads in the Agora Forums was 52,058. Even among the relevant subforums I devote my analysis to, that are explicitly designated to contain topics directly related to the marketplace, have 10 times as many threads as this subreddit during this period. Due to the concentrated nature of information on Darknet markets, the combination of the HTML of both the marketplace and the forum means that one can theoretically obtain a complete picture of the information customers would have had access to when deciding on a purchase. For this reason, I can completely characterize the information customers obtain from both formal (reviews) and informal (forum posts) means in this marketplace, which is what makes it such a uniquely interesting market to study.
The forum data has information on the subject text of individual thread, what topic the thread is in (General Discussion, Vendor Discussion,etc.) as well as the text of replies to the thread. In addition, each post is accompanied with information about the author, including the username of the author (which is the same as their username in the Agora marketplace), how active and experienced the user is, whether the author is classified as a seller by Agora, and overall favorability of the author's posts, as measured by “karma”, much like favorability measures employed on social media websites such as Reddit.
Like the marketplace data, these HTML scrapes are often incomplete due to bandwidth limitations of the Tor network; however, because the forums are cumulative (posts are typically not removed after they are posted, and the forum website stores all threads in the history of its existence), I am more likely to observe the vast majority of posts in the data. It is for this reason that I limit my analysis of the Agora marketplace to the calendar year 2014, even though the data extends to July 2015. This will allow me to compile a more complete set of threads in my time period of interest, due to one of the most complete scrapes occurring in January 2015. Since threads are uniquely identified with an iterative integer (i.e. the first thread on the forums would have ID 1, the second thread has ID 2 etc.), I can also directly calculate the percentage of threads I am able to observe in the data. By this calculation, I observe 35,420 of 52,058 (68.04%) of all the threads that have ever occurred on the forum. While some of the missing threads are due to the incompleteness of scrapes, it is likely that a substantial portion of missing threads are due to removal of a thread by moderators (for either being listed in the incorrect topic or violating the rules of the forum; spam is a common issue), so this should be considered a conservative lower bound on my coverage of the forum threads.
The basic unit of observation in this paper will be the time series of each listing on the marketplace. For each item, I observe every time it is sold, via the mandatory review, along with the date the transaction was completed (i.e. when the user reports they have received the product and completes the escrow transaction). From this, I can construct the number of transactions finalized each day, and use this as a proxy for the number of purchases each day. This follows the methodology used to measure sales in the modest literature that has studied Darknet markets. In order to reduce the impact of measurement error on the observed date of a sale, as well as to expedite computations, I aggregate the daily time series of sales data for each item to a weekly time series. I then incorporate vendor and item characteristics (price, average review rating, vendor rating, etc.) to construct a panel dataset of market listings over time, measured in weeks. As mentioned before, these time series will suffer from censoring because of the incomplete nature of the web crawls used to generate the dataset. I cannot be sure that an item that shows up for the first time in the web scrapes was not listed earlier on the marketplace, but simply missed by earlier scrapes. In this way, the listing time series' will suffer from random censoring, since the mechanism by which the series is censored (both left and right censorship are possible) is unrelated to any characteristics of the product itself. Because of this, the effect of the censoring on statistical inference will be limited, and I ignore it's impact on my estimates for coefficients.
Table (ref) contains summary statistics of my panel dataset. As is evident in the table, product sales are relatively sparse (about 1 sale every 3 weeks per item). Across our sample period, we have 50,000 unique product from which to draw inference, sold by 2,000 different vendors. Both vendor ratings and item ratings are extremely high on average. Prices show extraordinary dispersion; this is mostly due to a couple of outlier listings that probably were mistakes by the vendor (possibly meant to put the price in USD terms); when the maximum price is excluded, the standard deviation is reduced by a factor of 10. To avoid these outliers altering the estimations, I log-transform prices whenever they are input into a regression equation. Table (ref) also summarizes the \# of reviews by item / vendor, the average \# of mentions of a vendor on the forums, and the sentiment score variables we use later in the analysis. I now proceed to discuss how I take the rich text data from both reviews and forum threads, and turn it into useful covariates for regression analysis.
In the absence of verifiable information from sellers, buyers must seek alternative means of acquiring information on the quality or veracity of a particular product. I consider two avenues by which consumers acquire information. The first is through customer reviews. Since these reviews are required, by design they should convey a relatively broad cross-section of past customer experiences. In practice, reviews are heavily biased upwards, resulting in a coarsening of information available to consumers. There is an immense pressure on Agora to give a vendor a 5-star rating for a product you purchase, in part because of the potential danger one risks by giving a distributor of illegal goods one's address (even if it is just a nearby P.O. Box). Many explicitly request it on their listing page, much like finalizing early. In fact, in my sample 97.7% of reviews are 5 out 5 stars (one can rate a purchase from 0 to 5 stars). Vendors care a lot about this rating since a vendor's overall rating (an unweighted average of reviews on their products) is embedded into every active product listing by the vendor. In contrast, there are no widely distributed “averages” for the accompanying text required with each review. For this reason, I hypothesize that text, even in reviews with biased ratings, can provide valuable information to a prospective customer. In order to extract this information from the text, I use natural language processing methods to construct a “sentiment” variable associated with each review text that measures the relative positive or negative sentiment associated with a given product review, based on how certain words associate with customer sentiment of their purchase experience. After pre-processing the review text (removing stopwords, non alpha-numeric characters, as well as converting it to lowercase) using the Natural Language Processing Toolkit (NLTK) bird2009natural, I decompose a review text into a matrix of word counts, and normalize the rows to sum to 1 (so the matrix consists of rows of word frequencies in an observed review). I then use multinomial inverse regression (MNIR) taddy2013multinomial to relate these frequencies to positive or negative product sentiment. Specifically, I assume that text for customer reviews is generated from a sequence of i.i.d. draws of words $w_i$ from a multinomial \[w_i \sim Multinomial(\textbf{p}, m_i), p_j = \frac{e^{\alpha_j + \phi_j r_i + \epsilon_ij}}{\Sigma_k e^{\alpha_k + \phi_k r_i + \epsilon_ik}} \] whose probability vector p is a is determined by a latent measure of customer satisfaction of customer $i$, $r_i$. Here, $k$ indexes words across the entire vocabulary considered, while $\alpha_j$ and $\phi_j$ denote word-specific intercept and slope coefficients to relate each word's probability to $r$. I measure customer satisfaction with the star rating of a review each text is associated with. Even though this variable suffers from the bias and coarsening issues I just discussed above, this regression method should be able to identify strongly positive and negative text that accompanied all reviews, and give a better sense of whether, for example, the sample of 5-star reviews are really consistently associated with text as positive as the perfect rating given to vendors.
To calibrate this multinomial model's hyperparameters, I choose 10 regularization paths for the estimation (for computational tractability), and a hyperparameter choice of $\gamma=0$, which amounts to a complexity penalty equivalent to that found in lasso regression. I also choose to drop all words from the estimation procedure that do not appear at least 5 times in the set of $\approx 200,000$ reviews I study. I do this because it reduces the vocabulary I need to consider in the model(from about 23,000 to 6,000 unique stemmed words), which significantly eases the estimation computationally. At the same time, however, these “scarce” words are unlikely to be very informative about consumer sentiment and might result in extreme estimates for the score loadings $\boldsymbol{\phi}$ that are only based on a couple of observations. The choice of 5 words as the cutoff is an arbitrary choice I made that seemed reasonable in the tradeoff between limiting extreme knife-edge estimates for $\boldsymbol{\phi}$, and excessive censoring of text.
After estimating the multiomial model of word draws based on sentiment, I extract score loadings $\boldsymbol{\phi}$ associated with the star-rating for each word and apply these to the observed review text to construct the sufficient reduction projection scores $s_i= \Sigma_{k\in i} w_k\phi_k$, for each review. This score will capture all the available information in the text related to customer sentiment, in the sense that the review rating will now be orthogonal to the review text, conditional on the score variable taddy2013multinomial.
Table (ref) shows the 20 most negative and positive words from the estimation in terms of their associated scores, alongside the observed frequency of each word and the average rating of reviews that contain each of these words. Both the negative and positive words are stems we might expect; 3 out of 20 of the most negative words allude to the product being a “scam”, and the others mostly refer to either dishonesty on the part of the vendor or warn other customers to steer clear. In contrast, the positively scored words appear to signal praise for the vendor, or mention a positive experience with the delivery of the product. Notably, only 4 of the 40 words that occupy the extremes of the score spectrum actually explicitly refer to the quality characteristics of the product itself. The other high-scoring words appear to be concerned with the quality of the vendor themself. This may be because I am estimating sentiment from reviews across all types of product categories. Within each category, different words may be used to describe a product as low or high quality. I would not expect the word “medibud”, for example, to be associated with a positive review of prescription pills. At the same time, all products listed on Agora suffer from common concerns and issues with delivery of the goods or reliability of a vendor. In this sense, my model is limited in that it will likely have difficulty identifying sentiment concerning “product-specific quality” as opposed to “vendor-specific quality” \footnote{one could imagine training a MNIR regression model separately or reviews in each sub-category, but this approach would be limited when comparing to the forum data}, but at the same time this identification strategy will prove useful when I analyze vendor quality sentiment in the forums. Another limitation of my approach will be that I do not account for the estimation error from the MNIR model in my later regression analysis when I use the scores $s_i$ as inputs; I take the output of the model as the true estimates. Correcting for this would involve boostrapping over the sample of reviews, re-estimating the MNIR model, and then performing all subsequent analysis using the random sample of reviews. Due to the intense computational requirements of both the MNIR estimation and the later regression analysis I perform, it is excessively cumbersome to do this analysis multiple times at this stage, and I ignore any error from the MNIR model. For this reason, we might expect standard errors to be biased downward in the later analysis.
The sentiment score $s_i$ is the main covariate used in my analysis. Of the 6,013 score loadings I extract from the multinomial model, about 34% of them have non-zero estimates for their values. This is in part due to my choice of a strict lasso penalty, rather than a semi-concave one, but in addition due to the limited variation in the review rating variable. The vast majority of reviews are 5/5 stars. Nonetheless, there is substantial variation in the score sentiment variables I construct from the text using the MNIR model. Only 7% of reviews I consider have sentiment scores of zero, which would be the case when a review text is “uninformative” in regard to customer sentiment (as measured by not having any words with nonzero score magnitudes). Figure (ref) shows a boxplot of sentiment by the star rating of the review. While there is a clear monotonic relationship by number of stars given, there is substantial variation in sentiment within each star rating as well. In particular, we see that the most variation occurs in the 5-star category, as might be expected since it is by far the most populated. The range of the upper and lower hinges of the 5-star boxplot demonstrate that within 5-star reviews, there is enough variation in text sentiment to cover virtually the entire support of observed sentiment in 1-4 star reviews. This suggests that there is in fact a non-trivial amount of 5-star reviewers whose text aligns more with the language used in low rating reviews.
In addition to analyzing customer review text, I analyze the text of posts in each forum thread on the Agora forums. I now describe the process through which I analyze forum text as a complement to review text. The unit of measurement for analyzing forum text are forum posts which are classified as “mentions” that satisfy one of two requirements. The first is that the text in a post specifically includes a vendor's username in it's body. This is a direct mention. The second is if the post is located within a forum thread of which either:
The idea behind this latter identification strategy is that if a forum's title or first post has a vendor username, I assume the topic of the thread is the mentioned vendor, and specifically the vendor's quality. If a post mentions multiple vendors in the same body of text, for my main analysis, I drop the post, since it is impossible for me to evaluate within a forum post which positive or negative words are directed at which vendor. Later, as a robustness check, I rerun my baseline analysis by counting mentions of multiple vendors in a single post as separate but identical “reviews” of each individual vendor's quality. In practice this does not happen often. In order to reduce the noise I might introduce by including mentions that are not actually discussing with the quality of the vendor, I limit my analysis to forum posts within 4 sub-forums: General Discussion, Vendor Discussion, Product Offers, and Product Categories. The other sub-forums I exclude are: Referral Links, Newbie Section, Generic Randomness, German, Security discussion, Bugs, Philosophy, New features, and News. If forum posts are correctly categorized, it is unlikely that any of these excluded subforums would contain a vendor mention about the vendor's quality in relation to their products or customer service. Additionally, I exclude any vendor whose name is an English word, according to the Natural Language Toolkit's corpus of English words. I do this because it will be impossible to distinguish mentions that refer to the actual vendor in question or happen to just be using the English word the vendor is associated with. In my later analysis, to be consistent, I also throw out all reviews and product listings by vendors whose name is an English word as well, since these vendors may be discussed on the forums but my mention identification strategy would fail to identify them. About 5% of the vendors in my data (the set of users who listed a product at least once in the listings dataset in 2014) have an English word as a name. I also remove one specific user, whose name is “ebay”, who had posts associated with them in the forums that were actually discussing counterfeiting methods on the popular auction website eBay. Finally, I remove from my sample all posts by a vendor where they self-advertise (since consumers will presumably discount this information given the source). This constitutes a very small portion of all observed posts.
Using this definition of vendor mentions, I identify 61,679 unique posts that mention a vendor. Of these, 17,234, or 28%, are a direct mention, and rest are posts in forum threads whose subject line includes a vendor. Some examples of the titles of these forum threads are “californiadreamin vendor. legitimate or not? feedback?”, “beware goingpostal, they are scammers!!!!” and “current whitelist of opiate vendors (updated on the regular)”.
Like all choices, my choice of how to identify vendor mentions in the Agora forums is somewhat arbitrary and has its limitations. I cannot rule out that some threads might veer off-topic and discuss an entirely new subject, unrelated to the vendor's quality. Inversely, there may be some posts that are in fact reviewing a vendor's quality, but fail to name the vendor in their text body (perhaps a reply to a direct mention of a vendor). Additionally, there will be some posts I do not capture simply because a particular thread was never scraped during the data collection process.
A primary interest of this paper is to compare the relative impact of formal and informal sentiment on vendor quality. I take all 61,679 posts flagged as mentioning a specific vendor on the Agora forums, and convert the text to a vendor sentiment variable, analogously to how I did in the review text setting. However, in this setting I have no labeled dependent variables; there is not an associated numerical rating of a vendor with each forum mention. So, I treat these forum posts as coming from the same population as the customer reviews, and apply the MNIR model previously estimated on customer reviews to the forum posts. Specifically, I preprocess all post text as I did with reviews, treat each individual post as a unit of observation, and apply the score loadings $\boldsymbol{\phi}$ I estimated with the labeled review data to the sequence of words $w_p$ contained in each post, generating a vendor sentiment score for each forum post, $fs_p$. To my surprise, there is a smaller fraction of forum posts that are “uninformative” (i.e. $fs_p=0$); 3% of the forum posts are uninformative, while about 7% of formal reviews are uninformative. One can imagine that, even though posts mention a specific vendor, and are in one of the 4 appropriate subforums, some of these flagged posted are discussing an aspect of the vendor that is unrelated to consumer sentiment. There is a substantial amount of variation in the sentiment in forums as well. Figure (ref) shows the density function of informative (non-zero) customer sentiment scores in both the set of forum posts and customer reviews. The density function that includes all consumer sentiment is similar in appearance but with a large spike at zero, which makes the plot more difficult to interpret due to the scaling. The distributions of forum and review sentiment appear comparable in shape, but the distribution of forum sentiment is clearly shifted to the left of review sentiment. An interpretation of this is that sentiment on the forums is more negative towards vendors than on the formal reviews. This suggests that the forums may serve as an unrestricted setting where users can communicate their actual opinion on vendors more freely without the threat of blackmail that might accompany a very negative product review. Or it may be that the set of users who post on forums and those who actually purchase products simply differ in their attitudes or beliefs towards vendors.
Having generated scores $\{s_i\}$ and $\{s_p\}$ for reviews and forum posts, respectively, I (separately) standardize the non-zero scores in each setting to have mean 0 and standard deviation 1\footnote{I specifically excluded the zero-score text from the mean and standard deviation calculations because standardizing these reviews would cause them to take on a non-zero value, which is difficult to interpret since I know these scores are zero because they have no informative text in their messages. note that overall forum and review sentiment variables will still have mean 0 after this transformation, but a lower standard deviation than 1}. Doing so eases the interpretation of coefficient estimates for later analysis, so that a 1 standard deviation increase in (informative) sentiment will result in an additive $\beta $ or multiplicative $e^\beta$ increase in the dependent variable I consider, depending on whether I use a linear or exponential mean model. For each vendor, I then construct an aggregate average vendor sentiment variable for the history of reviews of vendor $v$ up until week $t$: \[\widetilde{s}_{v,t} = \frac{\Sigma_{i=1}^{n_v} s_{i,v} \mathbbm{1}_{\{t(i,v) < t\}}}{\Sigma_{i=1}^{n_v} \mathbbm{1}_{\{t(i,v) < t\}}} \] where $t(i,v)$ maps a review by consumer $i$ on a product sold by vendor $v$ to the week it was posted. Analogously, I construct for each vendor an average sentiment variable from forum mentions; that is, the mean of forum vendor sentiment scores $fs_{p,v}$ until week $t$: \[\widetilde{fs}_{v,t} = \frac{\Sigma_{p=1}^{m_v} fs_{p,v} \mathbbm{1}_{\{t(p,v) < t\}}}{\Sigma_{p=1}^{m_v} \mathbbm{1}_{\{t(p,v) < t\}}}\] where $m_v$ is the total number of forum mentions of a vendor in the sample period. Additionally, I construct total counts of the number of vendor's product reviews / vendor mentions up until time $t$ in the same manner (denoted $\widetilde{n}_{v,t}$ and $\widetilde{m}_{v,t}$, respectively). I choose to construct review sentiment variable's at the vendor level (rather than at the item level) so that average vendor sentiment between the two settings (reviews and forums) can be more directly comparable. Since discussion on forums mostly concerns vendors, rather than the specific items they sell, identifying specific mentions of products would have been challenging and probably yielded significantly less forum posts to work with. Because some vendors may not have a history of reviews or forum mentions at any given point in time, I always include a missing dummy when including these variables in a regression analysis.
Tables (ref) and (ref) contain regression output from regressing (standardized) sentiment in each setting on virtually all observables available. We see, in the review setting, that after adding product reviewed and review week fixed effects, there is monotonically increasing relationship between sentiment scores and stars given in the review (as should be the case). In addition, other interesting relationships include that buyers with a longer transaction history are overall more positive in their review sentiment. Also, the buyer's personal rating (vendors may voluntarily review buyers who buy from them) does not display a clear relationship, except that those without a perfect 5-star rating give more negative text feedback. As may be expected, reviews on no finalize early products are more positive than those that ask customers to finalize early (this effect is estimated from listings that within their listing time switch from no finalize early to finalize early, or vice versa). In regards to forum sentiment, we see that experienced forum members are less likely to be positive in their sentiment. This descriptive result is suggestive of experienced vendors being sophisticated in how they share information with the community: they may know to be “good” to vendors in their product reviews, yet on forums, they are less worried about how a vendor might punish their negative report, and can speak more openly. We find no correlation, after controlling for a variety of fixed effects (username of the poster, vendor they are mentioning, and week the post is published), of karma, a social media similar in spirit to Facebook likes, relating to how positive or negative forum post sentiment is.
I also include in these tables one model regressing number of words in a review or mention on observables. We see that non-5 star reviews on average have 3 more words, which may be expected since a review $<$ 5 stars represents a deviation from the default and so may indicate more thought is being given to these reviews. We also see some evidence of review fatigue: more experienced customers have, on average, less words in their reviews than newer buyers on the market.
A characteristic of a product listing that does not cleanly enter a regression equation is the listing text description of the product, along with the title of the listing. It is easy to see that changes in the description can in fact cause changes in sales: perhaps by editing the text to be less misleading about the product's quality, or making the title more engaging to consumers scrolling through a list of products when searching the Darknet. Unlike the previous section, where I was able to compare reviews against each other because of associated ratings, there are no labels on the listing text, so I cannot say which listings texts represent effective or ineffective advertisement of the product. Thus, I need an alternative methodology to MNIR in order to be able to successfully control for changes in listing text over the lifespan of a given product. As before, I preprocess the text of listings and then transform the title and description text (separately) into word frequency matrices for all of the sample listings. If all the information in the text relevant to consumers is in these word frequency matrices, then I could simply include the word matrices as additional controls in my regression to capture possible changes in sales due to changes in the phrasing of the listing. However, due to the massive size of these word frequency matrices, (14,000 unique word stems in titles, and 71,000 unique word stems in descriptions), this is computationally infeasible. To circumvent this, I decompose the word frequency matrices into principal components (PCs) that can explain a large portion of the variation in the listing text data. By using the lower-dimensional principal components as controls, I can approximately control for the word frequencies. Due the size and sparsity of these word frequency matrices, I use truncated singular value decomposition (TSVD) to retrieve the initial PCs of each listing text description. Furthermore, I do this principal component decomposition on the “joint” title-description matrix that is obtained by merging the title and description word frequency matrices along observations, so that all rows now sum to 2, and each contain the frequencies of words within titles and descriptions separately. The motivation for this is that, rather than extracting, for example, the first 100 PCs of text and description separately, then including the 200 PCs in my regression specification as controls, there are likely to be strong correlations between words present in the title and words present in the item description. So doing a PC decomposition on the joint matrix will require fewer PCs to explain the same amount of variation in the data. I chose to not just lump the title word counts into the counts of description words since, on average, description text is much longer than the title text, so the title text would contribute much less to the estimated principal components. the joint matrix approach I use more equitably weights the principle components explaining variation in the title and description listing text.
This paper limits its inference to the effects of our covariates of interest on expected product demand, which the Poisson distribution is found to be asymptotically robust to. This is the case even if the true underlying demand function differs from Poisson, as long as the true mean function shares the exponential mean functional form cameron2013count. An alternative interpretation is that the true underlying model of product demand has covariates that are multiplicative factors of product demand (as opposed to additive, as is necessary in the linear regression model). This is a fairly general assumption in this setting, and if satisfied, we can consistently estimate the multiplicative parameters $e^\beta$ and their standard errors, robust to misspecification. Given our large sample size (about 600,000 observations of weekly sales data at the item level), I am not too concerned with our estimates deviating greatly from their asymptotic counterparts.
A simple histogram of sales per week by item listing (Figure (ref)) clearly indicates that sales of individual products per week are exponentially distributed, as might be expected. In fact, when we compare the empirical distribution of observed sales against a Poisson distribution with mean parameter $\lambda$ equal to the sample mean of weekly item sales, we see that the Poisson distribution, even on the unconditional empirical distribution, appears to be a fairly good approximation. One notable feature of the data is that while it appears similar in shape to a Poisson distribution with the same mean, it also appears overdispersed, as is evident from the additional mass on the $0$ and $\geq7$ sales bins. Indeed, the sample variance is about 4 times as large as the sample mean (1.21 versus 0.28).
Given information on sales, sentiment, and other product/vendor characteristics, I estimate the relationship between my covariates of interest and product demand via Poisson regression. Based on the unconditional distribution of sales, modeling weekly product sales as Poisson appears to be a reasonably good approximation. In addition to this, I prefer Poisson regression (as opposed to other count data regression models) due to its closed form solution to incorporating fixed effects, which allows me to non-parametrically control for unobservable time-invariant product characteristics. My regression model for the baseline specification is the follows:
where $\mathbf{X_{v,t}}$ indicates a matrix of vendor controls, $\mathbf{X_{v,t}}$ indicates a matrix of product controls, $\delta_t$ and $ \mu_j$ are product and time fixed effects, respectively. I include a variety of controls to remove all supply-side variation available to consumers from the sales data when estimating my model of relating vendor sentiment to product demand. As I have stated before, by including variables for all decisions made by the vendor that are observable to the consumer, I can ensure that a change by the vendor to, for example, flag the listing as no finalize early, is not confounded with the measured effect of vendor sentiment on demand. These include:
By also including time fixed effects (so indicators for each unique week in the dataset), I am able to implicitly control for any changes in aggregate market conditions, such as the stability of the TOR network, the Bitcoin exchange rate, and seasonal changes in sales (on both the demand and supply side). Through controlling all of these alternative channels that may effect sales, I can then cautiously interpret coefficients from regressing sales of a listing on vendor sentiment as responses in demand (since responses in supply that effect a consumer's choice are already conditioned on).
I now introduce the main regression results of the paper. Table (ref) displays estimates from a Poisson panel regression of sales on vendor sentiment in both the forum and customer review setting, along with the (logged) \# of reviews & mentions. The panel regression is done with respect to each product listing (so it includes product fixed effects), along with cluster-robust standard errors at the item level to allow for arbitrary correlations across observations of an individual listing. I iteratively add more controls column-by-column to show the persistence of the relationships that are estimated. The main specification of interest is column (5), where I control for all relevant supply responses, in addition to time fixed effects, so that the coefficients may be interpreted as demand responses to sentiment (i.e. elasticities).
The table shows multiplicative effects of a 1 unit increase in each regressor. Both measures of sentiment are averages of individual reviews/post sentiment that are normalized to have standard deviation 1 and mean 0 (among non-zero scores). The interpretation of the coefficient of, for example, Avg. Vendor Review Sentiment, is that an increase (decrease) in vendor sentiment among all reviews by 1 standard deviation leads to an 8.75% increase (8.05% decrease) in quantity demanded for the offered product. I find in the baseline regression that vendor sentiment on forums and reviews are approximately equal in their influence on demand- an average increase in 1 standard deviation of forum posts mentioning vendors leads to a 10.7% increase in sales, which is not statistically significantly different from the impact of average vendor sentiment in reviews. Additionally, both have the expected sign - more positive vendor sentiment leads to higher demand. This verifies the hypothesis I had entering this paper- forum-based vendor sentiment has a substantial impact on consumer demand, and consumers in fact appear to value this information at least as highly as the information they glean from the text present in product reviews of a vendor. This is somewhat surprising, given the potential for mischaracterizations or lies in mentions on forums, and my somewhat crude method for identifying relevant forum posts. This underscores the potential importance of a designated space for Darknet market participants to freely exchange information on vendors.
Besides the sentiment measures, other variables for the most part have the expected sign: average item rating (on a 0-5 scale) is very positively associated with increased demand, price has a negative multiplicative (insignificant) effect on sales, and flagging one's post as no finalize early has a positive effect. Note that we cannot interpret the coefficient on price and no finalize early as effects on demand, since the response in sales to these changes could be endogenously driven by the vendor's decision to change the price / flag (i.e. low sales lead them to lower the price). Vendor rating, somewhat surprisingly, is insignificant when we include all of our additional controls. This may reflect the fact that the vendor review sentiment variable is a “sufficient statistic” for the vendor's average rating, and so its inclusion in the regression makes the vendor rating variable irrelevant. Conditional on scores $s_i$ , the review ratings $r_i$ should be orthogonal to the words contained in the review text. A more positive interpretation of this finding is that the star ratings in reviews across a vendor's products are informative insofar as they reflect the information contained within the text; no additional demand-relevant information is present in the numerical rating. This cannot be completely true however, given the strongly positive association between sales and the product-specific average rating. However. it may be the case that information on vendor-wide attributes (such as communication, customer service) is transmitted exclusively through review text, while product attributes (such as quality of the good) may not be completely conveyed in the review text. The negative relation of total count on past reviews with future sales is simply capturing the limited stock of vendors. Since past reviews exactly coincide with past sales, vendors with a fixed supply will mechanically have to sell less as past sales accumulate.
Having established the importance of overall average vendor sentiment on product demand, I now look to uncover some heterogeneous effects of vendor sentiment on product demand. Specifically, I test the hypotheses I proposed in section (ref). These hypotheses are formally tested in Table (ref), where I control for all the supply variation I considered in column (5) in Table (ref).
Overall, the largest takeaway from these differential effects is that consumer demand is significantly more influenced by vendor sentiment on the web when one increases the sample from which this average comes from. In this sense, there are definitely “increasing scales” to positive vendor sentiment if it continues to be perpetrated by more and more customers.
While I have shown that within the lifespan of a product listing, vendor changes in product characteristics can be appropriately controlled for, there may be other forms of response by the vendor. Namely, it may be the case that a vendor responds by pulling a listing altogether from the marketplace due to poor sentiment , and this may be distorting my interpretations of coefficients on vendor sentiment as demand responses. To investigate this potential confounder, I regress the total number of listings produced by each vendor on the sentiment variables discussed before, while controlling for vendor attributes. If it is found that vendors are not responsive to customer sentiment in text on the extensive margin, then it is unlikely this will be a significant confounder. Because I am now interested in evaluating weekly number of total listings, I can draw upon the GRAMS dataset in addition to the Agora scrapes. The GRAMS data will contain the uncensored count of listings for each vendor on Agora, so it should produce the more precise estimate of the true relation between vendor sentiment and number of posted listings on the Darknet.
Table (ref) shows the regression output of a Poisson regression of listings count on sentiment variables. I include a version of both fixed and random gamma vendor effects, and I draw my dependent variable from both the GRAMS and Agora HTML datasets. As is evident, there is no statistically significant vendor response in the number of listings by a vendor to text sentiment. Across the board, the regression using GRAMS data show zero effect of sentiment in both forums and reviews on a vendor's number of listings. And while the coefficient on vendor review sentiment on the Agora scrapes is significant at the 10% level, it in fact moves in the opposite direction one expects. The interpretation of the coefficient is that if positive sentiment about a vendor in reviews increases, they actually offer less products, rather than more. Given this counterintuitive coefficient, along with the estimated effect of zero everywhere else, and the noisy nature of the Agora data (since product counts may be incorrectly underestimated due to the web crawl timing out), I discount this finding as an outlier not reflective of the true relationship.
Here, I discuss some validity checks performed to confirm my results are not knife's edge in nature and survive a broad set of specifications. To begin, Tables (ref) & (ref) are the same specification as (ref) & (ref), discussed above, but this time estimated with random gamma-distributed effects instead of non-parametric fixed effects. We see that in this specification, qualitatively results are identical, and the multiplicative effects are on average slightly greater in magnitude. Nothing else substantially changes. In addition, we see that the $\alpha$ parameter of the regression, the multiplier for the variance of the conditional Poisson distribution, is significant and positive. In fact, this specification (Poisson with gamma random effects) is equivalent to a negative binomial regression, which allows for overdispersion in the count data, and so leads to more efficient estimation when overdispersion is present cameron2013count. These tables demonstrate that the main results of the paper are not a function of the Poisson regression assumption that the mean and variance of the conditional distribution are equal. \\ I also include a version of the regressions present in Tables (ref) & (ref) estimated via linear regression (maintaining product fixed effects), in Tables (ref) & (ref). While a linearly additive model for non-negative count data is questionable, these regressions also yield coefficients that are qualitatively similar to those estimated in the tables discussed above. They provide evidence that the estimates I find are not a result of the maximum-likelihood estimation procedure getting stuck in a local mode due to the high dimensionality of the coefficient parameter $\beta$ (I have about 180 different regressors, excluding fixed effects, in my main specification) since OLS yields a closed form solution for the global optimum.
Finally, in Table (ref), I estimate my main specification using the non-exclusive measure of forum mentions. Instead of limiting mentions to those that include only one specific vendor, I assign the posts that mention multiple vendors to each individual vendor as duplicates. The estimated effect is lower under this classification, possibly due to increased measurement error, but nonetheless significant at the 5% level.
I present evidence in this paper that communication between marketplace participants - as captured by sentiment variables on vendor quality- is an important influence of demand for products. While that result is not particularly groundbreaking, I find that consumer demand is equally influenced by communication on both formal and informal networks - namely, product reviews versus community forums. This may come as a surprise since well-regulated product reviews should be more reliable, but the evidence presented here suggests that discussion even in unstructured settings like forums should be an important determinant of product demand. In addition, I find some empirical evidence that a vendor's ability to commit to disclosure, by flagging their listings as no finalize early, dampens the effect of communication on demand. Furthermore, I find strong evidence that product demand is more responsive to customer communication as the number of messages grows, as may be expected in a Bayesian updating framework.
There are some limitations to interpreting these results, however. Since my dataset is based on imperfect web crawls of the Darknet, I cannot perfectly observe when a change in product characteristics occurs. While this missingness in random, it may be the case that I am not perfectly controlling supply side variation, and so cannot completely attribute the effect of sentiment on sales to demand. By aggregating to the weekly level, I attempt to minimize the influence of this, but also introduce new errors: namely, that some customers who bought during a week may have been exposed to a different listing page, depending on when they purchased and when a vendor updated their product page. This paper looks only at the effects on the Agora marketplace, but Agora is functioning alongside other Darknet markets, and their may be important cross-marketplace effects this paper does not explore. Additionally, I excluded from my analysis controlling for variation in the image shown on the product listing page. State-of-the-art machine learning algorithms exist to control for the information contained in an image, but these ultimately proved too cumbersome at this stage to include in this project.
Ultimately, this paper serves as a useful starting point for the study of a market that is exceptional for testing economic theory, due to its highly anonymous nature and independence from real-world regulations, but more work on the economic properties of Darknet markets should follow.