Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
110,660 characters · 30 sections · 24 citation commands
The Information Content of Taster's Valuation in Tea Auctions of India
Tea (Camellia sinensis) is a manufactured drink that is consumed across the world. The tea crop has rather specific agro-climatic requirements that are only available in tropical and subtropical climates. Tea production, therefore, is geographically limited to a few areas around the world and is highly sensitive to changes in growing conditions. Majority of the tea producing countries are located in the continent of Asia with China, India, Kenya, Sri Lanka and Vietnam being the top producers (in that order), accounting for around 78% of the world tea production and 73% of exports. World tea production is estimated at over 5 million tonnes in 2015, valued around Rs 1 trillion. It increased by 4 percent to 5.2 million tonnes in 2014 (see FAO1 for further details). China remains the largest tea producing country with an output of 2.1 million tonnes in 2014, accounting for more than 40 percent of the world total, while production in India, the second largest producer, remained flat at 1.2 million tonnes in 2014, contributing around 30% of would production.
The tea industry is one of the oldest organized industries in India with a large network of tea producers, retailers, distributors, auctioneers, exporters and packers. Interestingly, India is also the world's largest consumer of black tea with the domestic market consuming around 1,000 million kg of tea during 2016. India’s annual production of tea is around 1,200 million kgs and the market size is estimated to be approximately Rs 20,000 Crore. The Tea Industry in India also derives its importance by being one of the major foreign exchange earners and for playing a vital role towards employment generation as the industry is highly labour intensive. India exports around 225 million kg of tea and is the fourth largest exporter in the world with Russia being its largest importer. The annual value of tea exports from India is around \$800 million. The other major importers of Indian tea are Iran, UK, Pakistan and UAE.
Tea is heterogeneous both over season and region, and even intra-region. Varieties of tea: 90% of the tea produced in India is of CTC (Cut, Tear and Curl) variety followed by Orthodox (9%) and Green Tea (1%). CTC tea is largely graded as Broken Leaf, Dust and Fannings. Each of these grades has a dozen of sub-grades based on the size of the grain etc. Several factors influence the demand for tea, including the price and income variables, demographics such as age, education, occupation, and cultural background. Apart from consumption, other main drivers of international tea prices are trends and changes in per capita consumption, market access, the potential effects of pests and diseases on production, and changing dynamics between retailers, wholesalers and multinationals (Source: FAO1).
Tea demand is very price sensitive. Price elasticities for black tea vary between -0.32 and -0.80, which means that a 10 percent increase in black tea retail prices will lead to a decline in demand for black tea between 3.2 percent and 8 percent, according to FAO. The average weekly volatility of CTC grade prices in Siliguri is 7%. Though the prices seem to be volatile, the tea prices show a specific pattern. The prices are at a peak as the new crop arrivals begin, with an increase in supplies, the prices witness a decline, lasting till the end of the season.
Due to the variation of quality, tea within the same grades are sold at a very wide range of prices in auctions. For example on any given auction day, say Broken Pekoe (BP) variety would sell at a range starting from as low as Rs 60 to Rs 250 per kg, depending on the producer mark (brand), quality and demand, making standardization difficult. Moreover, a given quality of tea is not available throughout the year. \footnote{The details regarding the tea industry has been collected from FAO1, FAO2 and website FAO3.}
{\bf Overview of e-auctions of tea:} An e-auction is a primary marketing channel for selling tea to the highest bidder. The auction system serves two basic purposes. The first purpose is to facilitate price discovery by bringing the buyers and sellers to a common platform with broker’s intermediation. Buyers bid for lots of tea and each lot is sold to the winning bidder. The second purpose is that the auction system provides a guaranteed transaction protocol for the transaction. The transaction includes activities such as delivery of tea to the warehouse, storing, sampling, bidding and payment (see FAO4 for a discussion).
Since September 2016, the auctions are pan India. It means that a member registered with any tea trade association anywhere in India can directly participate in any e-auctions conducted by Tea Board. Earlier, the buyer registered with local tea association could only buy from the e-auctions taking place in respective centers. The tea auction system brings the buyers and sellers together, to determine the price through interactive competitive bidding on the basis of prior assessment of quality of tea. Manufactured tea is dispatched from various gardens/ estates to the auction centres for sale through the appointed auctioneers, on receipt of which, the warehouse keeper sends an arrival a ‘weighment report’ showing the date of arrival and other details pertaining to the tea including any damage or short receipt from the carriers. The tea is catalogued on the basis of their arrival dates within the framework of the respective Tea Trade Associations, the quantities are determined according to the rate of arrivals at a particular auction centres. Registered buyers, representing both the domestic trade and exporters receive samples of each lot of teas catalogued, which is generally distributed a week ahead of each sale enabling the buyers to taste, inform their principals and receive their orders well in time for sale. The auctioneers taste and value the tea for sale and these valuations are released to the traders. Guidelines for the price levels likely to be established when the tea is sold are formulated on the basis of these valuations and last sale price. J. Thomas & Co. Pvt. Ltd is the largest auctioneer in the world, handling over 200 million kg of tea a year, which is one-third of all tea auctioned in India.
Given the above complexities, the report aims to evaluate the feasibility of automating the pricing process to the extent of dispensing with the manual testing and valuation steps. The next section describes the data set we have used for our analysis. We first discuss the clustering exercise according to Grade and source in Section 3. Section 4 discusses the pattern of volumes over different months of the year. The pattern of salability of different tea lots is investigated in section 5. Section 6 discusses the Price to value ratio to gain some insight into the pricing pattern which finally is used fully in the pricing models developed in Section 7. A detailed analysis of the value - price causality question is made in Section 8. The final comments on the feasibility of automation and future plans are discussed in section 9.
In this report, we have used J-Thomas datasets on the weekly tea details for Kolkata Dust tea, Orthodox tea details, CTC tea details and Darjeeling tea details. Moreover, we have the e-auction statistics as a part and parcel of the dataset. We have used the data in 2018 for modelling, as the training data, and the 2019 data has been used for cross-validation. We begin our initial modelling assuming that the model, conditioned on the relevant factors, does not depend on the year of the auction. We shall later see in the cross-validation procedure, that our predictions are quite satisfactory to assert that our assumption was not falsified.
The e-auction statistics (2018-19) consists of the name of tea leaf type, total lots offered in auctions, total amount sold in packets and in quantity, and average price. For each such tea leaf type, detailed info on the weekly sale, total amount sold in packets and in quantity, and average price has also been provided. An initial overview of the characteristics of the data at hand, has been produced in our_earlier.
In the J-Thomas datasets, we find lot numbers (hence the difference between the maximum and minimum lot number would give us the number of lots offered), the categorical variable: the grade of the tea, number of packages offered, the valuation given by the agency, and finally the auction selling price.
In our dataset, we have 25 types of tea grades available names, namely,
However, there are only three main broad categories for tea dust: D1, PD and PD1 according to Wikipedia wiki. Hence our grades require clustering. Out of the 49 weeks of data we possess, 38 weeks of auctions are from 2018, from the period of January to December. However, since not every week had an auction, we thus have data on weeks 2-9, 18, 20, 22-38, 40-41,43-52. Along with this, the remaining data is on the first 12 weeks in 2019, (except the 11th week, in which no auction took place), which we shall be using for cross-validation. Hence, based on our training set, we club together the tea grades for which there are less than 38 observations in the entire dataset (as we require at least one observation per week on the average.) This clubbing is done by its proximity to the other tea grades, where the proximity is based on the qualitative similarity of the grades. Thus we obtain the following 14 clubbed grades till now.
Due to the only packet of GT Dust in the dataset, and that too remaining unsold, and further, due to its lack of immediate similarity from any of the existing tea grades, we conclude that GT Dust is a very rare category in this dataset, and hence we leave it out for the classification problem for now.
The number of clusters so formed in the previous subsection is still quite large, and we suspect that they have quite an inherent similarity in their characteristics and hence in their market appeal and corresponding auction transactions. To have an idea about this, we form a dissimilarity matrix among the grades to visualize the measure of degree of similarity across the clusters. We form the dissimilarity matrix based on the Volume Weighted Valuations and use the metric: $$\text{Dissimilarity} (d)=2(1-\rho^2)$$ where $\rho$ is the product moment correlation correlation coefficient between these time series of Volume Weighted median Valuations for different grades. Thus we create the dissimilarity matrices between these grades, and obtain Figure (ref), where the darker shade represents a larger value of dissimilarity, while a lighter shade shows that the clusters are similar.
One approach of clustering these is using hierarchical clustering based on this dissimilarity measure, to obtain the following clustering dendrogram in Figure (ref).
Analogously, considering Volume weighted mean valuations, we obtain Figure (ref).
However, note that correlation is invariant to change of scale and origin, hence if a particular grade has even twice a valuation than another, using correlation as a measure of clustering would essentially nullify the effect of even twice the valuation, and would land the two grades into the same cluster, in spite of the fact that the grades are quite distinct in their market characteristics.
Furthermore, we need to incorporate the idea that we have the data for multiple weeks, and the clustering should involve the clubbing for all the weeks combined.
Thus we use the following idea:
Then, we define a new similarity structure, where the $i$-th and the $j$-th grade's similarity is proportional to the number of weeks they have occurred in the same cluster by EM-GMM method. Thus, we form a similarity matrix with $(i,j)$-th entry of the matrix characterizing the similarity in terms of number of co-occurrence weeks. The Mosaic plot corresponding to these similarity matrices is given in Figure (ref) and Figure (ref). Finally based on this similarity matrix, we conduct a hierarchical clustering to obtain the clusters, shown in Figure (ref) and Figure (ref).
Thus we obtain the following 6 clusters based on grade:
while the GT Dust category has been left out due to lack of sufficient data points.
As mentioned before, according to Wikipedia wiki, the three main categories for dust type tea leaf are D1, PD and PD1. However, our clustering algorithm shows that Tea Grade D1 and PD1 are similar in market characteristics by putting them into same cluster. On this note, we consider two different clusterings which might be possible.
To find out which one of the above clusterings would be better for further analysis, we find out the proportion of total variation of both weekly price and weekly valuation which is explained by the clusters. Some of the obtained results for both of the clusterings are given in Table (ref). It was found that, for both weekly price and weekly valuation, about 50% variation of these variables over different lots are explained by the clusters alone. Also, using 8 clusters instead of 6 clusters increases this proportion of explained variation by at most 2%, which is not significant in contrast to the loss of simplicity of subsequent works. This diagnostic suggests us to stick with the original clustering, with 6 clusters as defined previously.
Now, note that, if we just cluster the tea based on their grades, then we are foregoing valuable information about the source from which the tea has been produced. This information would be very relevant for our subsequent discussions, and it would not be a good idea to get rid of it. Valuations about a tea grade depend on gardens they are originating from, as it utilizes the idea about the environment used for their nourishment, soil levels, and other significant factors. Hence the source of the tea dust is of utmost importance. However, there are 238 tea gardens (293 including their Clonal, Royal, Gold and Special variants) from which the tea has originated, and again as above, we suspect that maintaining track of each of the tea gardens would be intractable as well as redundant, as the tea also show similarity in characteristics, a significant factor of which might be geographical proximity. Hence we undergo clustering based on the volume-weighted median to obtain dendrogram given in Figure (ref). These dendrogram has again been clustered using the EM-GMM algorithm and then the time-based similarity matrix as in the preceding section.
Before going into the final source-based clusterings, we need to keep the following things in mind as well:
We attach here the maps of the tea-dust producing districts of West Bengal and Assam to give a visualization of the geographical proximities (in Figure (ref) and Figure (ref) ).
Thus we finally use the following 7 clusters based on the source of the tea dust.
Once the clustering for tea grades and sources are obtained, the next idea would be to create a temporal clustering. The data for 2018, presented in the form of weeks, are put into buckets of months. This is done to capture the seasonal variations in the market characteristics of tea dust. But again the exact week number might be too redundant an information, as the tea dust appearing in the market does not fluctuate as frequently as weeks, but might vary by seasons. Thus, to incorporate such a possibility of market dynamics, we calibrate the data of tea dust grades by months. This gives us the Table (ref). To obtain Table (ref), we find out the number of tea packets offered in a lot and multiply it with the net average weight of the tea packets to obtain the total amount (or volume) of tea offered. Then, for each cluster of grade, we find the proportion of its total volume which is offered during a specified month. This gives us a basis for clustering to analyze the supply side of market dynamics.
The mosaic plot for the above table has been included in Figure (ref) for a better visualization.
For analyzing the auction of tea packets, we should concern ourselves with the proportion of tea packets to be sold, and relate its valuation and several other characteristics to it. A successful attempt at predicting the probability of being sold (or being unsold) of an incoming tea packet based on its Grade, Source, Valuation and the current month, would give an insight for automating the auction process, alongside enabling an opportunity to study the effect of Valuation on determining the market characteristics of tea grades.
Primary inspection is made to see whether any particular type of tea grades are more likely to be sold at the auction than other types. Table (ref) shows the proportion of tea lots being sold and the total number of tea lots offered across different grades.
Note that, Table (ref) shows that Fine variant of a tea grade is offered more rarely than its original variant, and its probability of getting sold at the auction also increases. Also, the Special variant of any tea grade is rarer to be offered than its Fine variant, and for this reason, there is not a significant number of observations to conclude whether it increases or decreases the probability of getting sold at the auction.
To find out whether the probability of being sold significantly depends on the variant of tea grades, we perform a simple one-way Analysis of Variance model with the indicator of being sold as the response variable and the type of variant as our treatment variable. We obtain the results as shown in Table (ref). Clearly, the probability of being sold for Fine variants of tea grades are higher than regular variant, and the p-value is small indicating that there are a significant number of observations to support this. However, the proportion of being sold is possibly lower in Special variants than in Regular ones, but higher p-value indicates that there is not enough evidence to support this.
Similar to this, we tried to find out whether different variants of Source (for example, Clonal, Gold, Royal, etc.) affect the probability of getting sold. Again we perform an Analysis of Variance model, however, with the variant of the tea garden (or Source) as our treatment effect. We obtain the results as shown in Table (ref). From this, we note that if the tea packet has come from a Clonal tea garden, its selling probability is expected to be higher than Regular ones by 0.044, and the smaller value of p-value indicates evidence to support this claim. Similarly, the tea packets produced from the Gold type variant of Garden is expected to be 15% less probable to be sold at the auction. Also, with 95% confidence, we can say that the tea packets produced from the Royal type variant of Garden are 22% less likely to be sold.
To build a predictive model to predict whether a tea packet will be sold at the auction or not, based on its Valuation, Grade, Source, and the current time, we use three competing models.
We divide the total dataset of the year 2018 into training and cross-validation sets, with the training set containing 70% of the samples. The cross-validation set is used to select the model. The dataset of the year 2019 is kept as a testing set, which is used to evaluate the performance of the finally selected model for prediction. The percentage of sold tea lots is kept almost similar for both the training and testing sets about 81%. For each of the sets, we consider each combination of Valuation, Grade, Source, and Month, and obtain the proportion of sold tea lots among all tea lots offered under that combination. On the other hand, the predicted probability of being sold under that combination is estimated from the trained model. Both of these probabilities are visually and analytically compared against each other to assess the performance of the trained model for both sets. For analytical comparison, we use three measures as follows;
We began with attempting to fit a logistic regression model with the full data. The following results came up:
Hence this model fails miserably in predicting whether a packet would be sold.
We use a Generalized Additive Model with a binomial family, with a smooth cubic spline fitted on the Valuation of tea packets as a predictor. The following results came up:
We try using a mixture of logistic regressions to predict the selling potential of the packets. The model, in general, for a mixture of $S$ components, is given by logit_mix (and extended in logit_mix_2, logit_mix_3) \[ H(y|T,\mathbf{x},\mathbf{w},\mathbf{\Theta})=\sum_{s=1}^S \pi_s(\mathbf{w},\mathbf{\alpha})\text{Bi}(y|T,\theta_s(x)) \] where $\mathbf{w}$ stands for the concomitant variables on which the mixing proportions $\pi_s$ depend, $\text{Bi}(y|T,\theta_s(\textbf{x}))$ is the binomial distribution with number of trials equal to $T$ and success probability $\theta_s\in(0,1)$, is modelled by usual logistic modelling; $\text{logit} \left(\theta_s(\mathbf{x})\right)=\mathbf{x}^\top \mathbf{\beta}^s$. The concomitant variable is assumed to have a multinomial logit model, i.e. of the form $$\pi_s(\mathbf{w},\mathbf{\alpha})=\dfrac{e^{\mathbf{w}^\top\mathbf{\alpha_s}}}{\sum_{u=1}^S e^{\mathbf{w}^\top\mathbf{\alpha_u}}} \ \forall s$$
Here, we have used our independent variable $x$ to be the corresponding valuations, $T$ being the number of packets arriving, and success denoting the event that the packet is sold. The concomitant variables are the source, grade cluster and month of the packets.
Here we consider 2 to 5 component mixture for this. Based on the Bayesian Information Criterion, the 3 component mixture of logistic regression is chosen which yields a BIC value 6183.014.
Valuation by experts does provide a significant knowledge of what the final price of the transaction would be. The entire process runs with the base price being set at some fixed proportion of the valuations, and thus the following price of transaction revolves significantly across this measure. In this section, we would attempt to fit distributions over the ratio of price and valuations to account for its variability and shape of distribution curves. We have attempted this exercise with the natural logarithm of the ratio of price and volume, to have full support over the real numbers. Observe that if, $\log\left(\dfrac{X}{Y}\right)\sim N(\mu,\sigma^2)$, then $$\mathbb{E}\left(\dfrac{X}{Y}\right)=\exp\left(\mu+\dfrac{\sigma^2}{2}\right)$$ and $$\mathbb{V}\text{ar}\left(\dfrac{X}{Y}\right)=\exp\left(2\mu+\sigma^2\right)\left(e^{\sigma^2}-1\right)$$ The histograms of the ratio of price and valuations for various grade clusters as obtained from the data are given in Figure (ref) and Figure (ref).
We have attempted fitting a single normal distribution as suggested by the histogram. However, this results in a poor fitting of the data. Hence, we attempt to fit a mixture of two normal distributions to the data. The chi-squared goodness of fit yielded a p-value of 0.2274 and the Kolmogorov-Smirnov statistic yielded a p-value of 0.1588, which is quite satisfactory for our case. Thus we report the distribution of the ratio of price and valuation of Cluster 1 to be a mixture of two log-normal distributions with the following properties
The histogram shows a single modal distribution. We tried fitting a single lognormal distribution, our p-values for the chi-square goodness of fit statistic came out as 0.2665 and for One sample Kolmogorov Smirnov test came out as 0.6908, which suggests a reasonably good fit. Thus we report the ratio for this cluster to follow a an unimodal log-normal distribution with the following properties.
The same pattern follows as in Cluster 2, unimodal from histogram, but a single log-normal fits badly (p-value 2e-04). But fit with mixture of two log-normals give reasonably well fits (p-value 0.3855). Thus we report a mixture of two log-normal distributions with the following properties.
The histogram shows a unimodal and highly leptokurtic structure, thus we attempt to fit a single log-normal distribution. However, it fits badly. A mixture of two log-normal distributions fits reasonably better, yielding a p-value of 0.4998 for Pearsonian goodness of fit test, and a p-value of 0.324 for Kolmogorov Smirnov test. Hence we report a mixture of two log-normal distributions with the following properties. The closeness of the means and smaller variances in both the components explain the reason for a unimodal looking structure in the histogram.
In this case, as the histogram shows a bimodal shape, we try fitting with a 2 component mixture of log-normal distribution. It gives reasonably well p-value, 0.8231 for Pearson's chi-sqaured goodness of fit test and 0.9891 for Kolmogorov Smirnov's test. Hence we report the distribution of the ratio to be a mixture of two log-normal distributions with the following parameters.
The histogram for this cluster is characterized by its unimodality and slight positive skewness. A single log-normal yielded a p-value of 0.0104. Hence, we moved on to a mixture of two log-normal distributions, and this time, the p-value came out to be a higher value of 0.06225. We also tried to fit a three-component mixture of log-normal distributions, however, that does not increase the p-value by a significant amount. Thus we report a mixture of two log-normal distributions with the following properties, to maintain the simplicity of the underlying model.
To model the pricing system of the tea market, we consider modeling the demand side by consideration of the Valuation of tea grades, its grade, source, and the month in which the tea lots are available. On the other hand, to model the supply side of the market, we consider the volume of the tea lots as our main predictor. Therefore, our pricing model should include these variables.
To check whether the variant of the tea gardens (Clonal, Gold, Royal, Special, etc.) should be included in the pricing model, we simply fit a one-way Analysis of Variance model with Price as our response variable and the variant of tea garden as the possible treatment variable. It is found that this factor explains a sum of squares of 842405 with 4 degrees of freedom, yielding an F-statistic value of 102.41 and consequently extremely small p-value. Therefore, based on the data, we find sufficient evidence to incorporate this factor into our pricing model in order to have better predictability.
Before specifying a statistical model for the prediction of Valuation and Price of a tea packet based on several of its characteristics, it is extremely important to understand the limitations of how much we can do first. To this end, it is found that there may be several tea packets with exactly the same characteristics, (coming from the same garden, is of the same grade and comes in the same week), and between them, the Range of their Valuations and Prices are calculated. If we consider the empirical CDF of such all possible ranges, then as seen from Figure (ref), in most cases, those packets are subjected to exactly the same Valuation by the auctioneers, however, the final Prices at they are sold may be very different. This exploration suggests that if we simply do away with Valuation and output a single prediction of Price for tea packets based on its Grade, Source, and Week of the year when it is held for the auction, we can almost hope for $55\%$ accuracy in prediction. However, if we wish to have a $90\%$ accuracy in predicting price, we must allow approximately $30$ rupees of deviation from the actual price. Even with a more robust measure of dispersion, mean deviation about median, the conclusion of this exploration remains the same, as seen from Figure (ref).
Since there are lots of Grades, Source and possible weeks combination provided in the dataset, modeling a different prediction for each of these combinations would require a great number of parameters in the model. However, as we have already performed a reasonable clustering analysis of different tea grades and the source gardens, we wish to explore whether having a prediction for tea grades with similar cluster characteristics will be within the economical tolerance level for the auctioneers. Figure (ref) and (ref) shows the corresponding ECDFs of Ranges and Mean deviations about median of Valuations and Prices, for tea packets sharing same cluster characteristics. As seen from them, such a model that predicts at the cluster level would be inadequate in modeling either the valuation or the price soundly.
For ease of interpretability, we start with a simple linear model for pricing system with the grade, source clusters, month of availability, variant of source garden, the volume of the tea packet, and valuation of the tea packets as our predictors. The model is given by;
The results from this model are summarized in Table (ref).
Comparatively, proceeding with our main objective, we remove the valuation as predictor to see how much it affects our original linear pricing model. In this case, we find that a simple linear model with Price of the tea packets as response variable would make the residuals to be heteroscedastic. Therefore, we apply a variance stabilizing logarithmic transformation and use the natural logarithm of price of tea packets as response variable. Thus the model is given by;
where $\varepsilon\sim N(0,\sigma^2)$ independently and identically distributed. The results obtained are summarized in Table (ref)
Thus the valuation part is significant for the prediction of the prices, and the process cannot be automated by a simple linear or log-linear model approach.
From the simple statistical analysis with a linear model performed above, it seems Valuation is indeed a pertinent variable in the explanation of the Pricing system. However, we could not yet allege that the auctioneers' valuation had a causal impact on the price level, although it seemed to be an indispensable predictor. To evaluate the causal impact of valuation on price, we lay down a part of the causal graph structure in Figure (ref). Note that there may be some causation structure between the variables (for example, not all grades occur in all months, hence there should be an arrow from month to grade). Nevertheless, the variables of interest are the valuation and the final price, and the parents to them are known, assuming we subscribe to a linear understanding of time and causality. Thus all the other variables in the model are parents to both valuation and price. Furthermore, valuation could be a parent of price. But, if the valuation is indeed an 'educated guess' of price as was originally intended, then, since the guess is based on only these variables; conditioned on these variables, price and valuation should be independent trials from (possibly same) distribution. Hence, they should be independent, had valuation not been a cause of price. Hence we aim to test this independence.
Thus, we wish to test whether $$\text{Valuation}\perp\!\!\!\perp \text{Price} \mid (\underbrace{\text{ Source, Grade, Volume, Garden, Month}}_{\text{rest}}) ?$$ However, we are faced with the fact that conditioning on so many variables render fewer data to provide reliable estimates for any inference. Hence, we take a different strategy, which is generally often taken when the conditioning variables are continuous. We model the log of valuations by a linear function of the other variables, and similarly for the log of price. That is
If these models provide a good fit, then we can look at the correlation between the residuals to identify the presence of a causal link between valuation and price. This is because the residuals in both the models are independent of the rest, and for two variables $A, B$ and $C$, we have, $$A\perp\!\!\!\perp B |C \Leftrightarrow [A - f(C)]\perp\!\!\!\perp [B - g(C)] | C \Leftrightarrow A - f(C) \perp\!\!\!\perp B - g(C) $$ if $A-f(C)$ and $B-g(C)$ are independent of $C$ itself. This is akin to testing for the partial correlation between valuation and price to be 0.
The results that we obtain from the data are as follows:
This has widespread implications. First of all, this substantiates that valuation of the auctioneers is not as harmless as providing an educated guess for the price, but rather have a causal impact on the final price. This occurs, as the base price is visible to all potential buyers. Thus, to automate the process, whatever procedure we propose cannot be held against the standard with how the valuation predicts price, as the predictability is nested as a causal impact. Hence, to automate the process, a possible change in the auction mechanism is required, and the premises of checking the success of any such alternate mechanism with this data (where valuation has had a causal impact) would be flawed.
The linear analysis provides substantial evidence in favor of the price- valuation linkage. But this may be confounded with many other interactions present in the model. Recall that we have $6$ clusters for types of tea grades and $7$ clusters for source or tea gardens from where the packet came from. Together, considering their interactions, we have a $6\times 7 = 42$ element matrix. These cells are the hidden states in the model and can be interpreted as the proxies for demands for different types of tea in the market. These states, along with several other predictors will generate a common value about a single grade, garden, and week combination, which shall be the true value of the tea packet to both the auctioneers and the buyers. Finally, these prices and valuations will be characterized based on the common value and shall further be influenced by the particular volume of the lot, which would yield the final observations. The model that describes such a situation most closely is a variant of the Linear dynamical system model, popularly associated with Kalman filter, as described in kalman1960new and kalman1961new.
The mathematical framework is as follows;
where $Z_t$ is a $42\times 1$ vector, denoting the market condition. Here, no intercept term is used, since we wish the matrix $F$ to be interpreted as a transition matrix over the market condition vector $Z_t$. One can simply identify $Z_t$ to be a non-deterministic linear dynamic system. Let, $Z_t$ be denoted symbolically as;
$$ Z_t =
$$
where $a_{ij, t}$ is the latent market condition for the demand of tea in the $i$-th grade cluster and $j$-th source cluster, at time $t$. The next level of the model is;
where $W_{gt}$ is a scalar, which denotes the common value of the tea lot at the combination $g$ = (Garden, Grade) at time $t$, which depends on its previous observation, the current market state $Z_t$ and some exogenous control variables $X_{gt}$. In this case, the vector $G_g$ has a special structure such that;
$$G_g Z_t = \beta_1 a_{i_0, j_0, t} + \beta_2 \sum_{i \neq i_0} a_{i, j_0, t} + \beta_3 \sum_{j \neq j_0} a_{i_0, j, t}$$
where $\beta_1, \beta_2$ and $\beta_3$ are parameters to be estimated. Here, $g$ is a grade and garden combination such that the grade belongs to the $i_0$-th cluster, and the garden belongs to the $j_0$-th cluster of the source. This special structure means that the common value for a tea packet depends on the market condition of demand for that particular type of tea, as well as the market condition of its available substitutes, which shares either the same tea grade or the same source garden, as a potential cause for substitutability.
And finally, we have the model for the observations;
where $y_{igt}$ is the actual bivariate observation of Price and Valuation of $i$-th repeated measure in $g$-th group combination in $t$-th time, while $u_{igt}$ is some more exogenous variables, whose influences are incorporated only in the final stage and;
$$\mathbf{1}_2 =
$$
The reason we need to deviate from the standard Simple Adaptive Control Model meyntweedie (Page 40) or the popularly known Kalman Filter model (which allows for only two indices), is that we have more than one observations which are manifestations of the same state (i.e., three indices), and hence, a direct influence of the states do not account for the variability within the observations of the same state.
In the above specification of our three-stage model, we only observe the variables $y_{igt}, u_{igt}$ and $X_{gt}$, and the latent variables $Z_t, W_{gt}$ are unobservable. Hence, we can characterize the model with the specification of all the parameters and the unobservable latent variables, namely by the list of elements $(Z_t, W_{gt}, F, Q, \Phi_0, \Phi, G_g, H, R, \Gamma, S)$. Unfortunately, the above model is not identifiable, as the new set of elements given by $(\alpha Z_t, W_{gt}, F, \alpha^2 Q, \Phi_0, \Phi, \frac{1}{\alpha} G_g, H, R, \Gamma, S)$, also result in the exact same model. The crucial reason for this unidentifiablity is that the first equation contains no observable variable. For this reason, we require to pose a constraint on the model by specifying $\Vert Z_t \Vert = 1$, i.e. the vector $Z_t$'s are normalized for any $t = 0, 1, \dots T$.
Figure (ref) shows the Causal DAG diagram for the above three-stage latent hierarchical (TSLH) model, for a fixed time point $t$, the defining SCM for this are given by equations (ref), (ref) and (ref). Note that, there are $4$ exogenous variables at the second stage (as noted from the four parameters $H_1, H_2, H_3$ and $H_4$) and only one exogenous variable in the third stage. We use Gibb's sampler to obtain the estimates, as discussed in the following subsection. The description of these parameters, along with the estimated value from the dataset is given in table (ref).
We begin by writing the likelihood for the Three Stage Latent Hierarchical (TSLH) model, upto a proportionality constant.
Before obtaining the individual conditional distributions, the following observation will come in handy:
Note that, when $Y\sim \mathcal{N}_k(\mu,\Sigma)$ as in regression model, then $$f(\mathbf{y})\propto \exp\left(-\dfrac 12Y^\top \Sigma^{-1} Y + Y^\top \Sigma^{-1}\mu \right)$$ Therefore, we can simply identify the normal distribution based on these coefficients $\Sigma^{-1}$ and $\Sigma^{-1}\mu$, which is a reparametrization of the parameters of normal distribution. This reparametrization shall be helpful in identifying the conditional distributions in the subsequent calculations.
We first obtain the conditional distributions for the latent variables, $Z_t$ and $W_{gt}$ respectively.
which is a normal distribution with the above reparametrization where, $\Sigma^{-1}$ is the coefficient of quadratic term, and $\Sigma^{-1}\mu$ is the coefficient of the linear term.
Now, with $W_{gt}$, we have;
This leads to another normal distribution. Next, for the parameters,
$$Q \mid \text{rest} \propto \vert Q\vert^{-T/2} \exp\left[ -\dfrac{1}{2} \text{tr}\left( \left(\sum_t \xi_t \xi_t^{\top} \right) Q^{-1} \right) \right]$$
which leads to $\mathcal{W}^{-1}\left[ \sum_t \xi_t \xi_t^{\top}; T-43 \right]$ distribution, where $\mathcal{W}^{-1}$ is used to denoted Inverse Wishart distribution.
Similarly,
$$R \mid \text{rest} \propto R^{-\sum_{t}N_t /2} \exp\left[ -\dfrac{1}{2} \left(\sum_{g,t} e_{gt} e_{gt}^{\top} \right) R^{-1} \right]$$
which is same as the density function of Inverse Gamma distribution with shape $\alpha = \sum_t \dfrac{N_t}{2} - 1$, scale parameter $\beta = \dfrac{1}{2} \left(\sum_{g,t} e_{gt} e_{gt}^{\top} \right)$, upto a proportionality constant.
And finally,
$$S \mid \text{rest} \propto \vert S\vert^{-N/2} \exp\left[ -\dfrac{1}{2} \text{tr}\left( \left(\sum_{i, g, t} \epsilon_{igt} \epsilon_{igt}^{\top} \right) S^{-1} \right) \right]$$
which again leads to $\mathcal{W}^{-1}\left[ \sum_{i, g, t} \epsilon_{igt} \epsilon_{igt}^{\top}; N-3 \right]$ distribution.
Continuing,
therefore,
$$F \mid \text{rest} \sim \mathcal{MN}_{42\times 42}\left( \left[ \sum_t Z_{t-1}Z_{t-1}^{\top}\right]^{-1} \left[ \sum_t Z_t Z_{t-1}^{\top}\right], I, \left[ \sum_t Z_{t-1}Z_{t-1}^{\top}\right]^{-1} Q \right)$$
where $\mathcal{MN}$ stands for the matrix normal distribution. In other words, since the covariance matrix between the rows of the $F$ is $I$, the identity matrix, hence we can generate the rows of $F$ independently from multivariate normal distributions with mean vectors same as the rows of the mean matrix, and the same covariance matrix $\left[ \sum_t Z_{t-1}Z_{t-1}^{\top}\right]^{-1} Q$.
On a similar note,
Therefore,
$$\Gamma \mid \text{rest} \sim \mathcal{MN}_{2\times 2}\left( \left[\sum_{i, g, t} u_{igt} u_{igt}^{\top}\right]^{-1} \left[\sum_{i, g, t} \left( y_{igt} - W_{gt}\mathbf{1}_2 \right)u_{igt}^{\top}\right], I, \left[\sum_{i, g, t} u_{igt} u_{igt}^{\top}\right]^{-1}S \right)$$
To get conditional distribution of parameters corresponding to stage 2 of the model, we assume that, $G_g Z_t = \beta_1 \tilde{Z}_{1t} + \beta_2 \tilde{Z}_{2t} + \beta_3 \tilde{Z}_{3t}$, with $\tilde{Z}$ being the proper linear combination of latent state $Z_t$ that affects the common value $W_{gt}$. Let us also denote the vector of parameters,
$$\theta =
$$
Based on this, we have;
$$\theta \mid \text{rest} \sim \mathcal{MVN}\left( A^{-1}b; A^{-1}R \right)\text{ where }b = \left( \sum_{g, t} W_{gt}, \sum_{g, t} W_{gt} W_{g(t-1)}, \sum_{g, t} W_{gt} \tilde{Z}_{1t}, \dots \right)^\top $$
and
$$A =
$$
where $\mathcal{MVN}$ denotes the multivariate normal distribution.
The performance of the estimated model has been shown in Figure (ref) and in Figure (ref). Some of the residual diagnostics are shown in Figure (ref) and Figure (ref), from which it is obvious that the residuals in logarithm scale follow an approximate normal distribution, other than some outlying values in the tail, as well as the residuals in the original scale of price, shows a histogram of leptokurtic distribution which closely resembles a lognormal one. The estimated values of the parameters of TSLH model are shown in the table (ref).
To emphasize the goodness of fit for the Three Stage Latent Hierarchical (TSLH) model we obtain the following:
The results we obtain have the following notable implications:
Now, to answer the question about whether there is a significant direct effect from valuation to the price, (i.e. whether the dotted arrow in Figure (ref) exists or not) a very general approach is the conditional independence test, as discussed in pearl2016causal and AnIntroductiontoCausalInference. As shown in Figure (ref), $W_{gt}$ and $u_{igt}$ creates a fork with $\log(\text{Valuation}_{igt})$ and $\log(\text{Price}_{igt})$ nodes in the DAG. However we shall require the following theorem:
There are particularly two remarks to be made relating to the above theorem.
To test the causality from $\log(\text{Valuation}_{igt})$ to $\log(\text{Price}_{igt})$, we shall require a conditional independence test between these two variables conditioned on the value of $W_{gt}$ and $u_{igt}$. In the given model, such independence would hold if and only if $\Gamma_{pv} = 0$. Therefore, in view of the above theorem, using the posterior samples $W_{gt, k}^{(r)}$, we obtain the estimate of $\hat{\Gamma}$ as $1.006317$, and the $95\%$ confidence set turns out to be $(0.981592, 1.020155)$, which does not contain $0$, thereby, showing sufficient evidence against the null hypothesis of conditional independence.
On the other hand, based on the fitted model, let us consider the residuals from Valuation predicting component, and the residuals from Price predicting component (leaving Valuation as an explanatory variable), and denote their product moment correlation as $r_{pv}$. In other words,
$$r_{pv} = \text{cor}\left( \log(\text{Valuation}_{igt}) - W_{gt} - \Gamma_v \log(\text{Volume}_{igt}), \log(\text{Price}_{igt}) - W_{gt} - \Gamma_p \log(\text{Volume}_{igt}) \right)$$
Tracking this correlation where the latent variables and parameters are substituted by posterior samples obtained from the Gibbs sampler, we obtain the posterior mean of $r_{pv}$ as $0.7819911$. In contrast to that, the Pearson's correlation coefficient between the logarithms of those variables valuation and price is $0.9573915$, and the Pearson's correlation coefficient between those variables valuation and price without any transformation is $0.9609977$. Note that, the correlation between the residuals obtained from the causal linear model described before was $0.859$. Therefore, we find that, the linear model was enough to structurally model some of the dependence, while the TSLH model was further able to reduce the correlation by explaining temporal dependence structure within the data. However, there was still unexplained correlation, which was simply a manifestation of the causal relationship between auctioneers' valuation and the ultimate selling price at the auction.
The preceding sections show us the utmost significance of the manual valuation of the tea packets that come in, in predicting the final price level. Thus the hope of automating the entire procedure seems unrealistic.
However, since this valuation is based on the inherent characteristics of the tea dust packets, and the volume of packets that arrive, we strongly believe that the practice of using valuations to set base prices can be done away with. George and Hui, in their paper optimal, provide an ingenious way to estimate demand in the auction market, under the Independent Private Value Model second-price auctions. A generalization of this method, in this regard, to the Common Value (CV) auction_book case, where the optimal symmetric bidding strategies are not the bidder signals themselves, but a monotonic function of their signals, may be helpful in our case. Then, knowing the distribution of the bidder values, an optimal reserve price may be set to maximize the ex-ante expected revenue, which is a function of this distribution. Levin and Smith, in their paper disproof, have shown that under the non-IPV case, the optimal reserve price for the seller converges to her true value - here it's her manufacturing costs. Hence if the pool of bidders grow, then it would be safe for the seller to set the reserve price at her manufacturing costs. Collusion among the bidders is often a very practical problem to ponder about, and most methodologies fail under scenarios not robust to such behavior. For example, in second-price auctions, one source of asymmetric equilibrium, (when the distribution has a support $[0,\omega]$) is for one bidder to bid $\omega$ and the others to bid $0$(or the minimum possible price). This is often realized in real-life scenarios, e.g. spectrum auctions. This is a possible solution here, and given that several auctions occur regularly in this market, bidders can sequentially alternate the role of the highest bidder, and thus can all be better off, at the cost of the seller. Our pricing model provides a way to detect such behavior on the part of the bidders. Since our pricing model with valuations provides an excellent fit to the true prices, this can be used to detect collusion. As in the case of collusion, the final price of the transaction would be low compared to the expected transaction price, a large deviation from the predictions would indicate the presence of such collusion. Our prices, with the estimated parameters, approximately follow a normal distribution, thus a low p-value from this distribution could be used as an indication of collusion of bidders. Further research on the aforesaid aspects could bring exciting breakthroughs in the path of automation.