Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
62,688 characters · 17 sections · 42 citation commands
A Framework for Crop Price Forecasting in Emerging Economies by Analyzing the Quality of Time-series Data
India is an agriculture-based country where 54.6% of the total workforce is engaged in agricultural and allied sector activities, accounting for 17.1% of the country’s Gross Value Added (GVA)\footnote{Annual Report 2018-19, Ministry of Agriculture and Farmers Welfare, Government of India}. Hence, it becomes important for the government bodies associated with agriculture to estimate market factors and take suitable actions to benefit the farmers. Therefore, having a robust automated solution, especially in developing countries such as India, not only aids the government in taking decisions in a timely manner but also helps in positively affecting the large demographics. The price of crops is one such market factor that requires the attention of the government. Accurate crop price forecasting can be useful for the government to take proactive steps and decide various policy measures such as adjusting MSP (Minimum Support Price) so that farmers get a decent price for their produce, restricting the export price by imposing an MEP (Minimum Export Price), so that exporters are forced to sell locally, thus bringing down the crop prices. At the same time, it will also be useful for the farmer for making better decisions like when to sell their produce or when to harvest the crop.
The crop prices are affected due to several factors such as the area under cultivation for a particular crop, supply projection, government policies, consumer demands, supply chain aspects of producers for agriculture-based products, etc. Additionally, weather conditions also play an important factor since the majority of agricultural production in India is rain-fed. Therefore, the study of fluctuations in agricultural crop prices is interesting as well as an important problem to solve from the government's perspective.
Apart from the above-stated reasons, agricultural crop price forecasting is quite challenging due to many factors such as data quality issues, unreliability in future weather predictions, high fluctuation present in the historical crop price, crop price variations across neighboring marketplaces, etc. Moreover, the manually recorded data is prone to human-induced errors such as no data or wrong data entered for a certain day. Considering ML/DL based models, with a new price data arrival every day, updating the models might cause stability issues because of quality issues associated with the crop price data.
In this paper, we address the set of challenges present in deploying a real-world crop price prediction solution by evaluating its robustness and reliability over a continuous time-frame. We summarize our contributions as below:
Time-series forecasting has been an active research area and has been studied under various scenarios such as stock price prediction Bao2004ForecastingSP, energy load forecasting Amarasinghe2017DeepNN, traffic forecasting Li2017DiffusionCR, crop yield prediction Shahhosseini2020ForecastingCY, etc. To capture time-series specific properties, various categories of models were proposed such as (i) Univariate models, which can only model the endogenous variables, like MA, ARMA, ARIMA and its variants such as SARIMA (ii) Multivariate models which can model exogenous variables along with the endogenous variables like VAR models along with its variants such as elliptical Qiu2015RobustEO, structured VAR model Melnyk2016EstimatingSV, ARIMAX, SARIMAX, etc. Along with these time-series specific models, many classical machine learning models found their way into the problem of time-series forecasting such as support vector regression KIM2003307, LASSO LI2014996, gaussian processes rsta2011 and even recent deep learning techniques such as LSTMs SiamiNamini2018ForecastingEA. Apart from using single models, various ensemble models OliveiraT14,CerqueiraTPS17,ShenBAW13 were designed and tested to improve the accuracy of time-series forecasting problems.\\ Crop price prediction is one such instantiation of time-series forecasting problem which is being actively pursued by the research community in the recent past. For example, bame_2008 use exponential smoothing, ARIMA and spectral methods for predicting crop prices. Yercan2012 exploit the seasonal properties of the prices and recommend seasonal models like SARIMA for Tomato price prediction. Gloria2013 proposes spline-based interpolation techniques for Tomato price prediction. Nasira2012VegetablePP proposed usage of Back Propagating Neural Networks (BPNN) for forecasting vegetable prices. Hemageetha2013RadialBF modeled the problem of Tomato price forecasting using the radial basis function neural networks. Ouyang2019 proposed the usage of Long- and Short-Term Time-series Network (LSTNet) for modeling the prices of agricultural crops. In addition to these simple models, various ensemble models were also proposed for efficiently forecasting techniques in agriculture-related problems. Shahhosseini2020ForecastingCY used stacking based ensemble models for predicting crop yield. Xiong2018SeasonalFO,taylor2018forecasting decompose time-series data into a seasonal, trend, and remainder components and proposed individual model to forecast the trend, remainder and seasonal components and finally add them to get the final forecast value. However, most of the prior works focus on continuous retraining of the model as opposed to contextual model retraining.
Further, apart from concentrating on building complex models, researchers proposed various additional features that can help the price forecasting problem. chakraborty2016predicting analyzed new events along with the Agmarknet data to predict food prices and reported a significant improvement over the standard ARIMA model. Ma2018AnIP used collaborative filtering for imputing missing data values using data from neighboring marketplaces. madaan2019price performed price forecasting and anomaly detection in the crop prices by training a classifier model on the dataset of news articles covering hoarding related incidents. Zhang2020ForecastingAC proposed various features such as complexity, linearity, stationarity, periodicity in addition to a model based on horizon features to model agricultural prices. A pilot study\footnote{Price Predictions using Machine Learning (AI) for Soyabean and Onion, Atal Bihari Vajpayee Institute of Good Governance and Policy Analysis, 2019} was performed on soybean and onion for price prediction that uses features such as historical yield, production, exchange rate, historical rainfall, etc. Kantanantha2010 uses environmental factors like rainfall and temperature for price prediction of corn and soybean. However, these prior works do not consider features related to data quality aspects.
In our work, we explore the time-series forecasting models proposed in the literature as discussed above in the context of crop price prediction. Specifically, we analyze the impact of data quality of time-series data and properties of the time-series data such as trend on the downstream models and suggest a framework for continuous retraining of models upon arrival of new data and retrieval of models for the best forecast. Additionally, we also study the impact of building models for each market versus each crop and report our findings in the form of quantitative and qualitative results.
In this section, we briefly discuss various factors we use for modeling crop prices and provide necessary evidence to showcase their importance. In our analysis, we choose different regions from Karnataka state in India for analyzing the prices of Tomato and Maize.
Crop prices are affected due to many factors such as (i) bad weather like heavy rainfall may impact the routes \added{used for transporting crops to the marketplaces} (ii) temperature and humidity affect the shelf life of crops (iii) quantity of crops brought to the marketplaces, as a surplus supply generally brings the prices down and vice versa, which is in turn affected by the weather conditions. Therefore it becomes important to consider weather-related information while predicting prices of crops. We analyze weather parameters such as average humidity, total rainfall and average temperature on a daily basis. We use The Weather Channel APIs to access the weather data for the geo-locations of all the marketplaces. Fig. (ref)(a) shows the variations in the weather parameters for Davangere district in Karnataka state for the years 2017 and 2018.
Agmarknet\footnote{https://agmarknet.gov.in/}(Agricultural Marketing Information Network) website is maintained by the Government of India that maintains crop related information such as minimum, maximum, and the modal price per quintal at marketplace level. \added{The Agmarknet website also captures the marketplace arrival volume of each crop which is measured in quintal for every marketplace. Agmarknet website is updated every day except Sundays since marketplaces are closed and there is no arrival of crops on Sundays. Sometimes, even for the weekdays, data is not updated on the website due to some issues resulting in days with missing values.}
Fig. (ref)(b) shows the modal price variations for Tomato and Maize from Kolar and Davangere marketplaces in Karnataka. It can be observed that the variations in prices are much more pronounced in Tomato than in Maize.
Fig. (ref) shows the heat maps\footnote{The missing values have been imputed as provided in (ref)} \added{which visualize the variation in modal price per Quintal of Tomato and Maize} for marketplaces \added{located in different districts in Karnataka.} From the heat maps, we can infer that there is some seasonality in the prices that can be seen from the repeating patterns in any row over a period of 3 and a half years. Also, it can be seen that some of the marketplaces appear to have correlation amongst them.
\added{Given historical crop prices along with weather data and other Agmarknet data, we first perform data preprocessing steps to deal with missing values and outliers. We then perform statistical analysis on the crop price data to obtain insights into the data quality along with trend and variations present in the data. We identify a set of additional features with the help of the insights from the statistical analysis. Two different feature representation techniques are introduced to capture data quality aspects along with Agmarknet data and weather data. The data quality based feature representation and the context-based model selection strategies have been designed to deal with data quality issues in time-series forecasting. Fig. (ref) shows a high-level pipeline of our proposed framework using trend-based model selection strategy.}
\added{In the data quality analysis, we primarily focus on the identification of missing values and outliers present in the data. All the days (except Sundays) where there was no entry were considered as missing values in the Agmarknet data and filled using spline based technique junninen2004methods. Outliers in the data were detected using IQR (Interquartile range) method barbato2011features with threshold value as 1, and the outlier values were not modified. We attribute the presence of outliers to human errors while entering the data in a digital system.} Fig. (ref)(a) and (ref)(b) show the fraction of days with missing values and outliers for the considered marketplaces of Tomato and Maize respectively.
\added{In the statistical analysis,} we perform ADF, KPSS tests schlitzer1995testing and compare the test statistics with the 5 percent critical value. Based on these two tests, we infer if the data is non-stationary, strict-stationary, trend-stationary or difference-stationary. Time-series for tomato price is found to be strict stationary while that of maize is found to be non-stationary. \added{Further, the Agmarknet data is decomposed into trend, seasonal and residual components using additive seasonal decomposition techniques.}
We further analyze \added{the statistical properties of the} residual component or noise. The mean is quite close to zero for all marketplaces which is expected. The average standard deviation \added{of the residual component for} Tomato is 202.43 which is much higher than that of Maize, which is found to be 34.39, \added{indicating that there is} a higher fluctuation in Tomato prices than in Maize.
\added{Based on the insights from the data quality and statistical analysis, the following list of features is proposed to capture the data quality issues and variations present in the Agmarknet time-series data.}
Missing value flag (M): \added{From the missing value analysis, a missing value flag is added as a feature, encoded as a binary value of 1 for days with missing values and 0 for others.} \newline Outlier flag (O): \added{From outlier analysis, an outlier flag is added as a feature. The flag takes value 1 for all the days which were identified as outliers and 0 for the other days.} \newline Fourier transform based features (FT): FFT (Fast Fourier Transform) is performed on Agmarknet data, thereby transforming it into the frequency domain. Then, higher frequency components are removed from the FFT, following which Inverse FFT is performed to obtain time-series which is used as features fumi2013fourier. In our experiments, we retain first 3, 6, 9, and 100 frequencies to obtain 4 time-series inverse transforms, with the transform with more components closer to the real data. These additional features help in removing the noise component from the data and discover the underlying patterns.\newline Time-related features (T): Information such as day and month of crop arrival are considered as additional features inorder to capture the temporal seasonality in the prices. These features are encoded in numeric form.\newline \textbf{Statistical Indicators (SI)}: Statistical indicators such as moving averages, moving standard deviation of the price time-series are added as a part of the feature set. These indicators are useful in capturing recent variations in time-series for long-term forecasting and are usually used in problems related to stock price predictions. \added{In our experiments, we use the following indicators: 3 and 7 days Moving Averages, Exponential Moving Averages with Center of Mass 0.25 and 0.5, 20 days Moving Standard Deviation and Moving Average Convergence Divergence (MACD), as used in thomas2019time}.
Typically, historical crop prices are used to obtain the feature representation. In addition to the Agmarknet data ($AG$), we utilize additional information such as weather data ($WD$), and data quality-related features as discussed above ($DQ$) for constructing the final feature representation. All the individual features can be combined into single representation using the two definitions of $\mathcal{F}$ as described below:
\added{Next, we introduce different contexts to retrain and select models considering different factors such as data quality, model stability, and trend analysis. The procedure for model retraining, model management, and version control for different model selection strategies for a given context is also described.}
In real-world time-series data analysis, \added{automated way of model selection and model deployment process} is very important. \added{We propose} a framework for context-based model selection strategies for crop price forecasting. We also demonstrate the utility of inferring the context by analyzing the variations in recent crop prices along with issues related to model stability and data quality. Fig. (ref) illustrates a set of different model selection strategies for crop price forecasting.
\added{ In the next section, we evaluate and compare the different feature representations using continuous retraining strategy. Furthermore, we also show the importance of different context-based model selection strategies such as Model Stability, Data Quality Metrics, and Trend Analysis.}
We evaluate our crop price forecasting algorithm for \added{7} marketplaces of Tomato and \added{7} of Maize using the Agmarknet dataset. Table (ref) shows the train-test split of the datasets for Tomato and Maize. To evaluate the robustness of the crop price forecasting algorithms we first build the model using the historical data for 997 days from 01-JAN-2016 to 03-FEB-2019. The model is evaluated for 128 days from 04-FEB-2019 to 02-JUL-2019 and compared against the ground truth values.
To show the effectiveness of proposed features and context-based model selection strategies, we conduct experiments for multiple marketplaces for Tomato. For experimentation, we consider multiple factors such as model characteristics - univariate (ARIMA, SARIMA, and Prophet) vs multivariate (MLP), features used - only crop price vs additional features, model selection for retraining - continuous vs context-based. Following is the list of experiments conducted:
To show the generalization capability of the proposed method, we also perform similar experiments for multiple marketplaces for Maize. The metrics used to evaluate the model performance and model robustness are discussed in the following section.
The performance of forecasting models is measured using Root Mean Square Error (RMSE) and Mean Absolute Percentage Error (MAPE). In addition to measuring forecasting model performance, measuring the robustness and the reliability of the deployed system over time is also a very important criteria.
As mentioned in Table (ref), we initially use 997 days data for training the model and everyday we forecast 30 days ahead crop prices. For each following day, we add the new data point and follow the training and forecasting procedure. Evaluation is done over a period of 128 days and the metrics such as average RMSE and average MAPE are computed to understand the model performance over longer period of time. The average RMSE (AR) and average MAPE (AM) are computed over $p$ days as:
where $\hat{y}_i^{(j)}$ and $y_i^{(j)}$ represent the prediction and the ground truth for $j^{th}$ test data point of $i^{th}$ day respectively. The values of $h$ and $p$ are used as $30$ and $128$ respectively in our experiments. \added{While computing the evaluation metrics, we ignore the forecasted values when there is no ground truth data available because of missing value.}
To compare the reliability and robustness of different models graphically, we show the error variations and cumulative error distribution (CED) curves along with the average RMSE (AR) and average MAPE (AM) values for the different models. The error variation graph shows the variation of error for the forecasts, where high fluctuations in the graph including high peaks suggest that a model is not reliable. CED curve of error values help in comparing the percentage of test data points whose \added{prediction} error falls within a particular error threshold value. A \added{higher} percentage value suggests that the model is more robust.
Most of the existing techniques such as ARIMA, SARIMA, and Prophet use only historical time-series data. However, crop prices are highly dependent on the local weather conditions \added{and} market arrival \added{in addition to the} historical crop prices. Our multivariate feature representation captures marketplace specific characteristics such as weather parameters (WD) which capture the variation in the avg. temperature, avg. humidity, and the total rainfall on a daily level, along with historical crop prices and historical market place crop arrival data (AG). We compare our feature \added{concatenation based} representation (Sec. (ref)(1)) with existing techniques which only use the historical crop prices.
The comparative average RMSE and average MAPE evaluation metrics of the models mentioned above are shown in Table (ref) for several marketplaces for Tomato crop. It can be observed that the multivariate models have a lower average error for all the marketplaces when compared to the baseline models. It can also be observed that the addition of data quality related features has improved the performance of the model as compared to using only the Agmarknet and weather-based feature representation.
To qualitatively analyse the effectiveness of the proposed features, we graphically represent the error variation and CED curves in Fig.(ref). From Fig. (ref)(a), it can be observed that for most of the forecasts, AG+WD and AG+WD+DQ models have lower errors when compared to baseline models indicating that models build using proposed features are more reliable. From the CED curve in Fig. (ref)(b), it can be observed that the AG+WD model is robust when compared to baseline models since the forecast errors fall in lower MAPE range. Further, it can also be observed that AG+WD+DQ model is better than the AG+WD model indicating that considering data quality information is helpful.
\added{In the current setup, models are continiously retrained over time. Next, we evaluate various context-based model selection strategies using two different feature representations such as feature concatenation and time-series dependent feature representation.}
Continuous retraining strategy may get adversely affected due to the presence of data quality issues such as outliers in the crop prices. To deal with this, in data quality based model selection $M^{d}$, we dynamically retrieve the best performing model with respect to Avg RMSE from the model catalog when there is an outlier present in the recent data. Due to high fluctuations in the factors affecting crop prices, updating model daily can impact the model stability adversely. To address this issue, we enable stability based model selection ($M^{s}$), where best performing model (described in Sec. (ref)(2)) is retrieved from the catalog daily. The catalog is updated when the new model outperforms previous best performing model. Trend-based model selection happens at marketplace level $T^m$ and crop level $T^c$. In $T^{m}$, we build separate $T^{pos}$ and $T^{neg}$ models for each marketplace whereas in $T^{c}$ we just build one set of $T^{pos}$ and $T^{neg}$ models on the accumulated data of all the marketplaces and forecast. In our experiments, for trend-based model selection we use 7-day price window in Eq. (ref).
Tab. (ref) compares the evaluation metrics for different context-based model selection strategies using the best performing AG+WD+DQ model. It can be observed that context-based model selection strategies provide robust and reliable price forecasting when compared to Tab. (ref). From Table (ref), it can be observed that $T^{c}$ performs better than $T^{m}$ indicating that the model built on crop level data can better capture the variation in prices across marketplaces and provide better forecasting. We also evaluate the performance of crop level trend-based model selection strategy using time-series dependent feature representation ($T^c$ (LSTM)) and observe that though it performs better than $M^d$ and $M^s$, it is unable to outperform $T^c$ (MLP).
Fig. (ref) illustrates the detailed comparative error analysis of different strategies for Kolar marketplace. We can see from Fig. (ref)(a) that $M^s$ and $M^d$ have a relatively uniform error variations, while $T^c$ is observed to provide relatively lower MAPE% for most of the forecasts. Fig. (ref)(a) gives the CED curve for the aggregated errors of all the marketplaces for Tomato. It can be observed that context-based model selection techniques using trend analysis ($T^{m}$ (MLP), $T^{c}$ (MLP), $T^{c}$ (LSTM)) have a lower AOC (Area over curve) indicating a lower expected errorbi2003regression. Similarly, multivariate model settings (AG+WD+DQ, AG+WD) have lower AOC when compared to baseline models.
\added{
Techniques such as ARIMA, SARIMA, and Prophet require the entire time-series data to perform auto-regression or seasonal decomposition to build a regressor. These models have to be fit for every new data point to make new forecasts which is the same as continuous retraining strategy, hence they do not allow model management.
}
To show the generalization capabilities of the proposed approach, we performed crop price prediction for Maize as well. From Sec. (ref), it can be noted that the residual component in Maize has a lower deviation than Tomato and this can be attributed to the fact that Maize is a cereal crop with a more regulated market than Tomato, which is primarily a cash crop. Even though the baseline models perform reasonably well, we show that proposed approach is helpful even in such scenarios where scope of improvement is less. From Tab. (ref), it can be observed that except for Harihara, trend-based model selection strategies have the lowest error followed by AG+WD+DQ model with continuous retraining. Further, CED curves for aggregated error of all the marketplaces for Maize shown in Fig. (ref)(b) also suggests that the trend-based model selection strategy results in lower forecasting errors.
Crop price forecasting has been a well-studied problem in the time-series analysis domain over the years. In our experimentation, it was identified that the widely used time-series forecasting approaches (ARIMA, SARIMA, and Prophet) are affected significantly by the data quality issues, which is a predominant factor in emerging economies such as India. To address these shortcomings, we introduced a novel feature representation that captures the weather condition, historical marketplace arrival, and data quality information for robust crop price forecasting. Furthermore, we have also proposed a framework for model selection based on the context for robust crop price forecasting. Experiments indicate that the trend-based model selection strategy is useful compared to existing techniques especially in the case of highly fluctuating crop prices. In the future work, we want to focus on two aspects -- efficient data quality improvement and enhanced context-based model selection. To mitigate the data quality issues efficiently, we plan to explore techniques such as leveraging information from nearby market place to impute missing values as mentioned in Ma2018AnIP. Similarly, we plan to explore more context-based rules and usage of meta-learning techniques like model stacking to choose the best possible strategy from the proposed strategies and obtain better performance.