EconBase
← Back to paper

A Framework for Crop Price Forecasting in Emerging Economies by Analyzing the Quality of Time-series Data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

62,688 characters · 17 sections · 42 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Framework for Crop Price Forecasting in Emerging Economies by Analyzing the Quality of Time-series Data

abstractAccuracy of crop price forecasting techniques is important because it enables the supply chain planners and government bodies to take appropriate actions by estimating market factors such as demand and supply. In emerging economies such as India, the crop prices at marketplaces are manually entered every day, which can be prone to human-induced errors like the entry of incorrect data or entry of no data for many days. In addition to such human prone errors, the fluctuations in the prices itself make the creation of stable and robust forecasting solution a challenging task. Considering such complexities in crop price forecasting, in this paper, we present techniques to build robust crop price prediction models considering various features such as (i) historical price and market arrival quantity of crops, (ii) historical weather data that influence crop production and transportation, (iii) data quality-related features obtained by performing statistical analysis. We additionally propose a framework for context-based model selection and retraining considering factors such as model stability, data quality metrics, and trend analysis of crop prices. To show the efficacy of the proposed approach, we show experimental results on two crops - Tomato and Maize for 14 marketplaces in India and demonstrate that the proposed approach not only improves accuracy metrics significantly when compared against the standard forecasting techniques but also provides robust models.
IEEEkeywordsTime-series Data Analysis, Crop Price Prediction, Data Quality of Time-series, Context-based Model Selection

Introduction

India is an agriculture-based country where 54.6% of the total workforce is engaged in agricultural and allied sector activities, accounting for 17.1% of the country’s Gross Value Added (GVA)\footnote{Annual Report 2018-19, Ministry of Agriculture and Farmers Welfare, Government of India}. Hence, it becomes important for the government bodies associated with agriculture to estimate market factors and take suitable actions to benefit the farmers. Therefore, having a robust automated solution, especially in developing countries such as India, not only aids the government in taking decisions in a timely manner but also helps in positively affecting the large demographics. The price of crops is one such market factor that requires the attention of the government. Accurate crop price forecasting can be useful for the government to take proactive steps and decide various policy measures such as adjusting MSP (Minimum Support Price) so that farmers get a decent price for their produce, restricting the export price by imposing an MEP (Minimum Export Price), so that exporters are forced to sell locally, thus bringing down the crop prices. At the same time, it will also be useful for the farmer for making better decisions like when to sell their produce or when to harvest the crop.

The crop prices are affected due to several factors such as the area under cultivation for a particular crop, supply projection, government policies, consumer demands, supply chain aspects of producers for agriculture-based products, etc. Additionally, weather conditions also play an important factor since the majority of agricultural production in India is rain-fed. Therefore, the study of fluctuations in agricultural crop prices is interesting as well as an important problem to solve from the government's perspective.

Apart from the above-stated reasons, agricultural crop price forecasting is quite challenging due to many factors such as data quality issues, unreliability in future weather predictions, high fluctuation present in the historical crop price, crop price variations across neighboring marketplaces, etc. Moreover, the manually recorded data is prone to human-induced errors such as no data or wrong data entered for a certain day. Considering ML/DL based models, with a new price data arrival every day, updating the models might cause stability issues because of quality issues associated with the crop price data.

In this paper, we address the set of challenges present in deploying a real-world crop price prediction solution by evaluating its robustness and reliability over a continuous time-frame. We summarize our contributions as below:

itemize• We present the architecture of end-to-end pipeline for robust crop price prediction by analyzing historical marketplace data, weather data and data quality-related features. • We propose a framework for enabling context-based model selection strategies under different conditions such as identifying context based on data quality metrics, model stability, and trend of historical crop prices. • We experiment with various regression models and show the results for two crops namely, Tomato and Maize for 14 marketplaces in India. We additionally report interesting findings to illustrate the benefits in modeling trend specific models built on marketplace level as well as crop level for a robust crop price forecast.

Related Work

Time-series forecasting has been an active research area and has been studied under various scenarios such as stock price prediction Bao2004ForecastingSP, energy load forecasting Amarasinghe2017DeepNN, traffic forecasting Li2017DiffusionCR, crop yield prediction Shahhosseini2020ForecastingCY, etc. To capture time-series specific properties, various categories of models were proposed such as (i) Univariate models, which can only model the endogenous variables, like MA, ARMA, ARIMA and its variants such as SARIMA (ii) Multivariate models which can model exogenous variables along with the endogenous variables like VAR models along with its variants such as elliptical Qiu2015RobustEO, structured VAR model Melnyk2016EstimatingSV, ARIMAX, SARIMAX, etc. Along with these time-series specific models, many classical machine learning models found their way into the problem of time-series forecasting such as support vector regression KIM2003307, LASSO LI2014996, gaussian processes rsta2011 and even recent deep learning techniques such as LSTMs SiamiNamini2018ForecastingEA. Apart from using single models, various ensemble models OliveiraT14,CerqueiraTPS17,ShenBAW13 were designed and tested to improve the accuracy of time-series forecasting problems.\\ Crop price prediction is one such instantiation of time-series forecasting problem which is being actively pursued by the research community in the recent past. For example, bame_2008 use exponential smoothing, ARIMA and spectral methods for predicting crop prices. Yercan2012 exploit the seasonal properties of the prices and recommend seasonal models like SARIMA for Tomato price prediction. Gloria2013 proposes spline-based interpolation techniques for Tomato price prediction. Nasira2012VegetablePP proposed usage of Back Propagating Neural Networks (BPNN) for forecasting vegetable prices. Hemageetha2013RadialBF modeled the problem of Tomato price forecasting using the radial basis function neural networks. Ouyang2019 proposed the usage of Long- and Short-Term Time-series Network (LSTNet) for modeling the prices of agricultural crops. In addition to these simple models, various ensemble models were also proposed for efficiently forecasting techniques in agriculture-related problems. Shahhosseini2020ForecastingCY used stacking based ensemble models for predicting crop yield. Xiong2018SeasonalFO,taylor2018forecasting decompose time-series data into a seasonal, trend, and remainder components and proposed individual model to forecast the trend, remainder and seasonal components and finally add them to get the final forecast value. However, most of the prior works focus on continuous retraining of the model as opposed to contextual model retraining.

Further, apart from concentrating on building complex models, researchers proposed various additional features that can help the price forecasting problem. chakraborty2016predicting analyzed new events along with the Agmarknet data to predict food prices and reported a significant improvement over the standard ARIMA model. Ma2018AnIP used collaborative filtering for imputing missing data values using data from neighboring marketplaces. madaan2019price performed price forecasting and anomaly detection in the crop prices by training a classifier model on the dataset of news articles covering hoarding related incidents. Zhang2020ForecastingAC proposed various features such as complexity, linearity, stationarity, periodicity in addition to a model based on horizon features to model agricultural prices. A pilot study\footnote{Price Predictions using Machine Learning (AI) for Soyabean and Onion, Atal Bihari Vajpayee Institute of Good Governance and Policy Analysis, 2019} was performed on soybean and onion for price prediction that uses features such as historical yield, production, exchange rate, historical rainfall, etc. Kantanantha2010 uses environmental factors like rainfall and temperature for price prediction of corn and soybean. However, these prior works do not consider features related to data quality aspects.

In our work, we explore the time-series forecasting models proposed in the literature as discussed above in the context of crop price prediction. Specifically, we analyze the impact of data quality of time-series data and properties of the time-series data such as trend on the downstream models and suggest a framework for continuous retraining of models upon arrival of new data and retrieval of models for the best forecast. Additionally, we also study the impact of building models for each market versus each crop and report our findings in the form of quantitative and qualitative results.

figure*[figure* omitted — 504 chars of source]
figure*[figure* omitted — 488 chars of source]

Dataset

In this section, we briefly discuss various factors we use for modeling crop prices and provide necessary evidence to showcase their importance. In our analysis, we choose different regions from Karnataka state in India for analyzing the prices of Tomato and Maize.

Weather Data

Crop prices are affected due to many factors such as (i) bad weather like heavy rainfall may impact the routes \added{used for transporting crops to the marketplaces} (ii) temperature and humidity affect the shelf life of crops (iii) quantity of crops brought to the marketplaces, as a surplus supply generally brings the prices down and vice versa, which is in turn affected by the weather conditions. Therefore it becomes important to consider weather-related information while predicting prices of crops. We analyze weather parameters such as average humidity, total rainfall and average temperature on a daily basis. We use The Weather Channel APIs to access the weather data for the geo-locations of all the marketplaces. Fig. (ref)(a) shows the variations in the weather parameters for Davangere district in Karnataka state for the years 2017 and 2018.

Agmarknet Data

Agmarknet\footnote{https://agmarknet.gov.in/}(Agricultural Marketing Information Network) website is maintained by the Government of India that maintains crop related information such as minimum, maximum, and the modal price per quintal at marketplace level. \added{The Agmarknet website also captures the marketplace arrival volume of each crop which is measured in quintal for every marketplace. Agmarknet website is updated every day except Sundays since marketplaces are closed and there is no arrival of crops on Sundays. Sometimes, even for the weekdays, data is not updated on the website due to some issues resulting in days with missing values.}

Fig. (ref)(b) shows the modal price variations for Tomato and Maize from Kolar and Davangere marketplaces in Karnataka. It can be observed that the variations in prices are much more pronounced in Tomato than in Maize.

Fig. (ref) shows the heat maps\footnote{The missing values have been imputed as provided in (ref)} \added{which visualize the variation in modal price per Quintal of Tomato and Maize} for marketplaces \added{located in different districts in Karnataka.} From the heat maps, we can infer that there is some seasonality in the prices that can be seen from the repeating patterns in any row over a period of 3 and a half years. Also, it can be seen that some of the marketplaces appear to have correlation amongst them.

comment\added{ \subsection{Data Quality Information} As mentioned earlier, since the data is captured manually at each marketplace, it suffers from data quality issues such as missing values or outlier values on certain days. Therefore the Agmarknet data is analyzed to assess its quality and perform suitable processing before we proceed to the task of price prediction. We consider this data quality analysis as an additional source of information to construct new features which can then be provided to the forecasting models. }
figure*[figure* omitted — 267 chars of source]
figure[figure omitted — 206 chars of source]

\added{Proposed Method}

comment\added{In this section, we describe how the data quality and statistical analysis was performed on the Agmarknet data, how the features are represented to the model and finally how the context based model selection strategies were designed to address the shortcomings identified in data quality analysis.}

\added{Given historical crop prices along with weather data and other Agmarknet data, we first perform data preprocessing steps to deal with missing values and outliers. We then perform statistical analysis on the crop price data to obtain insights into the data quality along with trend and variations present in the data. We identify a set of additional features with the help of the insights from the statistical analysis. Two different feature representation techniques are introduced to capture data quality aspects along with Agmarknet data and weather data. The data quality based feature representation and the context-based model selection strategies have been designed to deal with data quality issues in time-series forecasting. Fig. (ref) shows a high-level pipeline of our proposed framework using trend-based model selection strategy.}

comment\deleted{We performed rigorous statistical analysis on the Agmarknet data with two objectives. Firstly, as explained in the previous sections, the Agmarknet data is entered manually into the system and suffers from quality issues such as missing values on certain days, presence of outlier data among others. Secondly, some aspects of time-series data, such as stationarity, cannot be ascertained by simply plotting the data. In this section, we discuss the pre-processing techniques and statistical tests along with the insights that helped us in constructing features in addition to the weather and Agmarknet data discussed in the previous section. Finally, we discuss the feature representations that were used for price forecasting.}

\added{Data Preprocessing and }Statistical Analysis

\added{In the data quality analysis, we primarily focus on the identification of missing values and outliers present in the data. All the days (except Sundays) where there was no entry were considered as missing values in the Agmarknet data and filled using spline based technique junninen2004methods. Outliers in the data were detected using IQR (Interquartile range) method barbato2011features with threshold value as 1, and the outlier values were not modified. We attribute the presence of outliers to human errors while entering the data in a digital system.} Fig. (ref)(a) and (ref)(b) show the fraction of days with missing values and outliers for the considered marketplaces of Tomato and Maize respectively.

\added{In the statistical analysis,} we perform ADF, KPSS tests schlitzer1995testing and compare the test statistics with the 5 percent critical value. Based on these two tests, we infer if the data is non-stationary, strict-stationary, trend-stationary or difference-stationary. Time-series for tomato price is found to be strict stationary while that of maize is found to be non-stationary. \added{Further, the Agmarknet data is decomposed into trend, seasonal and residual components using additive seasonal decomposition techniques.}

commentWe found the trend and seasonality strength by using equations (ref) and (ref). Here $R_t$ and $T_t$ refer to the residual and trend components respectively. \newline \begin{minipage}{.5\linewidth} \begin{equation} F_T = \max\left(0, 1 - \frac{Var(R_t)}{Var(T_t+R_t)}\right) \end{equation} \end{minipage} \begin{minipage}{.5\linewidth} \begin{equation} F_S = \max\left(0, 1 - \frac{Var(R_t)}{Var(S_{t}+R_t)}\right) \end{equation} \end{minipage}

We further analyze \added{the statistical properties of the} residual component or noise. The mean is quite close to zero for all marketplaces which is expected. The average standard deviation \added{of the residual component for} Tomato is 202.43 which is much higher than that of Maize, which is found to be 34.39, \added{indicating that there is} a higher fluctuation in Tomato prices than in Maize.

\added{Data Quality based Features}

\added{Based on the insights from the data quality and statistical analysis, the following list of features is proposed to capture the data quality issues and variations present in the Agmarknet time-series data.}

Missing value flag (M): \added{From the missing value analysis, a missing value flag is added as a feature, encoded as a binary value of 1 for days with missing values and 0 for others.} \newline Outlier flag (O): \added{From outlier analysis, an outlier flag is added as a feature. The flag takes value 1 for all the days which were identified as outliers and 0 for the other days.} \newline Fourier transform based features (FT): FFT (Fast Fourier Transform) is performed on Agmarknet data, thereby transforming it into the frequency domain. Then, higher frequency components are removed from the FFT, following which Inverse FFT is performed to obtain time-series which is used as features fumi2013fourier. In our experiments, we retain first 3, 6, 9, and 100 frequencies to obtain 4 time-series inverse transforms, with the transform with more components closer to the real data. These additional features help in removing the noise component from the data and discover the underlying patterns.\newline Time-related features (T): Information such as day and month of crop arrival are considered as additional features inorder to capture the temporal seasonality in the prices. These features are encoded in numeric form.\newline \textbf{Statistical Indicators (SI)}: Statistical indicators such as moving averages, moving standard deviation of the price time-series are added as a part of the feature set. These indicators are useful in capturing recent variations in time-series for long-term forecasting and are usually used in problems related to stock price predictions. \added{In our experiments, we use the following indicators: 3 and 7 days Moving Averages, Exponential Moving Averages with Center of Mass 0.25 and 0.5, 20 days Moving Standard Deviation and Moving Average Convergence Divergence (MACD), as used in thomas2019time}.

comment\begin{table}[t] \tiny \caption{List of statistical indicators used along with their definitions} \begin{tabular}{|l|l|} \hline MA7 & 7 days moving average\\ \hline MA21 & 21 days moving average \\ \hline 12ema & 12 days exponential moving average\\ \hline 26ema & 26 days exponential moving average\\ \hline MACD & Moving average convergence and divergence (12ema - 26ema)\\ \hline 20std & 20 days moving standard deviation \\ \hline Upper band & MA21+ 2*20std\\ \hline Lower band & MA21- 2*20std \\ \hline Momentum & Price-1 \\ \hline Log momentum & Logarithm of momentum to base 10 \\ \hline EMA & Exponential weighted mean \\ \hline \end{tabular} \end{table}

Feature Representation \added

commentFor training we have parameters called n{_}steps{_}in and n{_}steps{_}out shown as $p$ and $q$ in figure (ref) respectively. n{_}steps{_}in number of days are used to predict the prices for the next n{_}steps{_}out days. The feature representation is prepared as shown in figure (ref). Suppose we have m features represented as $f_1$, $f_2$ to $f_m$ (like arrival quantity, humidity etc.) along with arrival price $c$ for n days. Let us number the days from ${D^1}$ to ${D^n}$. Hence for each day $D^i$, we have feature set as ($x^i_1, x^i_2, ..., x^i_m, c^i $). This feature set is represented as $X\_D^i$ Now we take days from ${D^1}$ to ${D^p}$ and concatenate their features set $X\_D^1$, $X\_D^2$ to $X\_D^p$. This creates a data with m*\textit{n{_}steps{_}in} features and becomes ${X^1}$. For creating label $Y^1$, the crop prices from day ${D^{p+1}}$ to ${D^{p+q}}$ are concatenated. $X^1$ and $Y^1$ constitute one training data point. This method is followed for the entire for the entire dataset ans is shown in figure (ref).
comment\begin{figure} \caption{Figure showing preparation of training and test data} \end{figure}

Typically, historical crop prices are used to obtain the feature representation. In addition to the Agmarknet data ($AG$), we utilize additional information such as weather data ($WD$), and data quality-related features as discussed above ($DQ$) for constructing the final feature representation. All the individual features can be combined into single representation using the two definitions of $\mathcal{F}$ as described below:

enumerate• Feature Concatenation: \added{To construct a feature vector for $i^{th}$ day,} we concatenate the historical $k$ days of attributes and use it to train a regressor model for crop price forecasting. Let $f$ be the operator that creates the feature representation $\mathcal{F}(d_{i})$ for $i^{th}$ day. For $f_{AG}(d_{i, i-k})$ concatenate arrival quantity and crop price from $i^{th}$ day to $(i-k)^{th}$ day. Similarly, we compute the feature representation for weather data ($f_{WD}(d_{i, i-k})$), and feature representation for data quality ($f_{DQ}(d_{i, i-k})$) for $i^{th}$ day. The concatenated features for $i^{th}$ day is computed as: \begin{equation} \mathcal{F}(d_i) = [f_{AG}(d_{i, i-k}), f_{WD}(d_{i, i-k}), f_{DQ}(d_{i, i-k})] \end{equation} \added{The feature representation $\mathcal{F}(d)$ is a combination of WD, DQ, and AG based historical features. The final combined feature representation is of dimension size $k\cdot$$N_{WD}$ + $k\cdot$$N_{DQ}$ + $k\cdot$$N_{AG}$ dimensional, where $N_{WD}$, $N_{DQ}$ and $N_{AG}$ are the dimensions of daily weather data, data quality, and Agmarknet data respectively. The values of $N_{WD}$ is 3, $N_{DQ}$ is 13, and $N_{AG}$ is 2. To show the importance of the constructed feature vector, we use a Multilayer perceptron (MLP) regressor with 1 hidden layer containing 10 neurons with ReLU activation, using feature concatenation-based feature representation. To train the model we use lbfgs solver with a constant learning rate of 0.001 and squared error as the loss function. The MLP model takes as input the feature vector $\mathcal{F}(d_{i})$ and makes a forecast for the next 30 days i.e., ($i+1$) to ($i+30$) days.} \begin{figure}[tb] \centerline \caption{LSTM-based feature encoder for crop price forecasting} \end{figure} • Time-series dependent feature representation: To capture the time-series dependency between the different attributes, we use a stacked LSTM based feature encoder to represent the multivariate data SiamiNamini2018ForecastingEA. \added{The LSTM model has a stack of two LSTM layers to encode the input features and the output from the last time step of the second LSTM layer is considered as an input to a Dense layer which forecasts for 30 days. The LSTM layer encodes the sequential information from the input that captures the historical weather data, data quality-related features, and Agmarknet data through the recurrent network. Fig. (ref) shows the high-level steps for extracting time-series dependent feature representation and use it for crop price forecasting.} \added{In our experiments, the two LSTM layers are 100 and 100 units respectively with each layer having ReLU activation. To train the network, we use Adam optimizer with a learning rate of 0.001 and mean squared error as the loss function.}

\added{Next, we introduce different contexts to retrain and select models considering different factors such as data quality, model stability, and trend analysis. The procedure for model retraining, model management, and version control for different model selection strategies for a given context is also described.}

Context-based Model Selection Strategies

In real-world time-series data analysis, \added{automated way of model selection and model deployment process} is very important. \added{We propose} a framework for context-based model selection strategies for crop price forecasting. We also demonstrate the utility of inferring the context by analyzing the variations in recent crop prices along with issues related to model stability and data quality. Fig. (ref) illustrates a set of different model selection strategies for crop price forecasting.

figure*[figure* omitted — 270 chars of source]
enumerate• Continuous Retraining: The model deployment can be treated as a continuous process rather than deploying a model once. One of the most important factors is to enable continuous model enrichment such that the model captures the recent crop price variations while forecasting long-term crop prices. However, this continuous model enrichment process requires extra processing for model retraining. Furthermore, continuous retraining may get impacted because of data quality issues if there is no supervision. • Model Stability: To address the limitation of the continuous retraining process, we automatically analyze the model performance after performing model retraining. This is achieved by creating a model catalog that stores the set of models along with model performance-related attributes. The best performing crop price forecasting model is identified as $\arg\max_{m_{i} \in \mathbf{M_{c}}} \phi^{acc}(m_i,D^{val})$ where $m_{i}$ represents the a model from the model catalog ($M_{c}$). $\phi^{acc}$ computes the performance metrics on the validation dataset ($D^{val}$). To capture the recent price variations in the model, we put a constraint on the model catalog such that it stores only recently trained models. • Data Quality Metrics: The data quality check is very important for model deployment and the model selection process. We analyze the data quality by identifying outliers, consecutive missing values present in the recent time-series, and decide whether to use this data for retraining the model. This way a continuous condition is enabled to captures the data quality aspects first, and then perform the continuous retraining. This step could also help in reducing the number of continuous model enrichment steps and thereby reducing the computational cost overall in an automated manner. • Trend Analysis: There is a significant body of work in the space of time-series forecasting using trend, seasonality, seasonal variations, random or irregular movements for modeling time-series data. In this work, we determine trend in crop prices to infer the context. We analyze recent crop prices ($CP_{1,...,k}$) for $k$ days and determine whether there is a positive trend ($T^{pos}$) or a negative trend ($ T^{neg}$) by computing the trend score using: \begin{equation} \begin{split} T(CP_{1,...,k})= \begin{cases} T^{pos} & \sum_{i=1}^{k-1} \frac{CP_{i+1} - CP_{i}} {CP_{i}} > 0 \\ \\ T^{neg} & \sum_{i=1}^{k-1} \frac{CP_{i+1} - CP_{i}} {CP_{i}} <= 0 \\ \end{cases} \end{split} \end{equation} Two different model catalogs are maintained to capture the positive trend and the negative trend related information. To predict crop prices, a recently trained model is retrieved from the catalog based on the context \added{as shown in Fig. (ref)} and its effectiveness is discussed in Sec. (ref). The model catalogs are updated based on the trend analysis of the recent time-series data. \begin{comment} \begin{figure} \begin{minipage}[b]{0.24\textwidth} \caption*{(a)} \end{minipage} \begin{minipage}[b]{0.24\textwidth} \caption*{(b)} \end{minipage} \caption{Figure showing the selection of positive and negative trend models. (a) shows a sample crop price time-series. If the trend score of a data-point is negative (marked with red in (b)) a negative trend model is retrieved, otherwise if the trend score is positive (marked with green in (b)), a positive trend model is retrieved. } \end{figure} \end{comment}
figure[figure omitted — 352 chars of source]

\added{ In the next section, we evaluate and compare the different feature representations using continuous retraining strategy. Furthermore, we also show the importance of different context-based model selection strategies such as Model Stability, Data Quality Metrics, and Trend Analysis.}

Experiments

We evaluate our crop price forecasting algorithm for \added{7} marketplaces of Tomato and \added{7} of Maize using the Agmarknet dataset. Table (ref) shows the train-test split of the datasets for Tomato and Maize. To evaluate the robustness of the crop price forecasting algorithms we first build the model using the historical data for 997 days from 01-JAN-2016 to 03-FEB-2019. The model is evaluated for 128 days from 04-FEB-2019 to 02-JUL-2019 and compared against the ground truth values.

table[table omitted — 366 chars of source]

Experimental Setup

To show the effectiveness of proposed features and context-based model selection strategies, we conduct experiments for multiple marketplaces for Tomato. For experimentation, we consider multiple factors such as model characteristics - univariate (ARIMA, SARIMA, and Prophet) vs multivariate (MLP), features used - only crop price vs additional features, model selection for retraining - continuous vs context-based. Following is the list of experiments conducted:

itemize• Baseline Models: To establish baselines, we consider widely used univariate models such as ARIMA yunus2015arima, SARIMA vagropoulos2016comparison, and FB-PROPHET taylor2018forecasting. These models only take the crop price as input to make the forecast. • Importance of Proposed Features: To understand the effect of additional features such as AG, WD, DQ we consider multivariate models. We use feature concatenation method proposed in Sec (ref)(1) to prepare the feature and use MLP to conduct the experiments using continuous model retraining strategy. • Importance of Context-based Model Selection: To show the effectiveness of context-based model selection over continuous model retraining, we conduct multiple experiments. We evaluate different context-based model selection strategies as mentioned in Sec. (ref) such as model selection based on data quality ($M^{d}$), model stability ($M^{s}$), and trend-based model selections either at marketplace level ($T^{m}$) or at crop level ($T^{c}$). In these experiments, we use two different types of feature representations namely, feature concatenation and time-series dependent feature representation as mentioned in Sec (ref).
comment\begin{itemize} • ARIMA, SARIMA, and Prophet as baseline models which take as input only the crop price • MLP model which takes as input AG and WD data along with crop prices as input • MLP model which takes as input AG, WD and DQ data along with crop prices as input \end{itemize} In the second phase, we validate the effectiveness of the proposed context-based model retraining and conduct the following experiments \begin{itemize} • MLP with a continuous retraining • MLP with a data quality based criteria for model retrieval for retraining • MLP with model stability based criteria for model retrieval for retraining • MLP with trend-based criteria for model retrieval for retraining \end{itemize}

To show the generalization capability of the proposed method, we also perform similar experiments for multiple marketplaces for Maize. The metrics used to evaluate the model performance and model robustness are discussed in the following section.

Evaluation Metrics

The performance of forecasting models is measured using Root Mean Square Error (RMSE) and Mean Absolute Percentage Error (MAPE). In addition to measuring forecasting model performance, measuring the robustness and the reliability of the deployed system over time is also a very important criteria.

As mentioned in Table (ref), we initially use 997 days data for training the model and everyday we forecast 30 days ahead crop prices. For each following day, we add the new data point and follow the training and forecasting procedure. Evaluation is done over a period of 128 days and the metrics such as average RMSE and average MAPE are computed to understand the model performance over longer period of time. The average RMSE (AR) and average MAPE (AM) are computed over $p$ days as:

equation[equation omitted — 124 chars of source]
equation[equation omitted — 148 chars of source]

where $\hat{y}_i^{(j)}$ and $y_i^{(j)}$ represent the prediction and the ground truth for $j^{th}$ test data point of $i^{th}$ day respectively. The values of $h$ and $p$ are used as $30$ and $128$ respectively in our experiments. \added{While computing the evaluation metrics, we ignore the forecasted values when there is no ground truth data available because of missing value.}

To compare the reliability and robustness of different models graphically, we show the error variations and cumulative error distribution (CED) curves along with the average RMSE (AR) and average MAPE (AM) values for the different models. The error variation graph shows the variation of error for the forecasts, where high fluctuations in the graph including high peaks suggest that a model is not reliable. CED curve of error values help in comparing the percentage of test data points whose \added{prediction} error falls within a particular error threshold value. A \added{higher} percentage value suggests that the model is more robust.

commentFor the purpose of performance evaluation of our regressor models we used RMSE and MAPE. both of which are discussed in brief below Root mean square error (RMSE) RMSE is used to measure the overall deviation of the predicted values from the actual value. It is mathematically given by Eq. (ref). Here $y_i$ is the actual value, $\hat{y_i}$ is the predicted value of the ith datapoint, and n is the number of predictions. It is easy to see that RMSE would be smaller for the case in which the predicted value is closer to the actual value. Similarly a higher value of RMSE means that there is a higher deviation in predicted values form the actual ones. Mean absolute percentage error (MAPE) It is another measure of prediction accuracy for regression problems. It is used to express accuracy as a percentage of absolute deviation from the actual value and is described by equation (ref) \begin{minipage}{.5\linewidth} \begin{equation} RMSE = \sqrt{\frac{1}{n}\sum_{i=1}^{n}(\hat{y_i} -y_i)^2} \end{equation} \end{minipage} \begin{minipage}{.5\linewidth} \begin{equation} MAPE = \frac{100%}{n}\sum_{t=1}^{n}\left |\frac{y_t- \hat y_t}{y_t}\right| \end{equation} \end{minipage} From section (ref) we understand that a single prediction results in n{_}steps{_}out values making it multi target prediction. Assuming that we make g predictions with each prediction resulting in h number of targets, equation (ref) and (ref) are modified as: \begin{equation} RMSE =\frac{1}{p}\sum_{j=1}^{p}{ \sqrt{\frac{1}{h}\sum_{i=1}^{h}(\hat{y}_i^{(j)} -y_i^{(j)})^2}} \end{equation} \begin{equation} MAPE = \frac{100%}{p}\sum_{j=1}^{p}{\frac{1}{h}\sum_{i=1}^{h}\left |\frac{y_i^{(j)}- \hat y_i^{(j)}}{y_i^{(j)}}\right|} \end{equation}

Importance of Agmarknet, Weather and Data Quality related Features

commentWe performed 30 days ahead price forecasting using ARIMA, SARIMA and Prophettaylor2018forecasting. These univariate models used only the historical modal price information. \footnote{ Optimal parameters (p,d,q) for ARIMA and (p,d,q)(P,D,Q,m) for SARIMA) were obtained through auto_arima API available in python.} We then used MLP regressor under multivariate settings with additional features for weather conditions on the day of arrival(mentioned in (ref)) and arrival_quantity in tonnes. We prepared the feature representation as elaborated in (ref). We chose n{_}steps{_}out as 30 since we are interested in forecasting 30 days ahead prices. For n{_}steps{_}in, we chose 7 based on our initial experiments. The training set was restructured as given in section (ref) to obtain 956 training data points initially. This data was used to train our model was was then fed the next 7 days data to obtain 30 days forecast. As given in (ref) the model catalog is updated every day with the model trained on the updated training data. We performed the 30 days forecasting for 128 test points and recorded the error for each forecasting, with $n$ being 30 in equations (ref) and (ref). We also found the average errors according to equations (ref) and (ref). Table (ref) compares the average RMSE and MAPE error of the models mentioned above. We see that the multivariate model gives lower average error for all the marketplaces when compared to Prophet, ARIMA and SARIMA. Fig. (ref)(a) compares the MAPE error variation of the models for the 128 forecasts. We see that the multivariate model has a lowest error for most of the forecasts which shows that it is more reliable. Fig. (ref)(b) shows the pareto chart for the errors which shows that the multivariate model provides more robust forecasting with the error of the forecasts falling in lower MAPE range. Fig. (ref) shows the plots for marketplace Mysore.

Most of the existing techniques such as ARIMA, SARIMA, and Prophet use only historical time-series data. However, crop prices are highly dependent on the local weather conditions \added{and} market arrival \added{in addition to the} historical crop prices. Our multivariate feature representation captures marketplace specific characteristics such as weather parameters (WD) which capture the variation in the avg. temperature, avg. humidity, and the total rainfall on a daily level, along with historical crop prices and historical market place crop arrival data (AG). We compare our feature \added{concatenation based} representation (Sec. (ref)(1)) with existing techniques which only use the historical crop prices.

The comparative average RMSE and average MAPE evaluation metrics of the models mentioned above are shown in Table (ref) for several marketplaces for Tomato crop. It can be observed that the multivariate models have a lower average error for all the marketplaces when compared to the baseline models. It can also be observed that the addition of data quality related features has improved the performance of the model as compared to using only the Agmarknet and weather-based feature representation.

To qualitatively analyse the effectiveness of the proposed features, we graphically represent the error variation and CED curves in Fig.(ref). From Fig. (ref)(a), it can be observed that for most of the forecasts, AG+WD and AG+WD+DQ models have lower errors when compared to baseline models indicating that models build using proposed features are more reliable. From the CED curve in Fig. (ref)(b), it can be observed that the AG+WD model is robust when compared to baseline models since the forecast errors fall in lower MAPE range. Further, it can also be observed that AG+WD+DQ model is better than the AG+WD model indicating that considering data quality information is helpful.

\added{In the current setup, models are continiously retrained over time. Next, we evaluate various context-based model selection strategies using two different feature representations such as feature concatenation and time-series dependent feature representation.}

table*[table* omitted — 2,750 chars of source]
figure*[figure* omitted — 659 chars of source]

Importance of Context-based Model Selection

Continuous retraining strategy may get adversely affected due to the presence of data quality issues such as outliers in the crop prices. To deal with this, in data quality based model selection $M^{d}$, we dynamically retrieve the best performing model with respect to Avg RMSE from the model catalog when there is an outlier present in the recent data. Due to high fluctuations in the factors affecting crop prices, updating model daily can impact the model stability adversely. To address this issue, we enable stability based model selection ($M^{s}$), where best performing model (described in Sec. (ref)(2)) is retrieved from the catalog daily. The catalog is updated when the new model outperforms previous best performing model. Trend-based model selection happens at marketplace level $T^m$ and crop level $T^c$. In $T^{m}$, we build separate $T^{pos}$ and $T^{neg}$ models for each marketplace whereas in $T^{c}$ we just build one set of $T^{pos}$ and $T^{neg}$ models on the accumulated data of all the marketplaces and forecast. In our experiments, for trend-based model selection we use 7-day price window in Eq. (ref).

table*[table* omitted — 2,825 chars of source]
figure*[figure* omitted — 622 chars of source]

Tab. (ref) compares the evaluation metrics for different context-based model selection strategies using the best performing AG+WD+DQ model. It can be observed that context-based model selection strategies provide robust and reliable price forecasting when compared to Tab. (ref). From Table (ref), it can be observed that $T^{c}$ performs better than $T^{m}$ indicating that the model built on crop level data can better capture the variation in prices across marketplaces and provide better forecasting. We also evaluate the performance of crop level trend-based model selection strategy using time-series dependent feature representation ($T^c$ (LSTM)) and observe that though it performs better than $M^d$ and $M^s$, it is unable to outperform $T^c$ (MLP).

Fig. (ref) illustrates the detailed comparative error analysis of different strategies for Kolar marketplace. We can see from Fig. (ref)(a) that $M^s$ and $M^d$ have a relatively uniform error variations, while $T^c$ is observed to provide relatively lower MAPE% for most of the forecasts. Fig. (ref)(a) gives the CED curve for the aggregated errors of all the marketplaces for Tomato. It can be observed that context-based model selection techniques using trend analysis ($T^{m}$ (MLP), $T^{c}$ (MLP), $T^{c}$ (LSTM)) have a lower AOC (Area over curve) indicating a lower expected errorbi2003regression. Similarly, multivariate model settings (AG+WD+DQ, AG+WD) have lower AOC when compared to baseline models.

\added{

Techniques such as ARIMA, SARIMA, and Prophet require the entire time-series data to perform auto-regression or seasonal decomposition to build a regressor. These models have to be fit for every new data point to make new forecasts which is the same as continuous retraining strategy, hence they do not allow model management.

}

table*[table* omitted — 3,501 chars of source]
figure*[figure* omitted — 534 chars of source]

Generalization

To show the generalization capabilities of the proposed approach, we performed crop price prediction for Maize as well. From Sec. (ref), it can be noted that the residual component in Maize has a lower deviation than Tomato and this can be attributed to the fact that Maize is a cereal crop with a more regulated market than Tomato, which is primarily a cash crop. Even though the baseline models perform reasonably well, we show that proposed approach is helpful even in such scenarios where scope of improvement is less. From Tab. (ref), it can be observed that except for Harihara, trend-based model selection strategies have the lowest error followed by AG+WD+DQ model with continuous retraining. Further, CED curves for aggregated error of all the marketplaces for Maize shown in Fig. (ref)(b) also suggests that the trend-based model selection strategy results in lower forecasting errors.

Conclusions & Future Work

Crop price forecasting has been a well-studied problem in the time-series analysis domain over the years. In our experimentation, it was identified that the widely used time-series forecasting approaches (ARIMA, SARIMA, and Prophet) are affected significantly by the data quality issues, which is a predominant factor in emerging economies such as India. To address these shortcomings, we introduced a novel feature representation that captures the weather condition, historical marketplace arrival, and data quality information for robust crop price forecasting. Furthermore, we have also proposed a framework for model selection based on the context for robust crop price forecasting. Experiments indicate that the trend-based model selection strategy is useful compared to existing techniques especially in the case of highly fluctuating crop prices. In the future work, we want to focus on two aspects -- efficient data quality improvement and enhanced context-based model selection. To mitigate the data quality issues efficiently, we plan to explore techniques such as leveraging information from nearby market place to impute missing values as mentioned in Ma2018AnIP. Similarly, we plan to explore more context-based rules and usage of meta-learning techniques like model stacking to choose the best possible strategy from the proposed strategies and obtain better performance.