Determining the best length of the history of your timeseries data for timeseries forecasting
Introduction
This article examines the influence of the length of the available data history of a time series on the quality of the forecast for future periods. The results of simulation studies that are presented in this article address the importance of long time histories to performing good time series forecasting.
The results shown in this article are based on a simulation study where the available time history is continuously increased and the respective model quality is analyzed. Thus, the relationship between length of time history and model quality can be studied in a visual way.
Additionally, for each of the time series and the different simulation repetitions per time series, the length of the time series that delivers the best forecast quality is selected. Based on these results, the distribution of the optimal length of the time history over all-time series examples is shown and discussed.
The general assumption is that with increasing time history, the quality of the forecast model, which is trained on these data, improves (expressed by a decrease in MAPE). In general, you can expect to see a relationship between increasing data quantity (length of available time history) and forecast quality.
Presenting the conclusion in advance
This article shows results based on simulation studies on how different available lengths of time series data affect the forecast quality. The results correspond to intuition that, on average, a longer time history improves forecast quality.
In addition to these results, the optimal length of the time history is also analyzed. It will be shown that the individual optimal lengths of the time history can also be very short (from 1 to 12 months). This is especially true for time series where the underlying patterns change frequently and quickly and, therefore, the more recent history is more important than the older history.
Simulation procedure
The simulation procedure is as follows:
- For all time series that are available for analysis, the time history is truncated to the length of 1. Based on this 1 value data, a forecast for the next 12 periods is performed and evaluated against the true (untouched) data from the out-of-sample data.
- In the next step, the available time history is increased backwards to the last 2 months. Again, the forecast quality is assessed on the next 12 months.
- This procedure is iterated by adding an additional historic month to each iteration until a maximum data history of 48 months is reached.
This procedure is also iterated for different start points of the time series. This means that the above setup is “shifted” in the time series to avoid studying results which only depend on one start point.
The available length of the data history
The simulation procedure described above has been run on 788 time series from different industries, leisure, retail, steel manufacturing, paper production, and oil and gas production.
The outcomes from this simulation, on how the quality of the forecast model changes over the available history length, are shown in a line chart.
The general assumption is that with increasing time history, the quality of the forecast model, which is trained on these data, improves (expressed by a decrease in MAPE). In general, we expect to see a relationship between increasing data quantity (length of available time history) and forecast quality.
The X-axis represents the respective number of available history months. The Y-axis represents the median MAPE value over all-time series and all shift scenarios.
The graph below shows the same line plot as above; however, it has a different Y-axis scale that allows a more detailed insight into the characteristics of the line. It more truly visualizes the relative change in MAPE with the increasing time history.
From the graph, you can see that the MAPE decreases with more available months in the time history. Studying the course of the line in more detail, you can see the following features:
- There is a steep descent from month 2 to 4, which shows that at least a couple of months are necessary to be able to forecast the mean level of the series to some extent.
- The line decreases further until month 12, showing that the relative contribution of each additional month plays an important role here because only a short time history is available.
- From month 12 to month 13, again a steep descent can be seen, when the first repetition of the seasonal cycle starts. Starting from here, the seasonal specifics of the time series can be integrated into the model.
- The MAPE line decreases further, showing again a decrease after month 24 where a full second cycle is available in the data.
- This repeats after the third full season at month 36. The potential increase at the very edge of the X-axis might also be due to a data artifact and should not be important.
The analysis of the interquartile range of the MAPE statistics per length of time history also shows a decrease of the variability (6 months: 14.7%, 18 months: 12.9%, 30 months: 12.5%, 42 months: 12.2%). This indicates that with increasing length of the history the forecast error decreases on average and the stability of the results increases.
Interpretation
The results shown in the line plots here conform to intuition that, on average, with increasing data quantity in terms of available time history, the quality of the forecasting model increases.
It can also be seen that the availability of the two first full years indicates a strong contribution to model quality. Additional months still provide quality improvement, which is, however, relatively lower. Also, this finding is intuitive: the marginal effect on model quality is higher when only a few months are available for analysis.
Business case calculation
A fictional reference company is used to calculate a business case for the outcome of the different simulation scenarios. The change in forecast quality is transferred into a quantification of the avoidance of over- and under-forecasting and the respective profit is expressed in US dollars.
This is done to illustrate the effect of different data quality changes. The numbers in US dollars should be considered only as rough indications based on the assumption of the simulation scenarios and on the business case as described. In individual cases, these values and relationships are somewhat different. The business case allows you to compare the outcomes in different simulation scenarios and make the results and the findings much more visible.
The reference company Quality DataCom
The reference company “ Quality DataCom” operates in the communications business. The company has a business line that produces and sells electronic accessories for end user devices for both wholesale and retail sales. The company makes 10,000 products that are sold to customers. For each of these products, 1,000 units on average are sold per month. This leads to a total number of 10,000,000 units sold per month.
Quality DataCom currently has a mean absolute percentage error (MAPE) rate of 15%. This means that the forecast on average deviates by 1,500,000 units in total.
The benefit of improving forecast quality is to avoid out-of-stock situations, where a profit of $1 is lost per unit. Also overforecasting causes a loss of $1 due to excessive stock keeping and transportation as well as producing non-sellable goods.
Thus, the decrease in MAPE by one percentage point from 15% to 14% generates an additional profit for this company of 10,000,000 units multiplied by the difference in forecast errors of 1% equals a reduction of the forecast error by 100,000 units. Multiplying this by $1 results in a monthly profit of $100,000. Thus, a decrease in MAPE by 1 percentage point results in a profit of $1,200,000 per year.
Quantifying the value of a longer data history
Relating these results to the reference company shows that for each additional available month in time history, an additional monthly profit of $6,350 can be gained.
The table shows the averaged MAPE values plus the financial impact for the fictional reference company Quality DataCom. The values of the business case are shown in absolute values relative to the period of 1–6 months. Thus, having a time period of 7–12 months on average identifies the additional achieved MAPE of $61,561 in the business case. The last column shows the marginal gain when moving from one availability range to the other.
Analyzing the optimal length of the available time history
Simulation procedure
For each time series, forecasts on different lengths of the time history have been trained. Based on these results, for each time series the optimal history length that best forecasts for future periods can be identified.
This results in a distribution of time histories that deliver the model with the smallest MAPE for the particular time history and shift.
From the bar chart, you can see that
- The optimal history values range from 2 to 48, where each individual history length is represented with a reasonable frequency.
- The spikes at the start of the second, third, and fourth years of history indicate that in some of the simulations the availability of the information for the next seasonal cycle is important for the model.
There is no clear indication that the best forecasts are always performed with the longest available history.
- There is a (surprisingly) high accumulation of values in the range from 2 to 12. This means that there is a large proportion of about a third of the simulation cases (35.68%) where a time history of 12 months or less resulted in the best forecast.
- From the data in the table, you can see that in half (49.91%) of the simulation cases, the optimal length of the time history is less than 1.5 years. Around one-third of the simulation cases achieve their best results with a time history longer than 2.5 years.
Interpretation
The results closely mirror the background of the simulation data. For simulation cases where the optimal length is relatively short (up to 12 or 18 months), additional information from older time periods does not improve the forecast quality. Instead, it downgrades it.
From a business point of view, this is often the case with time series that are taken from retail sales, where the demand pattern shows a fluctuating behavior. Here, only the most recent months describe the most recent behavior. In this case, the more recent picture of the data in months 1 through 12 outperforms the additional benefit that would eventually be gained by repeating the seasonal cycle in the data.
Conclusion
This article has shown results based on simulation studies on how different available lengths of time series data affect the forecast quality. The results shown in the line chart correspond to intuition that, on average, a longer time history improves forecast quality. The graph additionally shows the decline in the marginal benefit of additional months if a longer time history is already available. Steps at the completion of another seasonal cycle can also be seen. They correspond to intuition as well.
In addition to these results, the optimal length of the time history was also analyzed. It has been shown that the individual optimal lengths of the time history can also be very short (from 1 to 12 months). This is especially true for time series where the underlying patterns change frequently and quickly and, therefore, the more recent history is more important than the older history.
Data relevancy
This finding also refers to the topic of “data relevancy”. Having the relevant data available for the analysis is also a feature of data quality. Data relevancy, however, is in most cases understood as the need to have the relevant data available like having access to additional data sources. Often the data quality problem is “not having the relevant data.”
The problem of data relevancy encountered here is somewhat different. With the length of the time history and the information from older periods, it might be that too much data and information are available, which is not as relevant for forecasting the future as more recent data.
It is, therefore, a task for the analyst to make sure that only the relevant segment is being used from the available data. While this is a typical task in predictive modeling where variable selection is performed, it is only rarely the case when the optimal length of the time history for time series forecasting is studied.
This article emphasized the need to carefully analyze the optimal length of the time history in time series forecasting. In order to facilitate this for individual analysis data, a part of the program that has been used to run these simulations is provided in the appendix D and can be downloaded.
Self-assessment of time series data
Refer to the SAS Communities article: Determine the best length of the history of your timeseries data using the %TS_HISTORY_CHECK macro where you find a SAS macro that allows to assess your individual time series data.
This macro can be used to analyze the optimal length of the history as shown above and the line chart for the relationship between the length of data history and the forecast quality.
Links and Downloads
Quantifying the Effect of Different Lengths of Data History in Time Series Forecasting
The data preparation for data science webinar contains more contributions around this topic.
SAS Communities article: Determine the best length of the history of your timeseries data using the %TS_HISTORY_CHECK macro
Webinar: Getting More Insight into Your Forecast Errors using Multivariate Statistics
Presentation #102 in my slide collection contains more visuals on this topic.
Chapters 20 and 21 in my SAS Press book “Data Quality for Analytics Using SAS” discuss these topics in more detail.
