From Models to Predictions
Fitting a model is only half the job. The practical goal of time series analysis is almost always forecasting: using a model calibrated on past data to generate predictions about future values. In this lesson we cover how to produce and evaluate those forecasts rigorously — including point forecasts, prediction intervals, accuracy metrics, exponential smoothing alternatives, and the modern Prophet library.
Forecasting is as much an art as a science. A good forecaster understands the assumptions behind each method, knows how to diagnose forecast failure, and chooses the right tool for the horizon and data frequency at hand.
Point forecast — A single best-guess value for each future time step. Simple to communicate, but gives no sense of uncertainty.
Interval forecast — A range (e.g. 95% prediction interval) within which the future value is expected to fall. Quantifies uncertainty, essential for decision-making.
A point forecast without an interval is an incomplete forecast.Prediction Intervals
For an ARIMA model, the forecast uncertainty grows with the horizon. At horizon h, the h-step-ahead prediction interval is:
For a random walk (ARIMA(0,1,0)), the forecast variance grows linearly with the horizon: σh2 = hσε2. This means prediction intervals fan out as a square root of h — a familiar feature when forecasting financial time series.
Forecast Accuracy Metrics
Accuracy is measured on a hold-out test set that the model never saw during training. The three most commonly reported metrics are MAE, RMSE, and MAPE.
No single metric is universally best. RMSE is preferred when large errors are especially costly; MAE when errors are roughly symmetric; MAPE when communicating to a non-technical audience. Always report multiple metrics and assess them alongside prediction interval coverage.
Exponential Smoothing Methods
Exponential smoothing methods form a family of forecasting algorithms that assign exponentially decreasing weights to past observations — recent observations matter more. They are often competitive with ARIMA and easier to tune.
Simple Exponential Smoothing (SES): Handles series with no trend or seasonality. One smoothing parameter α (0 < α < 1). Equivalent to ARIMA(0,1,1).
Holt’s Linear Trend Method: Adds a second equation tracking the trend with parameter β. Good for series with trend but no seasonality.
Holt–Winters (Triple Exponential Smoothing): Adds a seasonal component with parameter γ. The additive variant handles constant seasonal variation; the multiplicative variant handles growing seasonal swings.
Prophet: Forecasting at Scale
Prophet (released by Meta in 2017) is an open-source forecasting library designed for business time series with strong seasonality, multiple holiday effects, and missing data. It decomposes the series into trend, seasonality, and holidays/events, then adds a noise term.
Trend: Piecewise linear or logistic growth with automatic changepoint detection.
Seasonality: Fourier series with configurable period (yearly, weekly, daily).
Holidays: User-supplied event dates modeled with individual indicator terms.
Prophet is robust to outliers, handles gaps automatically, and requires minimal tuning — ideal for practitioners forecasting hundreds of series.
Cross-Validation for Time Series
Standard k-fold cross-validation shuffles observations randomly, which is invalid for time series because it leaks future information into training. The correct approach is walk-forward validation (also called time series cross-validation or expanding-window CV).
In walk-forward validation, you repeatedly train on data up to time t and evaluate on the next h observations, then advance the training window by one step and repeat. This mimics the real forecasting task and gives a distribution of forecast errors across many evaluation windows.
Expanding window: Training set grows at each step (Train on t=1..100, then 1..101, …). Uses all available history; best when data is scarce.
Rolling window: Training set has a fixed size that shifts forward (Train on t=1..100, then 2..101, …). Prevents distant past from dominating; better when the data-generating process changes over time.
Ensemble Forecasting
Combining forecasts from multiple models almost always outperforms any single model. This phenomenon — the forecast combination puzzle — is one of the most robust empirical findings in forecasting research. Even simple equal-weight averaging of ARIMA, exponential smoothing, and a naïve baseline often beats each individual method.
More sophisticated ensembles weight models by their recent accuracy (dynamic combination), use stacking via a meta-learner, or combine point forecasts with probabilistic models. The M4 Forecasting Competition (2018) demonstrated that hybrid statistical-ML ensembles set the state of the art on a benchmark of 100,000 series.
For K models with forecasts &hat;Y1,t+h, …, &hat;YK,t+h:
&hat;Yensemble = (1/K) ∑&hat;Yk,t+h- Point forecasts should always be accompanied by prediction intervals; uncertainty grows with the forecast horizon.
- MAE, RMSE, and MAPE each emphasize different aspects of forecast error; report multiple metrics evaluated on a held-out test set.
- Exponential smoothing (SES, Holt, Holt–Winters) provides simple and effective alternatives to ARIMA, especially for business series with trend and seasonality.
- Prophet handles business series with strong seasonality, holiday effects, and missing data automatically with minimal tuning.
- Walk-forward (expanding or rolling) cross-validation is the correct way to evaluate time series forecasting models; never use standard k-fold CV.
- Ensemble forecasts that combine multiple models nearly always outperform any single model; equal-weight averaging is a strong baseline.