What If TSF: Reframing Time Series Forecasting as Scenario-Guided Multimodal Forecasting
Recent advances in large language models (LLMs) have enabled time series forecasting to move beyond numerical observations and incorporate external information in multimodal settings. Such information can improve forecasting performance, but accuracy alone may not reveal whether models appropriately respond to it: a model may ignore relevant information, fail to distinguish scenarios with different implications, or overreact to irrelevant or weak signals. We introduce What If TSF (WIT), a benchmark for evaluating whether models effectively respond to future scenarios. WIT constructs controlled scenario sets by fixing the forecasting state while systematically varying future scenarios, enabling three complementary evaluations: Factual Comparison, which measures the predictive utility of factual scenario information; Relational Comparison, which evaluates whether forecasts satisfy expected direction, contrast, intensity, and restraint relations across scenarios; and Grounded Comparison, which assesses whether scenario-induced revision trajectories are empirically plausible relative to comparable real cases. Experiments show that factual accuracy gains do not consistently translate into appropriate responses to alternative scenarios, demonstrating the need to evaluate multimodal forecasting beyond predictive accuracy.