arXiv · 2609.36689
CHAIN: Calibrated LLM Forecasting via Causal-Temporal Hypergraph Inference
Abstract
Large language models have achieved significant progress in event forecasting, yet their probability outputs exhibit systematic calibration bias that varies heterogeneously across different domains and question types, undermining the trustworthiness of probabilistic outputs for decision-making under uncertainty. However, existing calibration methods typically correct probability outputs after prediction is complete, without modeling the structural sources of bias within the prediction process itself. To address this challenge, we decompose probabilistic prediction over causal-temporal hypergraphs into three stages, evidence weighting, evidence aggregation, and source fusion, and propose CHAIN, which designs stage-specific mechanisms to mitigate bias at each stage: (i) modulating the temporal decay function by causal topological distance, (ii) aggregating approximately independent causal chains via Noisy-OR after direction-aware deduplication, and (iii) driving adaptive fusion by causal coverage and directional balance. Experimental results on cross-domain forecasting benchmarks show CHAIN outperforms existing methods in expected calibration error, Brier score, and accuracy. Our project is available at https://github.com/QwenQKing/Chain.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wenjin Liu, Chenxi Wang, Yue Lu, Zhe Cui, Haoran Luo. 2026-09-29. CHAIN: Calibrated LLM Forecasting via Causal-Temporal Hypergraph Inference. https://arxiv.org/abs/2609.36689
Cite the original work for its findings. Save a collection to share your selection of sources.