Learning to Select Source-Traceable Evidence for Language-Model Prediction from Irregular Clinical Time Series
Numerical time-series models effectively process irregular electronic health record (EHR) trajectories, but do not expose which temporal patterns support each prediction as readable evidence. Existing text-based interfaces either serialize observations, preserving source traceability but offering limited clinical interpretation, or generate patient-level summaries that improve readability but can obscure links to source measurements. We introduce STEP-CTS (Source-Traceable Evidence for Prediction from Clinical Time-Series), which learns to select source-traceable text evidence for a language-model predictor. Multi-scale window statistics of each trajectory are verbalized as sets of deterministic threshold predicates, each set linked to its source window (e.g., "[5-7 h]: last temperature at least 38C"). An offline LLM, run once per unique predicate set without access to patient records, the prediction task, or outcome labels, attaches clinical concepts such as "fever" to their supporting predicates and abstains when none applies. Each predicate set, together with its concepts, forms an evidence unit. A learned Evidence Selector selects a fixed-size subset of the evidence units, which a pretrained clinical language-model encoder reads to make the prediction, with no patient-level text generation. Across three ICU benchmarks, STEP-CTS outperforms evaluated text-based baselines, improving AUPRC over the strongest by 5.1, 3.7, and 11.7 percentage points on P2012, MIMIC-III, and P2019, respectively, and is competitive with dedicated numerical time-series models. Ablations show that clinical concepts and learned evidence selection each contribute to predictive performance. In a blinded clinician study, the selected evidence is rated as traceable as deterministic serialization, near the ceiling of the scale, and above generated summaries on clinical interpretation.