arXiv2026
Numeric values in clinical narratives, such as heart rate, oxygen saturation, and pressure gradients, carry diagnostic meaning that Transformer models trained on generic text do not capture. Objective: We categorize numerical values in French pediatric intensive care unit (PICU) notes into eight physiological categories using CamemBERT-bio, under two constraints that make large-scale LLMs impractical: only 1,072 real, annotated clinical samples are available for this rare, single-site condition, and training must run on GPUs shared concurrently with other hospital workloads rather than a dedicated cluster. Methods: We compare fine-tuning CamemBERT-bio with Label Embedding for Self-Attention (LESA) against combining LESA with Xval, a magnitude-aware number embedding, under a multi-objective training loss. Results: Standard fine-tuning did not improve F1 score, but CamemBERT-bio + LESA raised it by over 13%, and adding Xval matched this gain while approaching GPT-4's performance. Conclusion: LESA and Xval let a compact encoder achieve reliable physiological value extraction under limited real data and shared hospital compute, offering a practical alternative to large-scale LLMs. Significance: Under limited-data and shared-compute constraints, this compact BERT-based language model remains effective without the resource trade-offs of trillion-parameter LLMs.