arXiv · 2601.18493
DisasterInsight: A Building-Centric Benchmark for Evaluating Vision--Language Models in Disaster Response
Abstract
Vision--language models (VLMs) show promise for disaster-response remote sensing, but existing benchmarks mainly emphasize scene-level or damage-centric assessment. To study this building-centric gap, we introduce \method{}, a diagnostic benchmark built on xBD, a pre/post-disaster satellite dataset with building-level damage labels. \method{} enriches building instances with OpenStreetMap-derived functional labels and contains 134{,}108 task-specific instruction records across 15 task types, spanning instance-level assessment, scene-level counting, multi-instance reasoning, and structured report generation. The benchmark supports RGB pre/post-disaster imagery, single- and multi-view instance formulations, and scene-level RGB/SAR diagnostic inputs. Experiments with general-domain and remote-sensing VLMs show that models perform better on visible damage cues than on building-function understanding, multi-instance reasoning, counting, and grounded reporting. Instruction tuning improves performance on several tasks but does not close this building-centric gap.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sara Tehrani, Yonghao Xu, Leif Haglund, Amanda Berg, Gulnaz Zhambulova, Michael Felsberg. 2026-09-18. DisasterInsight: A Building-Centric Benchmark for Evaluating Vision--Language Models in Disaster Response. https://arxiv.org/abs/2601.18493
Cite the original work for its findings. Save a collection to share your selection of sources.