Search arXiv⌕ Search

arXiv · 2610.02270

Reliability Stress Tests and Decision-Time Routing for Chest X-ray Vision-Language Models

Abstract

Medical vision-language model (VLM) evaluation is sensitive to workflow design, prompting strategy, and benchmark construction, yet most studies treat these factors in isolation. We introduce a reliability stress test for chest X-ray interpretation built on two balanced datasets (a private report-backed set and a curated MIMIC subset). Three medical VLMs (CheXagent, MedGemma-4B, and MedGemma-27B) are evaluated across three prompt styles and two workflows (single-VLM and multi-agent), producing 36 configurations. We show that exact-match accuracy alone can overstate the effectiveness of conservative models that default to "Normal" predictions. Diagnostic reliability also depends heavily on model family and scale: multi-agent reasoning helps some configurations but hurts others. Building on these observations, we propose a decision-time routing framework that selectively escalates to multi-agent inference only when beneficial, improving the cost-quality trade-off over fixed workflows. Our results highlight the need for evaluation protocols that jointly consider prompt sensitivity, failure-mode diversity, and workflow choice before clinical deployment.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xinye Yang, Zhusi Zhong, Scott Collins, Grayson Baird, Xuyu Wang, Zhicheng Jiao. 2026-10-01. Reliability Stress Tests and Decision-Time Routing for Chest X-ray Vision-Language Models. https://doi.org/10.1109/chase69719.2026.00073

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

FloVMos: Optical Flow-based Medical Video Mosaicking

Biomedical imaging modalities often require a trade-off among resolution, field of view (FOV), and acquisition speed. Video mosaicking offers a strategy to overcome this limitation by computationally stitching sequential high-resolution frames into a wide-FOV composite. However, existing methods struggle with non-rigid deformations, and modality-specific artifacts arising in clinical and research imaging. Here, we present FloVMos, a generalizable, optical-flow-based deep learning framework for real-time video mosaicking across diverse biomedical imaging modalities. FloVMos achieves robust, pixel-level registration by fine-tuning an optical flow model on synthetic training data with ground-truth deformation fields. We introduce a pipeline for generating this training data, simulating realistic tissue motion and imaging distortions from existing mosaics or raw videos. Our automated synthetic data generation and optical flow model training based on this data allow users to adapt FloVMos to different imaging modalities. To demonstrate this, we applied FloVMos to seven diverse imaging modalities: reflection confocal microscopy, open-top light-sheet microscopy, fetoscopy, laparoscopy, dermoscopy, sparse spectral microscopy, and endoscopy. FloVMos outperforms conventional baselines in accuracy, robustness, and speed for all the tested modalities. This adaptable and training-efficient framework enables large-area visualization with real-time performance and may support broader use of video-based biomedical imaging in research and clinical workflows.

eess.IV↗

Underwater imaging without color distortions requires RAW capture

1. Consumer cameras are not designed to be scientific instruments, yet they are increasingly used in science not just to see, but to measure. This is possible because camera sensors respond approximately linearly to the light they receive, a property that can, in principle, support quantitative measurements of color. 2. Whether that quantitative measurement potential is preserved depends on what happens to the image after capture. Minimally processed sensor images, commonly known as RAW, retain the approximately linear response needed. Instead of RAW, however, the default output of most cameras is a heavily processed image optimized for visual appearance, compact storage, and immediate use, typically saved as a JPEG. 3. The convenience offered by an in-camera processed image comes at a scientific cost. RAW capture requires more storage and processing effort but leaves calibration choices open. Processed images reduce those demands at acquisition and in post-processing but transfer irreversible decisions about color manipulation from the researcher to the camera. 4. The preservation of RAW sensor data is necessary but not sufficient for quantitative color measurement. Pixel values must also be calibrated to account for the light conditions under which the image was acquired. Underwater, this calibration is more challenging than in clear air, because light is absorbed and scattered as it travels through water, resulting in effects that depend on both wavelength and distance. Consequently, color distortions vary across the scene and generally cannot be corrected with a single global color transformation like white balancing; instead, calibration requires information about the imaging distance. We explain the consequences for quantitative imaging and introduce the 3P protocol for distortion-free color imaging underwater.

eess.IV↗

OASIS: Optimized Lightweight Autoencoder System for Distributed In-Sensor computing

In-sensor computing, which integrates computation directly within the sensor, has emerged as a promising paradigm for machine vision applications such as AR/VR and smart home systems. By processing data on-chip before transmission, it alleviates the bandwidth bottleneck caused by high-resolution, high-frame-rate image transmission. We propose a system architecture that integrates a CMOS image sensor (CIS) with a logic chip via advanced packaging, where the logic chip executes a lightweight encoder that aggressively compresses activations before off-chip transmission. However, conventional deep neural network (DNN) partitioning strategies require multiple layers before achieving meaningful dimensionality reduction, limiting bandwidth savings. To address this, we propose a dual-branch autoencoder-based vision architecture trained with a triple loss function combining task-specific, entropy, and reconstruction objectives to produce compact, Huffman-codable representations. The encoder is prototyped on a FPGA to provide hardware-validated energy measurements, complementing circuit-simulated CIS front-end characterization. Our approach achieves a four-order-of-magnitude reduction in transmitted data volume compared to raw input images, resulting in $2{\times}{-}4.5{\times}$ energy savings. We evaluate on CNN and ViT-based models across diverse tasks, achieving state-of-the-art accuracy with energy efficiency of up to 22.7 TOPS/W.

eess.IV↗