Search arXiv⌕ Search

arXiv · 2609.31789

MammoClaw: Towards Skill-Evolving Agent Harness for Breast Cancer Mammography Analysis

Abstract

In this work, we explore MammoClaw, a training-free agent framework that leverages frozen MLLMs for mammography analysis. To support agentic investigation, we equip the agent with lightweight mammography-specific tools for targeted image analysis, including ROI, paired-view, and contralateral-breast examination. MammoClaw iteratively gathers evidence through these tools, while skill evolution enables non-parametric adaptation by transforming failed trajectories into reusable guidance for later runs. We evaluate the framework on BI-RADS assessment and breast density estimation tasks. In our experiments, we find that tools alone do not reliably improve performance, whereas evolved skills can improve tool-use behavior and performance in some settings. Beyond these results, MammoClaw enables transparent inspection of evidence acquisition, tool interactions, and failure modes, facilitating the analysis and auditing of agent behavior. We view this work as an exploratory study of training-free, self-evolving agentic approaches for mammography and hope it provides a concrete starting point for future work on mammography-specific tools and self-evolution mechanisms. We release our code at https://krishnakanthnakka.github.io/mammoclaw.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Krishna Kanth Nakka. 2026-09-25. MammoClaw: Towards Skill-Evolving Agent Harness for Breast Cancer Mammography Analysis. https://arxiv.org/abs/2609.31789

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Active Data Acquisition with Side Information via Discrete Diffusion Priors

Acquiring data is costly: higher measurement fidelity costs power and storage and risks collecting irrelevant content, while aggressive cost reduction can discard information that later analysis needs. We address this trade-off with an information-theoretic framework that acquires data relevant to a broad set of tasks rather than to one model. A mask policy, conditioned on side information, chooses which pixels to measure so as to maximize the mutual information between a discrete image and its partial observation under a budget; since the image entropy does not depend on the mask, this is equivalent to minimizing the conditional entropy. A frozen discrete denoising diffusion model (D3PM) supplies the posterior, and we use it in two ways: as an entropy surrogate for training a one-shot mask generator, and as the criterion for sequential greedy acquisition. The one-shot generator outperforms random masks only with care, including an unbiased gradient estimator for binary masks. With sequential acquisition, on MNIST the prior makes $8\times$ fewer errors than random at a $10\%$ budget, and on CIFAR-10 it gains $0.9$--$3.4$~dB. On fastMRI, our proposed technique using a static mask outperforms the well-known methods such as variable density and LOUPE.

eess.IV↗

Q-Probe: Scaling Image Quality Assessment to High Resolution via Context-Aware Agentic Probing

Reinforcement Learning (RL) has empowered Multimodal Large Language Models (MLLMs) to achieve superior human preference alignment in Image Quality Assessment (IQA). However, existing RL-based IQA models typically rely on coarse-grained global views, failing to capture subtle local degradations in high-resolution scenarios. While emerging "Thinking with Images" paradigms enable multi-scale visual perception via zoom-in mechanisms, their direct adaptation to IQA induces spurious "cropping-implies-degradation" biases and misinterprets natural depth-of-field as artifacts. To address these challenges, we propose Q-Probe, the first agentic IQA framework designed to scale IQA to high resolution via context-aware probing. First, we construct Vista-Bench, a pioneering benchmark tailored for fine-grained local degradation analysis in high-resolution IQA settings. Furthermore, we propose a three-stage training paradigm that progressively aligns the model with human preferences, while simultaneously eliminating causal bias through a novel context-aware cropping strategy. Extensive experiments demonstrate that Q-Probe achieves state-of-the-art performance in high-resolution settings while maintaining superior efficacy across resolution scales.

eess.IV↗

Optimizing Multiple Feature Types for Image Inpainting in the Linear and Nonlinear Setting

Inpainting-based compression stores a carefully optimized subset of the full image data and reconstructs the missing data by inpainting. The quality of these lossy codecs depends decisively on the stored data. So far, these data consist almost exclusively of pixel locations along with their grayscale or color values. In the present paper, we present a general theory and a practical framework that allows to incorporate arbitrary features which can be described by linear or nonlinear equations. This includes e.g. derivatives of arbitrary order or local integrals. Our features can be combined with linear or nonlinear inpainting operators. Moreover, we present an algorithm that automatically optimizes the location and the type of the selected feature. The approach of allowing different types of optimized features turns inpainting-based compression into a more general, versatile and powerful paradigm. Our experiments report a consistent quality gain when increasing the number of feature types from 1 to 5. With the same amount of stored data, the average peak signal-to-noise improvement is 2.76 dB for harmonic (homogeneous diffusion) inpainting, and 1.82 dB for edge-enhancing diffusion inpainting.

eess.IV↗