Towards Fairness under Label Bias in Image Segmentation: Impact, Measurement and Mitigation
Biases in annotation pipelines may introduce label bias: when annotation quality or style varies systematically across demographic subgroups, causing disparities in performance and auditability. Label bias in image segmentation remains underexplored, as even detecting it typically requires clean, unbiased annotations, which are not readily available. We present a Confident Learning framework adapted to segmentation for auditing label bias directly in the training data without a clean, unbiased ground truth. By comparing the provided training labels to the model's confident predictions, we measure the prevalence and direction of suspected label errors; where standard overlap metrics like Dice fail. We further show that label bias influences subgroup separability in the encoder's feature space, an artifact we leverage for bias mitigation rather than suppressing it. We evaluate three datasets spanning from synthetic to real-life bias and across experimental conditions, showing how our framework produces proxy audit signals without clean labels and mitigates disparities with a correctly identified cleaner reference group.