FRAME: Separating sampling variation from performance disparities in medical image analysis
Fairness audits of medical imaging models commonly report the largest performance difference between demographic subgroups and treat any positive value as evidence of bias. Yet a perfectly fair model also produces a positive difference, which grows as subgroups shrink. Some mitigation methods remove demographic information from models because they assume that it produces the difference. Here we introduce Fair-model Reference And Mechanism Evaluation (FRAME). Its first step compares each reported difference with the difference expected from a perfectly fair model at the same subgroup counts. Its second step tests whether candidate causes of any excess over this fair-model reference change the difference when they are injected into model features. We evaluated 36 encoders on nine datasets in three modalities and audited 89 published differences across six modalities. For 10 frozen encoders and 13 chest radiograph findings, the reference is a median 41% of the race difference and 22% of the age difference. It is exceeded by 40 of 53 published differences in sensitivity or false positive rate and by one of 36 in the area under the receiver operating characteristic curve (AUROC). Adding race information to the features of two encoders does not change the race difference significantly. At matched disease performance, mitigation reduces the race difference by less than its spread across three pretraining seeds (medians 0.005 and 0.012) and the age difference by more (0.008 and 0.006). Image-text pretraining raises the worst-group AUROC for race by a median 0.049 over self-supervised pretraining. Reporting the fair-model reference beside every new and published subgroup difference could direct mitigation to the differences above it.