TY - RPRT TI - Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects AU - Emad Alharbi PY - 2026 UR - https://arxiv.org/abs/2608.28626 ID - 2608.28626 ER -