RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning
Dense image captioning is critical for vision-language pretraining and for text-to-image generation, but scaling expert-quality annotations is prohibitively expensive. While synthesizing captions from strong vision-language models (VLMs) is a practical alternative, supervised distillation often yields limited output diversity and weak generalization. Reinforcement learning (RL) could overcome these, but its successes have so far been concentrated in verifiable domains that rely on deterministic checkers---a luxury not available in open-ended captioning. We address this verification bottleneck with RubiCap, an RL framework that derives fine-grained, sample-specific rewards from LLM-written rubrics. For each image, an LLM rubric writer compares captions from a diverse committee of VLMs to identify consensus strengths and diagnose the current policy's deficiencies. These findings are converted into explicit evaluation criteria, enabling an LLM judge to decompose quality assessment and replace coarse scalar rewards with structured, multi-faceted assessments. Across extensive benchmarks, RubiCap achieves the highest win rates on CapArena, outperforming supervised distillation, prior RL methods, human-expert annotations, and GPT-4V-augmented outputs. On CaptionQA, it shows superior word efficiency: our 7B model matches Qwen2.5-VL-32B-Instruct, and our 3B model surpasses its 7B counterpart. Remarkably, using a compact RubiCap-3B as a captioner produces stronger pretrained VLMs than those trained on captions from proprietary models.