arXiv · 2609.26056
CricRAG: Retrieval Augmented Vision-Language Models for Personalized Cricket Coaching
Abstract
Vision-Language Models (VLMs) offer promising capabilities for automated sports coaching but face a fundamental limitation: they implicitly compare against professional standards, making their feedback impractical for developing players. We present CricRAG, a retrieval-augmented framework that aligns VLMs with skill-appropriate benchmarks for personalized cricket coaching. Our key insight is that by retrieving similar-but-better techniques as reference points, we can guide VLMs to provide developmentally appropriate feedback that mirrors human coaching practices. We contribute: (1) a labelled dataset of 288 cricket technique videos spanning multiple skill levels, (2) an efficient motion retrieval pipeline using contrastive learning that achieves 78% top-3 retrieval accuracy, (3) a frame sampling technique that reduces inference costs, and (4) a retrieval-augmented approach that significantly improves feedback alignment with coaching principles, achieving up to 94% agreement with professional assessments compared to 67% without retrieval context.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Agamdeep Singh, Sujit PB, Mayank Vatsa. 2026-08-08. CricRAG: Retrieval Augmented Vision-Language Models for Personalized Cricket Coaching. https://arxiv.org/abs/2609.26056
Cite the original work for its findings. Save a collection to share your selection of sources.