Skip to content
← All publications
Preprint 2026

SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy

Ismael Elsharkawi, Ahmed Sait, Silvio Giancola, Bernard Ghanem, Hossam Sharara, Abdelrahman Eldesokey

SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy

Abstract

Introduces a benchmark for whether soccer-video VLMs are actually grounded in what's on screen, not just pattern-matching to the right event label. Annotates soccer clips across 13 event types with layered visual cues and extends spatio-temporal attribution methods to score attention alignment. Finds that even accurate models rarely exceed 50% grounding and consistently underuse temporal information — accuracy and true visual grounding diverge.

Multi-Modal LLMsVideo Understanding