Skip to content
← All publications
Preprint 2026

CounterCount: A Diagnostic Framework for Counting Bias in Vision Language Models

Reem Alzahrani, Hassan Alshanqiti, Bushra Bin Hemid, Zaid Alyafeai, Abdelrahman Eldesokey, Bernard Ghanem

CounterCount: A Diagnostic Framework for Counting Bias in Vision Language Models

Abstract

Tests whether vision-language models actually look at the image when counting objects, using paired factual/counterfactual images with edited counts and localized evidence annotations. Models do well on factual images but degrade under counterfactual edits, revealing reliance on learned object-count priors — traced to underweighted attention on count-relevant visual tokens, which an inference-time attention-reweighting strategy partially corrects.

Multi-Modal LLMsEvaluation