R.I.P.
๐ป
Ghosted
Separating Clicks from Baits: Using Large Language Models to Detect Misleading YouTube Thumbnails
July 26, 2026 ยท Grace Period ยท ๐ Proceedings of the International AAAI Conference on Web and Social Media, 2027
Authors
Wajiha Naveed, Muhammad Muneeb Pervez, Zaeem Mohtashim Khan, Zafar Ayyub Qazi, Zartash Afzal Uzmi
arXiv ID
2607.23739
Category
cs.SI: Social & Info Networks
Citations
0
Venue
Proceedings of the International AAAI Conference on Web and Social Media, 2027
Abstract
Misleading video thumbnails on platforms like YouTube are a pervasive problem, undermining user trust and platform integrity. This paper proposes a novel multi-modal detection pipeline that uses Large Language Models (LLMs) to flag misleading thumbnails. We first construct a comprehensive dataset of 2,843 videos from eight countries, including 1,359 misleading thumbnail videos that collectively amassed over 7.6 billion views, providing a unique cross-cultural perspective on this global issue. Our detection pipeline integrates video-to-text descriptions, thumbnail images, and subtitle transcripts to holistically analyze content and flag misleading thumbnails. Through extensive experimentation and prompt engineering, we evaluate the performance of four frontier-level LLMs, including GPT-4o, GPT-4o Mini, Claude 3.5 Sonnet, and Gemini-1.5 Flash. We further evaluate open-weight vision-language models, LLaVA-v1.5 and Qwen2.5-VL-7B-Instruct, to assess the generalizability of our approach beyond proprietary systems. Our findings show the effectiveness of LLMs in identifying misleading thumbnails, with Claude 3.5 Sonnet consistently showing strong performance, achieving an accuracy of 93.8%, precision over 92%, and recall exceeding 94% in certain scenarios. Beyond evaluating detection performance, we conducted a careful failure analysis to understand when LLMs fail in identifying misleading thumbnails. We discuss the implications of our findings for content moderation, user experience, and the ethical considerations of deploying such systems at scale. Our findings pave the way for more transparent, trustworthy video platforms and stronger content integrity for audiences worldwide.
Community Contributions
Found the code? Know the venue? Think something is wrong? Let us know!
๐ Similar Papers
In the same crypt โ Social & Info Networks
R.I.P.
๐ป
Ghosted
Fake News Detection on Social Media: A Data Mining Perspective
R.I.P.
๐ป
Ghosted
Natural Scales in Geographical Patterns
R.I.P.
๐ป
Ghosted
Representation Learning on Graphs: Methods and Applications
R.I.P.
๐ป
Ghosted
The COVID-19 Social Media Infodemic
R.I.P.
๐ป
Ghosted