Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models

December 01, 2022 · Declared Dead · 🏛 Conference of the European Chapter of the Association for Computational Linguistics

Repo contents: .gitmodules, CLIP, OFA, README.md, detectron2, mae, moco-v3, segmenter, vlp_probe

Authors Zhuowan Li, Cihang Xie, Benjamin Van Durme, Alan Yuille arXiv ID 2212.00281 Category cs.CV: Computer Vision Cross-listed cs.CL Citations 2 Venue Conference of the European Chapter of the Association for Computational Linguistics Repository https://github.com/Lizw14/visual_probing ⭐ 1 Last Checked 1 month ago

Abstract

Despite the impressive advancements achieved through vision-and-language pretraining, it remains unclear whether this joint learning paradigm can help understand each individual modality. In this work, we conduct a comparative analysis of the visual representations in existing vision-and-language models and vision-only models by probing a broad range of tasks, aiming to assess the quality of the learned representations in a nuanced manner. Interestingly, our empirical observations suggest that vision-and-language models are better at label prediction tasks like object and attribute prediction, while vision-only models are stronger at dense prediction tasks that require more localized information. We hope our study sheds light on the role of language in visual learning, and serves as an empirical guide for various pretrained models. Code will be released at https://github.com/Lizw14/visual_probing

📄 View on arXiv 🌐 View on ar5iv 📑 PDF 💻 Repository 🎉 Report Code Found

Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

📜 Similar Papers

In the same crypt — Computer Vision

🌅 🌅 Old Age

Deep Residual Learning for Image Recognition

Kaiming He, Xiangyu Zhang, ... (+2 more)

cs.CV 🏛 CVPR 📚 220.4K cites 10 years ago

🌅 🌅 Old Age

Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

Shaoqing Ren, Kaiming He, ... (+2 more)

cs.CV 🏛 IEEE TPAMI 📚 70.4K cites 10 years ago

R.I.P. 👻 Ghosted

You Only Look Once: Unified, Real-Time Object Detection

Joseph Redmon, Santosh Divvala, ... (+2 more)

cs.CV 🏛 CVPR 📚 43.4K cites 10 years ago

🌅 🌅 Old Age

SSD: Single Shot MultiBox Detector

Wei Liu, Dragomir Anguelov, ... (+5 more)

cs.CV 🏛 ECCV 📚 33.8K cites 10 years ago

🌅 🌅 Old Age

Squeeze-and-Excitation Networks

Jie Hu, Li Shen, ... (+3 more)

cs.CV 🏛 CVPR 📚 32.3K cites 8 years ago

R.I.P. 👻 Ghosted

Rethinking the Inception Architecture for Computer Vision

Christian Szegedy, Vincent Vanhoucke, ... (+3 more)

cs.CV 🏛 CVPR 📚 30.2K cites 10 years ago

Died the same way — 🦴 Skeleton Repo

R.I.P. 🦴 Skeleton Repo

EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification

Patrick Helber, Benjamin Bischke, ... (+2 more)

cs.CV 🏛 J.STAEORS 📚 2.4K cites 8 years ago

R.I.P. 🦴 Skeleton Repo

Deep Learning for 3D Point Clouds: A Survey

Yulan Guo, Hanyun Wang, ... (+4 more)

cs.CV 🏛 IEEE TPAMI 📚 2.1K cites 6 years ago

R.I.P. 🦴 Skeleton Repo

Adversarial Examples: Attacks and Defenses for Deep Learning

Xiaoyong Yuan, Pan He, ... (+2 more)

cs.LG 🏛 IEEE TNNLS 📚 1.8K cites 8 years ago

R.I.P. 🦴 Skeleton Repo

Neural Style Transfer: A Review

Yongcheng Jing, Yezhou Yang, ... (+4 more)

cs.CV 🏛 IEEE TVCG 📚 828 cites 8 years ago