CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation

July 24, 2025 · Declared Dead · 🏛 ACM Multimedia

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Hyunwoo Oh, SeungJu Cha, Kwanyoung Lee, Si-Woo Kim, Dong-Jin Kim arXiv ID 2507.18750 Category cs.MM: Multimedia Cross-listed cs.SD, eess.AS Citations 2 Venue ACM Multimedia Last Checked 3 months ago

Abstract

We propose CatchPhrase, a novel audio-to-image generation framework designed to mitigate semantic misalignment between audio inputs and generated images. While recent advances in multi-modal encoders have enabled progress in cross-modal generation, ambiguity stemming from homographs and auditory illusions continues to hinder accurate alignment. To address this issue, CatchPhrase generates enriched cross-modal semantic prompts (EXPrompt Mining) from weak class labels by leveraging large language models (LLMs) and audio captioning models (ACMs). To address both class-level and instance-level misalignment, we apply multi-modal filtering and retrieval to select the most semantically aligned prompt for each audio sample (EXPrompt Selector). A lightweight mapping network is then trained to adapt pre-trained text-to-image generation models to audio input. Extensive experiments on multiple audio classification datasets demonstrate that CatchPhrase improves audio-to-image alignment and consistently enhances generation quality by mitigating semantic misalignment.

📄 View on arXiv 🌐 View on ar5iv 📑 PDF 🎉 Report Code Found

Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

📜 Similar Papers

In the same crypt — Multimedia

🌅 🌅 Old Age

Quality Assessment of In-the-Wild Videos

Dingquan Li, Tingting Jiang, Ming Jiang

cs.MM 🏛 ACM MM 📚 375 cites 6 years ago

R.I.P. 👻 Ghosted

Viewport-Adaptive Navigable 360-Degree Video Delivery

Xavier Corbillon, Gwendal Simon, ... (+2 more)

cs.MM 🏛 ICC 📚 328 cites 9 years ago

📚 📚 The Cartographer

A Comprehensive Survey on Cross-modal Retrieval

Kaiye Wang, Qiyue Yin, ... (+3 more)

cs.MM 🏛 arXiv 📚 322 cites 9 years ago

📚 📚 The Cartographer

An Overview of Cross-media Retrieval: Concepts, Methodologies, Benchmarks and Challenges

Yuxin Peng, Xin Huang, Yunzhen Zhao

cs.MM 🏛 IEEE TCSVT 📚 309 cites 9 years ago

R.I.P. 👻 Ghosted

A Convolutional Neural Network Approach for Post-Processing in HEVC Intra Coding

Yuanying Dai, Dong Liu, Feng Wu

cs.MM 🏛 ICMM 📚 305 cites 9 years ago

R.I.P. 👻 Ghosted

Video Generation From Text

Yitong Li, Martin Renqiang Min, ... (+3 more)

cs.MM 🏛 AAAI 📚 300 cites 8 years ago

Died the same way — 👻 Ghosted

R.I.P. 👻 Ghosted

Federated Learning: Strategies for Improving Communication Efficiency

Jakub Konečný, H. Brendan McMahan, ... (+4 more)

cs.LG 🏛 arXiv 📚 5.2K cites 9 years ago

R.I.P. 👻 Ghosted

In-Datacenter Performance Analysis of a Tensor Processing Unit

Norman P. Jouppi, Cliff Young, ... (+73 more)

cs.AR 🏛 ISCA 📚 5.1K cites 9 years ago

R.I.P. 👻 Ghosted

Deep Convolutional Neural Networks for Computer-Aided Detection: CNN Architectures, Dataset Characteristics and Transfer Learning

Hoo-Chang Shin, Holger R. Roth, ... (+7 more)

cs.CV 🏛 IEEE TMI 📚 4.9K cites 10 years ago

R.I.P. 👻 Ghosted

Explanation in Artificial Intelligence: Insights from the Social Sciences

Tim Miller

cs.AI 🏛 AI 📚 4.9K cites 8 years ago