| 151 |
Parse, Search, and Confirmation: Training-Free Aerial Vision-and-Dialog Navigation with Chain-of-Thought Reasoning and Structured Spatial Memory
Yu Qi, Hongyu Li, ... (+7 more)
|
|
cs.CV
|
0 |
17 days ago |
| 152 |
GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents
Suhaas Garre, Emily Ritchie, ... (+2 more)
|
|
cs.CV
|
0 |
17 days ago |
| 153 |
SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception
Mingjie Xie, Guangjun He, ... (+6 more)
|
|
cs.CV
|
0 |
17 days ago |
| 154 |
GIRAF: Towards Generalizable Human Interactions with Articulated Objects
Xiaohan Zhang, Sebastian Starke, ... (+4 more)
|
|
cs.CV
|
0 |
22 days ago |
| 155 |
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF
Eric Zhu, Abhinav Shrivastava, Soumik Mukhopadhyay
|
|
cs.LG
|
0 |
22 days ago |
| 156 |
Why Fake ? Unveiling the Semantic Vocabulary of Deepfake Detectors
Vazgken Vanian, Alexandros Doumanoglou, Dimitris Zarpalas
|
|
cs.CV
|
0 |
22 days ago |
| 157 |
Andha-Dhun: A First Look at Audio Descriptions in Hindi
Ritabrata Chakraborty, Divy Kala, ... (+4 more)
|
|
cs.CV
|
0 |
23 days ago |
| 158 |
Breaking Spurious Correlations via Generative Randomization and Cross-Variant Self-Supervised Learning
Suraj Yadav, Anjaneya Sharma, Siddharth Yadav
|
|
cs.CV
|
0 |
23 days ago |
| 159 |
CaT-GS: Efficient 3DGS Rendering for Large Scale Scenes via Inter-frame Caching and Tile Scheduling
Tingjia Zhang, Bo Chen, ... (+3 more)
|
|
cs.CV
|
0 |
10 days ago |
| 160 |
Test-Time Scaling for World Action Models via Zero-Shot Geometric Evaluation
Zesen Zhao, Minkyoung Cho, ... (+5 more)
|
|
cs.RO
|
0 |
10 days ago |