| 251 |
Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models
Yue Han, Chong Li, ... (+6 more)
|
|
cs.CV
|
0 |
2 months ago |
| 252 |
Towards Metric-Agnostic Trajectory Forecasting
Markus Knoche, Daan de Geus, Bastian Leibe
|
|
cs.CV
|
0 |
2 months ago |
| 253 |
LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models
Arpita Nema, Hanwei Zhu, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 254 |
AutoSpeed: Annotation-Free Stage-Adaptive Motion Speed Learning for Robot Manipulation
Qingda Hu, Ziheng Qiu, ... (+3 more)
|
|
cs.RO
|
0 |
2 months ago |
| 255 |
AVSR-Diff: Scale-Agnostic Diffusion Priors for Temporally Consistent Arbitrary-Scale Video Super-Resolution
Geunhyuk Youk, Jeonghyeok Do, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 256 |
Condensing Large-Scale Datasets Directly with Minimal Information Loss
Xinyi Shang, Peng Sun, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 257 |
MG-RWKV: Multi-Grained Context-Aware RWKV for Temporal Forgery Localization
Jingchen Ni, Cangjin Yu, ... (+7 more)
|
|
cs.CV
|
0 |
2 months ago |
| 258 |
DeWorldSG: Depth-Aware 3D Semantic Scene Graph Generation via World-Model Priors
Seok-Young Kim, Abdelrahman Elskhawy, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |
| 259 |
Improving Sparse-View 3DGS Generalization via Flat Minima Optimization
Kangmin Seo, Sangeek Hyun, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 260 |
MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment
Peiyuan Zhu, Shaoan Xie, ... (+6 more)
|
|
cs.CV
|
0 |
2 months ago |
| 261 |
GKDT: General Keypoint Detection Transformer
Changsheng Lu, Yuxin Chen, ... (+6 more)
|
|
cs.CV
|
0 |
2 months ago |
| 262 |
Towards Memory-Efficient Autoregressive Video Generation via Instance-Specific Parametric Absorption
Xiaomeng Fu, Jia Li, ... (+6 more)
|
|
cs.CV
|
0 |
2 months ago |
| 263 |
AdaBoosting Text Prompts for Vision-Language Models
Seokhee Jin, Changhwan Sung, ... (+3 more)
|
|
cs.LG
|
0 |
2 months ago |
| 264 |
Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts
Taewook Kang, Taeheon Kim, ... (+2 more)
|
|
cs.RO
|
0 |
2 months ago |
| 265 |
Not All Prediction Targets Keep Training-Free Diffusion Guidance on the Manifold
Yunsung Lee, Hyeongmin Lee
|
|
cs.CV
|
0 |
2 months ago |
| 266 |
Token-level Response-visual Attention Guidance for Multimodal LLMs Knowledge Distillation
Jaehyun Jang, Eunseop Yoon, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 267 |
EPO: Boosting 3D Foundation Models with Edge-based Pose Optimization
Mattia D'Urso, Christian Sormann, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 268 |
Caption Bottleneck Models
Seref Baris Cagliyan, Umut Ozdemir, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 269 |
BrainFIBRE: A Foundation Model via Information Decomposition for Brain Microstructure
Zijian Dong, Yi Lin, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 270 |
SPECSIA: Stylization Dataset for Novel-View Enhancement in Drawing-based 3D Animation
Kyuwon Kim, Sunjae Yoon, Chang D. Yoo
|
|
cs.CV
|
0 |
2 months ago |
| 271 |
HieDG: A Hierarchical Discrete Geometry-Guided Framework for Multi-Animal Tracking
Chenxun Deng, Zhongde Zhang, ... (+8 more)
|
|
cs.CV
|
0 |
2 months ago |
| 272 |
GenSP: Consistent Spherical Parameterization via Learning Shape Generative Models
Sai Karthikey Pentapati, Shashank Gupta, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 273 |
StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning
Yuan Qing, Chengzhi Mao, Boqing Gong
|
|
cs.CV
|
0 |
2 months ago |
| 274 |
Multi-scale Mixture of World Models for Embodied Agents in Evolving Environments
Jinwoo Jang, Daniel J. Rho, ... (+3 more)
|
|
cs.AI
|
0 |
2 months ago |
| 275 |
VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement
Seohyun Lee, Seoung Choi, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 276 |
Information-Regularized Attention for Visual-Centric Reasoning
Guohao Sun, Xiaofang Wang, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 277 |
HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding
Ji Ha Jang, Hayeon Kim, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 278 |
MedCAGD: Context-Aware Gated Decoder for Efficient Medical Image Segmentation
Saad Wazir, Patrick Dominique Vibild, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 279 |
Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising
Tianci Liu, Zihan Dong, ... (+6 more)
|
|
cs.AI
|
0 |
2 months ago |
| 280 |
The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models
Adeel Yousaf, Soumik Ghosh, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 281 |
Vitality-Aware Compression for Efficient Image-to-Shape Diffusion Transformers
Jaeah Lee, Hyunjin Kim, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 282 |
Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval
Jingjing Zhang, Lei Zhang, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 283 |
MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts
Nuoyan Zhou, Zhijun Tu, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |
| 284 |
SFDATrack: Generalized Source-Free Domain Adaptive Tracking Under Adverse Weather Conditions
Siyuan Yao, Ziqi Wang, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 285 |
ReShift: Aha-Moment-Driven Reasoning-Level Backdoor Attacks on Vision-Language Models
Zhihao Dou, Qinjian Zhao, ... (+2 more)
|
|
cs.CR
|
0 |
2 months ago |
| 286 |
Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs
Yoonhyung Park, Minji Kim, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 287 |
OnPoint: Offline-to-Online Multi-Level Distillation for Point-Supervised Online Temporal Action Localization
Sakib Reza, Gauri Jagatap, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 288 |
VOCA: Visual Odometry with Codec Awareness
Nouri Alexander Hilscher, Mateo de Mayo, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 289 |
DriftScope: Measuring The Hidden Effects of Diffusion Model Adaptation
Héctor Laria, Yiping Han, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |
| 290 |
Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting
Qian Ma, S M Rayeed, ... (+3 more)
|
|
cs.CL
|
0 |
2 months ago |
| 291 |
Progressive Pose-Guided 4D Animal Reconstruction from Monocular Video
Siyuan Li, Weiying Chen, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 292 |
A Mechanism-Driven Theory of Phase Transitions in Active Learning
Julia Machnio, Mads Nielsen, Mostafa Mehdipour Ghazi
|
|
cs.CV
|
0 |
2 months ago |
| 293 |
Benchmarking Federated Learning and Knowledge Distillation for Point Cloud Classification
Aizierjiang Aiersilan
|
|
cs.GR
|
0 |
2 months ago |
| 294 |
Segmenting, Fast and Slow: Real-Time Open-Vocabulary Video Instance Segmentation with Dual-Path Processing
Luca Barsellotti, Martin Sundermeyer, ... (+6 more)
|
|
cs.CV
|
0 |
2 months ago |
| 295 |
Lost in the Tail: Addressing Geographic Imbalance in Urban Visual Place Recognition
Zhiyao Shu, Jiacheng Yang, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 296 |
FaceMoE: Mixture of Experts for Low-Resolution Face Recognition
Kartik Narayan, Vishal M. Patel
|
|
cs.CV
|
0 |
2 months ago |
| 297 |
Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers
Anh Nguyen, Ngan Nguyen, ... (+12 more)
|
|
cs.CV
|
0 |
2 months ago |
| 298 |
LUNA: Learning Universal 3D Human Animation Beyond Skinning
Peng Li, Rawal Khirodkar, ... (+7 more)
|
|
cs.CV
|
0 |
2 months ago |
| 299 |
No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs
Haojian Huang, Harold Haodong Chen, ... (+7 more)
|
|
cs.CV
|
0 |
2 months ago |
| 300 |
DriveWeaver: Point-Conditioned Video Inpainting for Controllable Vehicle Insertion in Autonomous Driving Simulation
Junzhe Jiang, Zipei Ma, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |