| 101 |
TextDS: Parameter-Efficient Representation Alignment for Scene Text Detection under Distribution Shifts
Boyuan Chen, Zichen Dang, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 102 |
MLVC: Multi-platform Learned Video Codec for Real-World Deployment
Tanel Pärnamaa, Martin Lumiste, ... (+4 more)
|
|
eess.IV
|
0 |
2 months ago |
| 103 |
EMOSH: Expressive Motion and Shape Disentanglement for Human Animation
Dongbin Zhang, Hao Liu, ... (+6 more)
|
|
cs.CV
|
0 |
2 months ago |
| 104 |
Latent Visual Diffusion Reasoning with Monte Carlo Tree Search
Xirui Teng, Nan Xi, Junsong Yuan
|
|
cs.CV
|
0 |
2 months ago |
| 105 |
There and Back Again: A Flexible-Frame Transformer for Multi-Exposure Fusion
Lishen Qu, Yao Liu, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |
| 106 |
Improving Adversarial Robustness via Activation Amplification and Attenuation
Taïga Gonçalves, Yongsong Huang, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 107 |
MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations
Hejia Chen, Haoxian Zhang, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |
| 108 |
SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models
Ruoyu Wang, Jialun Liu, ... (+6 more)
|
|
cs.CV
|
0 |
2 months ago |
| 109 |
MASS: Motion-Aligned Selective Scan for Refinement in Flow-Based Video Frame Interpolation
Jun-Sang Yoo, Seung-Won Jung
|
|
cs.CV
|
0 |
2 months ago |
| 110 |
Understanding Cross-Rig Generalization in Automotive Perception: a Multi-Rig Benchmark and Rig Variation Metrics
Tim Alexander Bader, Tim Dieter Eberhardt, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 111 |
Tessellating The Earth
Daniel Cher, Hamza Iqbal, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 112 |
SemCityLoc: Aerial 6DoF Localization Using Semantic 3D City Models
Jingfeng Mao, Xuyang Chen, ... (+7 more)
|
|
cs.CV
|
0 |
2 months ago |
| 113 |
Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models
Shravan Venkatraman, Ritesh Thawkar, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |
| 114 |
DnA: Denoising Attention for Visual Tasks
Ron Campos, Subhajit Maity, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 115 |
Don't Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance
Pradhaan S Bhat, Rishubh Parihar, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 116 |
SAM2Matting: Generalized Image and Video Matting
Ruiqi Shen, Guangquan Jie, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 117 |
ViQ: Text-Aligned Visual Quantized Representations at Any Resolution
Xumin Yu, Zuyan Liu, ... (+6 more)
|
|
cs.CV
|
0 |
2 months ago |
| 118 |
See & Sniff: Learning Visuo-Olfactory Representations
Seongyu Kim, Seungwoo Lee, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 119 |
E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation
Wen Ye, Peiyan Li, ... (+8 more)
|
|
cs.RO
|
0 |
2 months ago |
| 120 |
Geometric Gradient Rectification for Safe Open-Set Semi-Supervised Learning
Jiahe Chen, Qian Shao, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |
| 121 |
Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE
Haoyou Deng, Keyu Yan, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |
| 122 |
PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation
Kexu Cheng, Zicheng Liu, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 123 |
Multi-modality Image Fusion under Adverse Weather: Mask-Guided Feature Restoration and Interaction
Xilai Li, Xiaosong Li, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 124 |
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP
Sicheng Zhang, Muzammal Naseer, ... (+6 more)
|
|
cs.CV
|
0 |
2 months ago |
| 125 |
ResilPhase: Plug-and-Play Phase Mapping and Noise-Resilient Macro-Trajectory Extrapolation for Diffusion Acceleration
Qicheng Zhao, Yu Li, ... (+2 more)
|
|
cs.AI
|
0 |
2 months ago |
| 126 |
Capacity-Controlled Multi-View Stylization of 3D Gaussian Splatting
Zhihao Wen, Yixin Yang, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |
| 127 |
LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing
Xinyu Wang, Chongbo Zhao, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 128 |
Dual-Prior Guided Null-Space Learning with Mixture-of-Splines for Arbitrary Medical Slice Super-Resolution
Haofei Song, Siyuan Xu, ... (+4 more)
|
|
eess.IV
|
0 |
2 months ago |
| 129 |
TaskTok: Delving into Task Tokens for Task-driven Image Restoration
Hongjae Lee, Sojung Kang, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 130 |
LogicIR: Logic Gate Networks for Image Restoration
Hongjae Lee, Myungjun Son, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 131 |
DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues
Geng Li, Yuxin Peng
|
|
cs.CV
|
0 |
2 months ago |
| 132 |
PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing
Shengbin Guo, Shaokang He, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |
| 133 |
From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP
Zhixing Li, Yinan Yu
|
|
cs.CV
|
0 |
2 months ago |
| 134 |
Forget, Anticipate and Adapt: Test Time Training for Long Videos
Rajat Modi, Sebastian Noel, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 135 |
Rethinking Training & Inference for Forecasting: Linking Winner-Take-All back to GMMs
Qiyuan Wu, Katie Z Luo, ... (+3 more)
|
|
cs.LG
|
0 |
2 months ago |
| 136 |
Staying VIGILant: Mitigating Visual Laziness via Counterfactual Visual Alignment in MLLMs
Xi Xiao, Chen Liu, ... (+10 more)
|
|
cs.CV
|
0 |
2 months ago |
| 137 |
Layer-Specific Prompt Fusion Discovery via Differentiable Search in Vision Foundation Models
Xi Xiao, Xingjian Li, ... (+8 more)
|
|
cs.CV
|
0 |
2 months ago |
| 138 |
MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation
Yang Chen, Xiaowei Xu, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |
| 139 |
In-context Region-based Drag: Drag Any Region to Any Shape
Jiacheng Sui, Tianyu Hao, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 140 |
$S^{2}$-FracMix: Label-Preserving Self-Saliency Mixup Augmentation
Khawar Islam, Arif Mahmood, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 141 |
UniTeD: Unified Temporal Diffusion for Joint Perception and Planning in Autonomous Driving
Bo Zhao, Xinting Zhao, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 142 |
Falcon: Functional Assembly and Language for Compositional Reasoning in X-ray
Yonathan Michael, Mohamad Alansari, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 143 |
FeVOS: Foresight Expression Video Object Segmentation
Kehan Lan, Kaining Ying, Henghui Ding
|
|
cs.CV
|
0 |
2 months ago |
| 144 |
H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks
Seulgi Jeong, Yunseong Cho, Sanghun Park
|
|
cs.CV
|
0 |
2 months ago |
| 145 |
Distill on a Diet: Efficient Knowledge Distillation via Learnable Data Pruning
Yifan Wu, Yiqi Wang, ... (+6 more)
|
|
cs.LG
|
0 |
2 months ago |
| 146 |
C3-Bench: A Context-Aware Change Captioning Benchmark
Jae-Woo Kim, Hyeongbeom Kim, Ue-Hwan Kim
|
|
cs.CV
|
0 |
2 months ago |
| 147 |
Geometry-Anchored Transport Framework for Exemplar-Free Class-Incremental Learning
Hongye Xu, Bartosz Krawczyk
|
|
cs.LG
|
0 |
2 months ago |
| 148 |
Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity
Heethanjan Kanagalingam, Thenukan Pathmanathan, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 149 |
Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation
Atin Pothiraj, Jaemin Cho, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 150 |
OrthoTrack: Continuous 6-DoF UAV Trajectory Estimation Anchored in Public Orthophotos
Oussema Dhaouadi, Zuria Bauer, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |