| 101 |
HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration
Jiaxin Li, Yuxiang Wu, ... (+12 more)
|
|
cs.CV
|
0 |
1 month ago |
| 102 |
Toward Robust In-Context Segmentation via Concept Guidance
Zhigang Chen, Xiawu Zheng, Rongrong Ji
|
|
cs.CV
|
0 |
1 month ago |
| 103 |
Monocular Avatar Reconstruction via Cascaded Diffusion Priors and UV-Space Differentiable Shading
Hong Li, Minqi Meng, ... (+11 more)
|
|
cs.CV
|
0 |
1 month ago |
| 104 |
MixTTA: Low-Rank Cross-Channel Mixing for Reliable Test-Time Adaptation
Mansoo Jung, Youngwook Kim, Jungwoo Lee
|
|
cs.LG
|
0 |
1 month ago |
| 105 |
TextDS: Parameter-Efficient Representation Alignment for Scene Text Detection under Distribution Shifts
Boyuan Chen, Zichen Dang, ... (+3 more)
|
|
cs.CV
|
0 |
1 month ago |
| 106 |
MLVC: Multi-platform Learned Video Codec for Real-World Deployment
Tanel Pärnamaa, Martin Lumiste, ... (+4 more)
|
|
eess.IV
|
0 |
1 month ago |
| 107 |
EMOSH: Expressive Motion and Shape Disentanglement for Human Animation
Dongbin Zhang, Hao Liu, ... (+6 more)
|
|
cs.CV
|
0 |
1 month ago |
| 108 |
Latent Visual Diffusion Reasoning with Monte Carlo Tree Search
Xirui Teng, Nan Xi, Junsong Yuan
|
|
cs.CV
|
0 |
1 month ago |
| 109 |
There and Back Again: A Flexible-Frame Transformer for Multi-Exposure Fusion
Lishen Qu, Yao Liu, ... (+5 more)
|
|
cs.CV
|
0 |
1 month ago |
| 110 |
Improving Adversarial Robustness via Activation Amplification and Attenuation
Taïga Gonçalves, Yongsong Huang, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 111 |
MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations
Hejia Chen, Haoxian Zhang, ... (+5 more)
|
|
cs.CV
|
0 |
1 month ago |
| 112 |
SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models
Ruoyu Wang, Jialun Liu, ... (+6 more)
|
|
cs.CV
|
0 |
1 month ago |
| 113 |
MASS: Motion-Aligned Selective Scan for Refinement in Flow-Based Video Frame Interpolation
Jun-Sang Yoo, Seung-Won Jung
|
|
cs.CV
|
0 |
1 month ago |
| 114 |
Understanding Cross-Rig Generalization in Automotive Perception: a Multi-Rig Benchmark and Rig Variation Metrics
Tim Alexander Bader, Tim Dieter Eberhardt, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 115 |
Tessellating The Earth
Daniel Cher, Hamza Iqbal, ... (+3 more)
|
|
cs.CV
|
0 |
1 month ago |
| 116 |
SemCityLoc: Aerial 6DoF Localization Using Semantic 3D City Models
Jingfeng Mao, Xuyang Chen, ... (+7 more)
|
|
cs.CV
|
0 |
1 month ago |
| 117 |
Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models
Shravan Venkatraman, Ritesh Thawkar, ... (+5 more)
|
|
cs.CV
|
0 |
1 month ago |
| 118 |
DnA: Denoising Attention for Visual Tasks
Ron Campos, Subhajit Maity, ... (+3 more)
|
|
cs.CV
|
0 |
1 month ago |
| 119 |
Don't Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance
Pradhaan S Bhat, Rishubh Parihar, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 120 |
SAM2Matting: Generalized Image and Video Matting
Ruiqi Shen, Guangquan Jie, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 121 |
ViQ: Text-Aligned Visual Quantized Representations at Any Resolution
Xumin Yu, Zuyan Liu, ... (+6 more)
|
|
cs.CV
|
0 |
1 month ago |
| 122 |
See & Sniff: Learning Visuo-Olfactory Representations
Seongyu Kim, Seungwoo Lee, ... (+3 more)
|
|
cs.CV
|
0 |
1 month ago |
| 123 |
E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation
Wen Ye, Peiyan Li, ... (+8 more)
|
|
cs.RO
|
0 |
1 month ago |
| 124 |
Geometric Gradient Rectification for Safe Open-Set Semi-Supervised Learning
Jiahe Chen, Qian Shao, ... (+5 more)
|
|
cs.CV
|
0 |
1 month ago |
| 125 |
Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE
Haoyou Deng, Keyu Yan, ... (+5 more)
|
|
cs.CV
|
0 |
1 month ago |
| 126 |
PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation
Kexu Cheng, Zicheng Liu, ... (+3 more)
|
|
cs.CV
|
0 |
1 month ago |
| 127 |
Multi-modality Image Fusion under Adverse Weather: Mask-Guided Feature Restoration and Interaction
Xilai Li, Xiaosong Li, ... (+4 more)
|
|
cs.CV
|
0 |
1 month ago |
| 128 |
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP
Sicheng Zhang, Muzammal Naseer, ... (+6 more)
|
|
cs.CV
|
0 |
1 month ago |
| 129 |
ResilPhase: Plug-and-Play Phase Mapping and Noise-Resilient Macro-Trajectory Extrapolation for Diffusion Acceleration
Qicheng Zhao, Yu Li, ... (+2 more)
|
|
cs.AI
|
0 |
1 month ago |
| 130 |
Capacity-Controlled Multi-View Stylization of 3D Gaussian Splatting
Zhihao Wen, Yixin Yang, ... (+5 more)
|
|
cs.CV
|
0 |
1 month ago |
| 131 |
LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing
Xinyu Wang, Chongbo Zhao, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 132 |
Dual-Prior Guided Null-Space Learning with Mixture-of-Splines for Arbitrary Medical Slice Super-Resolution
Haofei Song, Siyuan Xu, ... (+4 more)
|
|
eess.IV
|
0 |
1 month ago |
| 133 |
TaskTok: Delving into Task Tokens for Task-driven Image Restoration
Hongjae Lee, Sojung Kang, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 134 |
LogicIR: Logic Gate Networks for Image Restoration
Hongjae Lee, Myungjun Son, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 135 |
DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues
Geng Li, Yuxin Peng
|
|
cs.CV
|
0 |
1 month ago |
| 136 |
PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing
Shengbin Guo, Shaokang He, ... (+5 more)
|
|
cs.CV
|
0 |
1 month ago |
| 137 |
From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP
Zhixing Li, Yinan Yu
|
|
cs.CV
|
0 |
1 month ago |
| 138 |
Forget, Anticipate and Adapt: Test Time Training for Long Videos
Rajat Modi, Sebastian Noel, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 139 |
Rethinking Training & Inference for Forecasting: Linking Winner-Take-All back to GMMs
Qiyuan Wu, Katie Z Luo, ... (+3 more)
|
|
cs.LG
|
0 |
1 month ago |
| 140 |
Staying VIGILant: Mitigating Visual Laziness via Counterfactual Visual Alignment in MLLMs
Xi Xiao, Chen Liu, ... (+10 more)
|
|
cs.CV
|
0 |
1 month ago |
| 141 |
Layer-Specific Prompt Fusion Discovery via Differentiable Search in Vision Foundation Models
Xi Xiao, Xingjian Li, ... (+8 more)
|
|
cs.CV
|
0 |
1 month ago |
| 142 |
MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation
Yang Chen, Xiaowei Xu, ... (+5 more)
|
|
cs.CV
|
0 |
1 month ago |
| 143 |
In-context Region-based Drag: Drag Any Region to Any Shape
Jiacheng Sui, Tianyu Hao, ... (+3 more)
|
|
cs.CV
|
0 |
1 month ago |
| 144 |
$S^{2}$-FracMix: Label-Preserving Self-Saliency Mixup Augmentation
Khawar Islam, Arif Mahmood, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 145 |
UniTeD: Unified Temporal Diffusion for Joint Perception and Planning in Autonomous Driving
Bo Zhao, Xinting Zhao, ... (+3 more)
|
|
cs.CV
|
0 |
1 month ago |
| 146 |
Falcon: Functional Assembly and Language for Compositional Reasoning in X-ray
Yonathan Michael, Mohamad Alansari, ... (+3 more)
|
|
cs.CV
|
0 |
1 month ago |
| 147 |
FeVOS: Foresight Expression Video Object Segmentation
Kehan Lan, Kaining Ying, Henghui Ding
|
|
cs.CV
|
0 |
1 month ago |
| 148 |
H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks
Seulgi Jeong, Yunseong Cho, Sanghun Park
|
|
cs.CV
|
0 |
1 month ago |
| 149 |
Distill on a Diet: Efficient Knowledge Distillation via Learnable Data Pruning
Yifan Wu, Yiqi Wang, ... (+6 more)
|
|
cs.LG
|
0 |
1 month ago |
| 150 |
C3-Bench: A Context-Aware Change Captioning Benchmark
Jae-Woo Kim, Hyeongbeom Kim, Ue-Hwan Kim
|
|
cs.CV
|
0 |
1 month ago |