| 1 |
VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition
Tanush Yadav, Mohammadreza Salehi, ... (+7 more)
|
|
cs.CV
|
0 |
2 months ago |
| 2 |
M\textsuperscript{4}Fuse: Lightweight State-Space MoE with a Cross-Scale Gating Bridge for Brain Tumor Segmentation
Meihua Zhou, Xinyu Tong, Li Yang
|
|
cs.CV
|
0 |
2 months ago |
| 3 |
Fine-Tuning Impairs the Balancedness of Foundation Models in Long-tailed Personalized Federated Learning
Shihao Hou, Chikai Shang, ... (+6 more)
|
|
cs.CV
|
0 |
2 months ago |
| 4 |
From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs
Le Zhang, Jihan Yang, ... (+12 more)
|
|
cs.CV
|
0 |
2 months ago |
| 5 |
CADFS: A Big CAD Program Dataset and Framework for Computer-Aided Design with Large Language Models
Vladislav Pyatov, Gleb Bobrovskikh, ... (+7 more)
|
|
cs.CV
|
0 |
2 months ago |
| 6 |
High-Fidelity Mobile Avatars with Pruned Local Blendshapes
Youyi Zhan, He Wang, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 7 |
Profile-Specific 3DMM Regression from a Single Lateral Face Image
Taiki Kanaya, Hideo Saito
|
|
cs.CV
|
0 |
2 months ago |
| 8 |
Multi-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and Reasoning
Sixian Zhang, Yiyao Wang, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 9 |
TrajRAG: Retrieving Geometric-Semantic Experience for Zero-Shot Object Navigation
Yiyao Wang, Sixian Zhang, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 10 |
Act2See: Emergent Active Visual Perception for Video Reasoning
Martin Q. Ma, Yuxiao Qu, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |
| 11 |
Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection
Tianxiao Li, Zhenglin Huang, ... (+11 more)
|
|
cs.CV
|
0 |
2 months ago |
| 12 |
Towards Visual Query Localization in the 3D World
Liang Peng, Bohan Tan, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |
| 13 |
Decision Boundary-aware Generation for Long-tailed Learning
Jiacheng Yang, Ruichi Zhang, ... (+6 more)
|
|
cs.CV
|
0 |
2 months ago |
| 14 |
Registration-Free Learnable Multi-View Capture of Faces in Dense Semantic Correspondence
Panagiotis P. Filntisis, George Retsinas, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 15 |
VISTA: Video Interaction Spatio-Temporal Analysis Benchmark
Alejandro Aparcedo, Akash Kumar, ... (+6 more)
|
|
cs.CV
|
0 |
2 months ago |
| 16 |
Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
Muyang Li, Yucheng Liu, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 17 |
Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs
Jingze Wu, Quan Zhang, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 18 |
CUE: Concept-Aware Multi-Label Expansion to Mitigate Concept Confusion in Long-Tailed Learning
Ruichi Zhang, Chikai Shang, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |
| 19 |
A Large-Scale Study on the Accuracy vs Cost Trade-offs of Training and Evaluation Settings in Fine-Grained Image Recognition
Edwin Arkel Rios, Augusto Christian Surya, ... (+6 more)
|
|
cs.CV
|
0 |
2 months ago |
| 20 |
CMAG: Concept-Scaffolded Retrieval for Marketplace Avatar Generation
Rajeev Goel, Jason Ding, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 21 |
Dance Across Shifts: Forward-Facilitation Continual Test-Time Adaptation through Dynamic Style Bridging
Zhilin Zhu, Yabin Wang, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 22 |
RT-Splatting: Joint Reflection-Transmission Modeling with Gaussian Splatting
Ji Shi, Xianghua Ying, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 23 |
SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning
Pawat Chunhachatrachai, Gueter Josmy Faure, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 24 |
Best Segmentation Buddies for Image-Shape Correspondence
Itai Lang, Dongwei Lyu, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 25 |
View-Aware Semantic Alignment for Aerial-Ground Person Re-Identification
Quan Zhang, Zeqiang Cai, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |
| 26 |
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
Boyuan Sun, Bowen Yin, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 27 |
A More Word-like Image Tokenization for MLLMs
Hyun Lee, Hyemin Jeong, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |
| 28 |
UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation
Tianhao Han, Haoyang Zhang, ... (+6 more)
|
|
cs.CV
|
0 |
2 months ago |
| 29 |
StyleText: A Large-Scale Dataset and Benchmark for Stylized Scene Text Inpainting
Aleksandr Simonyan, Nipun Jindal
|
|
cs.CV
|
0 |
2 months ago |
| 30 |
Is VLA Reasoning Faithful? Probing Safety of Chain-of-Causation
Nicanor Mayumu, Xiaoheng Deng, Patrick Mukala
|
|
cs.AI
|
0 |
2 months ago |
| 31 |
RHINO: Reconstructing Human Interactions with Novel Objects from Monocular Videos
Lixin Xue, Chengwei Zheng, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |
| 32 |
HAD: Hallucination-Aware Diffusion Priors for 3D Reconstruction
Xi Liu, Weiwei Sun, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 33 |
Metric-Guided Feature Fusion of Visual Foundation Models for Segmentation Tasks
Yachan Guo, JoseLuis Gomez Zurita, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 34 |
Cross-modal Affinity-aligned Multimodal Learning Analytics for Predicting Student Collaboration Satisfaction in Game-Based Learning
Wen-Hsin Tsai, Chia-Ming Lee, Yuk-Ying Tung
|
|
cs.LG
|
0 |
2 months ago |
| 35 |
ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned Representation
Qing Huang, Zhipei Xu, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 36 |
Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?
Renye Yan, Jikang Cheng, ... (+8 more)
|
|
cs.CV
|
0 |
2 months ago |
| 37 |
Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models
Yujun Tong, Dongliang Chang, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 38 |
How to Choose Your Teacher for Fine Grained Image Recognition
Oswin Gosal, Edwin Arkel Rios, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 39 |
Pretraining Objective Matters in Extreme Low-Data FGVC: A Backbone-Controlled Study
Alexander Hackett, Srikanth Thudumu, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 40 |
U-SEG: Uncertainty in SEGmentation -- A systematic multi-variable exploration
Michael Smith, Frank P. Ferrie
|
|
cs.CV
|
0 |
2 months ago |
| 41 |
CT-DegradBench: A Physics-Informed Benchmark for CT Degradation Detection and Severity Estimation
Yousra Nabila Taifour, Marouane Tliba, ... (+10 more)
|
|
cs.CV
|
0 |
2 months ago |
| 42 |
VGGT-$Ω$
Jianyuan Wang, Minghao Chen, ... (+8 more)
|
|
cs.CV
|
0 |
2 months ago |
| 43 |
SuperADD: Training-free Class-agnostic Anomaly Segmentation -- CVPR 2026 VAND 4.0 Workshop Challenge Industrial Track
Lukas Roming, Felix Lehnerer, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |
| 44 |
Global Structure-from-Motion Meets Feedforward Reconstruction
Linfei Pan, Johannes Schönberge, Marc Pollefeys
|
|
cs.CV
|
0 |
2 months ago |
| 45 |
Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot Learning
Shuai Yi, Yixiong Zou, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 46 |
OMGTex: One-stage Multi-style Facial Texture Reconstruction without Geometry Guidance
Zitong Xiao, Yuda Qiu, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 47 |
From Contrast to Consistency: Rethinking Event-based Continuous-Time Optical Flow Estimation
Rui Hu, Song Wu, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 48 |
ComPose: A Unified Completion-Pose Framework for Robust Category-Level Object Pose Estimation
Huan Ren, Yihan Chen, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 49 |
Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video Diffusion
Ting-Hsuan Chen, Ying-Huan Chen, ... (+11 more)
|
|
cs.CV
|
0 |
2 months ago |
| 50 |
MTLLFM: Multimodal-Temporal Laughter Localization: UR-FUNNY-Temporal and SMILE-Temporal Benchmarks with an Adaptive Multimodal Fusion Model
Eyal Hanania, Nadav Kirsch, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |