| 101 |
Evaluating Reasoning Fidelity in Visual Text Generation
Jiajun Hong, Jiawei Zhou
|
|
cs.CV
|
0 |
1 month ago |
| 102 |
Ultra-Fast Neural Video Compression
Jiahao Li, Wenxuan Xie, ... (+5 more)
|
|
cs.CV
|
0 |
1 month ago |
| 103 |
Efficient and Training-Free Single-Image Diffusion Models
Haojun Qiu, Kiriakos N. Kutulakos, David B. Lindell
|
|
cs.CV
|
0 |
1 month ago |
| 104 |
A Cookbook of 3D Vision: Data, Learning Paradigms, and Application
Hongyang Du, Zongxia Li, ... (+9 more)
|
|
cs.CV
|
0 |
1 month ago |
| 105 |
Semantic Constraint Synthesis for Adaptive Trajectory Optimization via Large Language Models
Eleanor Brosius, Yuji Takubo, ... (+3 more)
|
|
math.OC
|
0 |
1 month ago |
| 106 |
Reflection Separation from a Single Image via Joint Latent Diffusion
Zheng-Hui Huang, Zhixiang Wang, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 107 |
Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking
Zekun Qi, Xuchuan Chen, ... (+11 more)
|
|
cs.RO
|
0 |
1 month ago |
| 108 |
Demo2Tutorial: From Human Experience to Multimodal Software Tutorials
Zechen Bai, Zhiheng Chen, ... (+6 more)
|
|
cs.CV
|
0 |
1 month ago |
| 109 |
Does Language Shift Break Medical Vision-Language Models? Indonesian Radiology Visual Question Answering Case Study
Pieter Christy Yan Yudhistira, Dzaki Rafif Malik, Novanto Yudistira
|
|
cs.CL
|
0 |
1 month ago |
| 110 |
Beyond Single Solution: Multi-Hypothesis Collaborative Deep Unfolding Network for Image Compressive Sensing
Wenxue Cui, Hualin Li, ... (+4 more)
|
|
cs.CV
|
0 |
1 month ago |
| 111 |
Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching
Hao Zhong, Muzhi Zhu, ... (+9 more)
|
|
cs.CV
|
0 |
1 month ago |
| 112 |
AvatarMix: Identity-Preserving Cross-Avatar Composition for Outfit Personalization
Zhaorong Wang, Yoshihiro Kanamori, Yuki Endo
|
|
cs.CV
|
0 |
1 month ago |
| 113 |
PersistGS: Differentiable Physics for Object Permanence in 4D Gaussian Splatting
Adrian Ramlal, John S. Zelek
|
|
cs.CV
|
0 |
1 month ago |
| 114 |
MARIO: Motion-Augmented Real-Time Multi-Sensor Inertial Odometry
Yiquan Li, Taeyoung Yeon, ... (+4 more)
|
|
cs.RO
|
0 |
1 month ago |
| 115 |
Hand Trajectory Fusion for Egocentric Natural Language Query Grounding
Enmin Zhong, Carlos R. del-Blanco, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 116 |
Latent Space Reinforcement Learning for Inverse Material Estimation in Food Fracture Simulation
Adrian Ramlal, Yuhao Chen, John S. Zelek
|
|
cs.CV
|
0 |
1 month ago |
| 117 |
VinQA: Visual Elements Interleaved Long-form Answer Generation for Real-World Multimodal Document QA
Young Rok Jang, Hyesoo Kong, ... (+4 more)
|
|
cs.CV
|
0 |
1 month ago |
| 118 |
Post-Launch Capability Expansion of Vision-Language Models via Prompting for On-Orbit Spacecraft Inspection
Nicholas A. Welsh, Lennon J. Shikhman, ... (+4 more)
|
|
cs.LG
|
0 |
1 month ago |
| 119 |
One Layer's Trash is Another Layer's Treasure: Adaptive Layer-wise Visual Token Selection in LVLMs
Yongru Chen, Kai Zhang, ... (+5 more)
|
|
cs.CV
|
0 |
1 month ago |
| 120 |
FloVerse: Floor Plan-Guided Multi-Modal Navigation
Weiqi Huang, Shuangyi Dong, ... (+4 more)
|
|
cs.RO
|
0 |
1 month ago |
| 121 |
Reasoning for Mobile User Experience with Multimodal LLMs: Task, Benchmark, and Approach
Ruichao Mao, Zhou Fang, ... (+10 more)
|
|
cs.AI
|
0 |
1 month ago |
| 122 |
Context-Aware Feature-Fusion for Co-occurring Object Detection in Autonomous Driving
Binay Kumar Singh, Niels Da Vitoria Lobo
|
|
cs.CV
|
0 |
1 month ago |
| 123 |
Metadata-Aware Multi-Prompt Reasoning for Zero-Shot Accident Understanding
Tarandeep Singh, Soumyanetra Pal, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 124 |
AutoMine Solution for AV2 2026 Scenario Mining Challenge
Songliang Cao, Jiele Zhao, ... (+11 more)
|
|
cs.AI
|
0 |
1 month ago |
| 125 |
Information-Theoretic Decomposition for Multimodal Interaction Learning
Zequn Yang, Yake Wei, ... (+3 more)
|
|
cs.LG
|
0 |
1 month ago |
| 126 |
MOFA-VTON: More Fashion Possibilities with Fine-Grained Adaptations in Virtual Try-On
Xiaoyu Han, Chenyang Wang, ... (+4 more)
|
|
cs.CV
|
0 |
1 month ago |
| 127 |
Monte Carlo Pass Search: Using Trajectory Generation for 3D Counterfactual Pass Evaluation in Football
Andrew Kang, Priya Narasimhan
|
|
cs.AI
|
0 |
1 month ago |
| 128 |
DB-3DME: From Dataset to Benchmark for Human-aligned Automatic 3D Mesh Evaluation
Nanshan Jia, Zhenyu Zhao, ... (+3 more)
|
|
cs.CV
|
0 |
1 month ago |
| 129 |
The Impact of VAE Design on Latent Pose Representations for Diffusion-based Sign Language Production
Guilhem Fauré, Mostafa Sadeghi, ... (+2 more)
|
|
cs.AI
|
0 |
1 month ago |
| 130 |
Reference-Free Assessment of Physical Consistency in World Model-based Video Generation
Yun Oh, Sukmin Yun
|
|
cs.AI
|
0 |
1 month ago |
| 131 |
Zero-Shot Vision-Language Models for Classroom Engagement Recognition: A Benchmark Study of Prompt Sensitivity and Cross-Dataset Generalization
Aman Goyal, Kshama Nitin Shah, Kemmannu Vineet Venkatesh Rao
|
|
cs.CV
|
0 |
1 month ago |
| 132 |
Beyond Flat Labels: Level-Restricted Contrastive Learning for Hierarchical Fine-Grained Vision Classification
Zhiyuan Tao, Srikumar Sastry, ... (+10 more)
|
|
cs.CV
|
0 |
1 month ago |
| 133 |
Scene-Level Heterogeneous Physics Simulation with 3D Gaussian Splats
Xiaoyang Liu, Shangzhe Wu, Kai Han
|
|
cs.GR
|
0 |
1 month ago |
| 134 |
Adversarial Domain Prompt Tuning and Generation for Single Domain Generalization
Zhipeng Xu, De Cheng, ... (+4 more)
|
|
cs.CV
|
0 |
1 month ago |
| 135 |
BIT-Nav: Brain-Inspired Trajectory Memory for Embodied Navigation
Rithvik Jonna, Aakash Gurram, ... (+3 more)
|
|
cs.RO
|
0 |
1 month ago |
| 136 |
Linear Recurrent Unit with Semantic Modulation for Image Super-Resolution
Mingyu Choi, Woo Kyoung Han, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 137 |
SCR-Guided Difficulty-Aware Optimization for Infrared Small Target Detection
Yunus Sevim, Behçet Uğur Töreyin
|
|
cs.CV
|
0 |
1 month ago |
| 138 |
InTrain: Intrinsic Trainability for Zero-Cost Neural Architecture Search
Qinqin Zhou, Fuhai Chen, ... (+4 more)
|
|
cs.LG
|
0 |
1 month ago |
| 139 |
SIR: Structured Image Representations for Explainable Robot Learning
Paul Mattes, Jan Schwab, ... (+6 more)
|
|
cs.RO
|
0 |
1 month ago |
| 140 |
Enhancing Part-Level Point Grounding for Any Open-Source MLLMs
Jin-Cheng Jhang, Fu-En Wang, ... (+5 more)
|
|
cs.CV
|
0 |
1 month ago |
| 141 |
TrafficAlign: Aligning Large Language Models for Traffic Scenario Generation
Zhi Tu, Liangkun Niu, Tianyi Zhang
|
|
cs.CV
|
0 |
1 month ago |
| 142 |
IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial Fusion
Lizhou Lin, Songpengcheng Xia, ... (+4 more)
|
|
cs.CV
|
0 |
1 month ago |
| 143 |
Perceptual 3D Simulation With Physical World Modeling
Wanhee Lee, Klemen Kotar, ... (+3 more)
|
|
cs.CV
|
0 |
1 month ago |
| 144 |
Dataset Usage Inference without Shadow Models or Held-out Data
Wojciech Łapacz, Stanisław Pawlak, ... (+3 more)
|
|
cs.LG
|
0 |
1 month ago |
| 145 |
GRAFT: Graph-Based Affordance Transfer via Part Correspondence
Mengying Lin, Utkarsh Mishra, ... (+2 more)
|
|
cs.RO
|
0 |
1 month ago |
| 146 |
Universal Guideline-Driven Image Clustering via a Hybrid LLM Agent
Wenliang Zhong, Rob Barton, ... (+8 more)
|
|
cs.CV
|
0 |
1 month ago |
| 147 |
Mirror Illusion Art
Xiaopei Zhu, Zeyuan Li, ... (+2 more)
|
|
cs.CV
|
0 |
28 days ago |
| 148 |
How Much Future Helps? A Controlled Study of Future-Privileged Supervision for Causal Egocentric Gaze Estimation
Jia Li, Wenjie Zhao, ... (+6 more)
|
|
cs.CV
|
0 |
29 days ago |
| 149 |
Local Motion Matters: A Deconstruct-Recompose Paradigm for Reinforcement Learning Pre-training from Videos
Jinwen Wang, Youfang Lin, ... (+3 more)
|
|
cs.LG
|
0 |
29 days ago |
| 150 |
Phantom: A Unified Face-Swap Deepfake Protection Framework with Latent and Spatial Constraints
Jungkon Kim, Cheolseung Jung, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |