| 151 |
Auditing MCQA Benchmarks through Probability Landscapes
Minsoo Song, Chanjun Park
|
|
cs.CL
|
0 |
16 days ago |
| 152 |
Will the User Ever Know? Covert Indirect Prompt Injection on Tool-Using LLM Agents
Yunseok Lee, Yunji Kim, Woojin Lee
|
|
cs.AI
|
0 |
16 days ago |
| 153 |
Answer Probing-Guided Search for Diverse Solution Exploration of LLMs
Yi Fang, Que Shen, ... (+7 more)
|
|
cs.AI
|
0 |
16 days ago |
| 154 |
Beyond Token-Level Guidance: Inference-Time Alignment of Specialized LLMs via Cross-Family Representation Steering
Jin Gan, Xin Li, Jun Luo
|
|
cs.CL
|
0 |
16 days ago |
| 155 |
Lazy Grounding: Attacking Search Agents with Factual Evidence
Yulin Zhang, Yukun Huang, ... (+5 more)
|
|
cs.CL
|
0 |
16 days ago |
| 156 |
Dynamic Hub-and-Spoke Memory for Streaming Video Understanding
Xinru Jiang, Lin Zhao, ... (+8 more)
|
|
cs.CV
|
0 |
16 days ago |
| 157 |
PEARL: Front-Loading Relational Chains for Multi-Hop Table Retrieval
Subeen Ho, Hyeongu Kang, ... (+2 more)
|
|
cs.IR
|
0 |
16 days ago |
| 158 |
Using Prosody to Predict Syntactic Structure
Junghyun Min, Alex Warstadt, ... (+3 more)
|
|
cs.CL
|
0 |
16 days ago |
| 159 |
Beyond Surface Forms: Symbolic Edits as a Test for Logical Reasoning with LLMs
Ramya Keerthy Thatikonda, Wray Buntine, Ehsan Shareghi
|
|
cs.CL
|
0 |
16 days ago |
| 160 |
Strong Drafts Need Compact Memories: Long-Context Speculative Decoding with Compressed KV Cache
Tong Yuan, Chengxi Liao, Zeyi Wen
|
|
cs.LG
|
0 |
16 days ago |
| 161 |
DSEffi-Bench: Demystifying Large Language Models' Capability in Efficient Data Science Code Generation
Zhihao Gong, Junzhe Yu, ... (+4 more)
|
|
cs.SE
|
0 |
16 days ago |
| 162 |
Quantifying and Mitigating Korean Jamo-Level Typographical Vulnerabilities in Large Language Models
Seojin Lee, Hwanhee Lee
|
|
cs.CL
|
0 |
16 days ago |
| 163 |
LaMoC: Loss-Aware Modular Compression for LLMs
Mohanad Odema, Jacob Song
|
|
cs.AI
|
0 |
16 days ago |
| 164 |
When Models Hear What They Expect: Diagnosing Prosodic Heuristics in Multimodal Sarcasm Detection
Yongjian Chen, Pengfei Wei, ... (+3 more)
|
|
cs.CL
|
0 |
16 days ago |
| 165 |
ALTSTEER: Selective Safety Steering for Moving Beyond Hard Refusals to Constructive Alternatives
Hoejoon Kwon, Byeonggeuk Lim, ... (+2 more)
|
|
cs.CL
|
0 |
16 days ago |
| 166 |
Doc-REFRAG: Rethinking Multimodal Document Retrieval-Augmented Generation
Ruofan Hu, Shengyang Xu, ... (+6 more)
|
|
cs.IR
|
0 |
16 days ago |
| 167 |
Reactivating Test-Time Scaling for Plane Geometry Problem Solving
Xiaoqiang Kang, Shengen Wu, ... (+6 more)
|
|
cs.CL
|
0 |
16 days ago |
| 168 |
CAST: Critique-Aware Supervision for Training Reliable Long-Horizon Tool-Calling Agents
Amir Saeidi, Zehua Zhang, ... (+7 more)
|
|
cs.CL
|
0 |
16 days ago |
| 169 |
Rethinking Language's Role in Efficient VLA for Autonomous Vehicles: Toward Smarter, Trustworthy Driving
Tongfei Guo, Lili Su
|
|
cs.RO
|
0 |
16 days ago |
| 170 |
COGTRL: Training LLMs for Scientific Discovery Assistance using Cognitive Traces via Reinforcement Learning
Shrinidhi Kumbhar Santosh Mashetty Divij Handa Kevin Coutinho, Siddharth Sambhaji Ghule, Chitta Baral
|
|
cs.CL
|
0 |
16 days ago |
| 171 |
Budget-Aware Compression Pipeline for Single-GPU LLM Inference: Methods, Trade-offs, and Coupling Effects
Hongyu Yu, Yifei Shen
|
|
cs.CL
|
0 |
17 days ago |
| 172 |
How do World Models and Policies Compose in LLM Agents? A Joint Spectral and Behavioral Account
Ruize Xu, Xiao Yu, ... (+3 more)
|
|
cs.LG
|
0 |
17 days ago |
| 173 |
"Act Like a 5th Grader" is Not Enough: Bounding Knowledge in LLM-Based User Simulators
Krisztian Balog, Arild Michel Bakken
|
|
cs.CL
|
0 |
17 days ago |
| 174 |
Interpreting and Steering for Safe and Correct Code Generation
Hao Yan, Ziyu Yao
|
|
cs.AI
|
0 |
17 days ago |
| 175 |
Small Language Models as Judges for Rubric-Based Reinforcement Learning
Fengyu Xie, Yilun Zhao, ... (+3 more)
|
|
cs.CL
|
0 |
17 days ago |
| 176 |
TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models
Apoorva Kulkarni, Kaousheik Jayakumar, ... (+4 more)
|
|
cs.SD
|
0 |
17 days ago |
| 177 |
Generating Clinical Vignettes that Preserve Cognitive Formulations
Amit Oren, Nimrod Hertz-Palmor, ... (+2 more)
|
|
cs.CL
|
0 |
17 days ago |
| 178 |
AutoCRAT: Within-trajectory Joint Control of Stochasticity and Compute for LLM Reasoning
Hanjun Luo, Qiushi Liu, ... (+9 more)
|
|
cs.AI
|
0 |
17 days ago |
| 179 |
Evolutionary Soups: Evolving Mixture-of-Experts for Multi-Objective LLM Alignment
Lingxiao Kong, Steffen Staab, ... (+3 more)
|
|
cs.CL
|
0 |
17 days ago |
| 180 |
RIDGE: Region-Informed Derivative-Guided Evidence Selection for Long Video Understanding
Shanqing Xu, Meng Luo, ... (+8 more)
|
|
cs.CV
|
0 |
17 days ago |
| 181 |
XQDT: eXplainable and Quantitative Data-Text Alignment Metric with Feedback Signals
Kun Efimov-Zhang, Yifei Song, Claire Gardent
|
|
cs.CL
|
0 |
17 days ago |
| 182 |
Compression-Aware Abstention: Teaching LLMs to Refuse When KV-Compression Masks Remove Answer Evidence
Mohammadali Khodabandehlou, Bhaskar Krishnamachari
|
|
cs.CL
|
0 |
17 days ago |
| 183 |
When Less is More: Understanding When Token Filtering Helps and Fails in AI-generated Text Detection
Xiaoyang Han, Lvxiaowei Xu, Ming Cai
|
|
cs.CL
|
0 |
17 days ago |
| 184 |
REIGN: Refurbished Embeddings with Integrated Guidance Networks for Efficient Context-Length Scaling
Devrim Çavuşoğlu, Emre Akbaş
|
|
cs.CL
|
0 |
17 days ago |
| 185 |
En-ViMedNER: An English-Vietnamese Parallel Biomedical Corpus with UMLS Semantic Type Annotations
Nhu Vo, Phuong Nguyen, ... (+5 more)
|
|
cs.CL
|
0 |
17 days ago |
| 186 |
VibeJam: A User Study Platform for Web Development with Agents
Nishant Balepur, Connor Baumler, ... (+4 more)
|
|
cs.CL
|
0 |
17 days ago |
| 187 |
Check The Scoreboard: An Analysis of Scoring Schemes on Multiple-Choice Evaluation
Nishant Balepur, Paiheng Xu, ... (+4 more)
|
|
cs.CL
|
0 |
17 days ago |
| 188 |
Improving Argument Saliency Coverage in Small LLMs for Long Legal Opinion Summarization via Sequence-Level Distillation
Mohamed Elaraby, Ahmed Elhady, Diane Litman
|
|
cs.CL
|
0 |
17 days ago |
| 189 |
ManGo: Manga Active Narrative Grounding Optimization
Hao Qiu, Junyan Wang, ... (+5 more)
|
|
cs.CL
|
0 |
17 days ago |
| 190 |
ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models
Shaghayegh Kolli, Sina Emami, ... (+5 more)
|
|
cs.CV
|
0 |
17 days ago |
| 191 |
SkillForge: Compositional Skill Synthesis with Verification-in-the-Loop for Generating Formally Verified Dafny Programs
Yanming Liu, Xinyue Peng, ... (+3 more)
|
|
cs.CL
|
0 |
17 days ago |
| 192 |
EVAR: Evidence-Validated Hypothesis Admission for Budget-Aware Narrative Reasoning
Peilin Liu, Zhiquan Ji, Jinglong Ping
|
|
cs.CL
|
0 |
17 days ago |
| 193 |
You Know What I Mean: A Benchmark for Agentic Conversational Reference Grounding
Karen Fuchs, Uri Katz, Yoav Goldberg
|
|
cs.CL
|
0 |
17 days ago |
| 194 |
A^2Agent: Action-Aware Reinforcement Learning for Repository-Level Code Localization Agents
Doyeon Kim, Suyoung Bae, ... (+2 more)
|
|
cs.CL
|
0 |
17 days ago |
| 195 |
HiVe: Beyond Static Prompts for Multitask Learning via Hierarchy-based Vertical Mixture-of-Experts
HyeonJik Bae, Minyeol Kim, Susik Yoon
|
|
cs.CL
|
0 |
17 days ago |
| 196 |
Higher-Dimensional Rotary Position Embedding
Yixing Li, Ruobing Xie, ... (+4 more)
|
|
cs.LG
|
0 |
17 days ago |
| 197 |
Last Step Matters: Early Uncertainty Cannot Predict Failure in Long-Horizon Agents
Zongyue Li, Chengyue Yu, ... (+4 more)
|
|
cs.LG
|
0 |
17 days ago |
| 198 |
A Target-Centric Survey of Quantization-Aware Training
Jiamin Song, Mengjie Zhao, ... (+7 more)
|
|
cs.LG
|
0 |
17 days ago |
| 199 |
ACTD: Anchor-Based Cross-Tokenizer Distillation with Residual Regularization
Huiyi Zhang, Zijian Li, ... (+5 more)
|
|
cs.CL
|
0 |
17 days ago |
| 200 |
PrivBench: A Holistic and Modular Benchmarking Platform for Evaluating Text-to-Text Privatization
Stephen Meisenbacher, Andreea-Elena Bodea, ... (+4 more)
|
|
cs.CL
|
0 |
17 days ago |