Artificial Intelligence

Papers filed under cs.AI on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,961 to 4,020 of 15,235

  1. OPSDL: On-Policy Self-Distillation for Long-Context Language Models

    Xinsen Zhang, Zhenkai Ding, Tianjun Pan +4

    cs.CLcs.AIarXiv:2604.17535v12026
  2. TAAL: Mitigating Early Beam Pruning in Generative Recommendation via Temporal Autoregressive Alignment

    Lianjie Li, Zhiying Tu, Dianhui Chu +1

    cs.IRcs.AIarXiv:2608.29179v12026
  3. HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation

    Pei Liu, Xin Liu, Ruoyu Yao +4

    cs.CLcs.AIarXiv:2504.12330v12025
  4. MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

    Yixing Jiang, Kameron C. Black, Gloria Geng +4

    cs.LGcs.AIcs.MAarXiv:2501.14654v22025
  5. SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems

    Yunhao Feng, Yifan Ding, Yingshui Tan +6

    cs.CRcs.AIarXiv:2604.06811v22026
  6. Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models

    Yijia Shao, Yucheng Jiang, Theodore A. Kanell +3

    cs.CLcs.AIarXiv:2402.14207v22024
  7. Auditing and Mitigating Privacy Leakage in Cloud-Edge Collaborative Decoding

    Kejia Zhang, Tianyuan Zou, Zixuan GU +1

    cs.CRcs.AIarXiv:2608.29111v12026
  8. Interpreting CLIP with Hierarchical Sparse Autoencoders

    Vladimir Zaigrajew, Hubert Baniecki, Przemyslaw Biecek

    cs.CVcs.AIcs.LGarXiv:2502.20578v22025
  9. Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

    Siyu Yuan, Zehui Chen, Zhiheng Xi +3

    cs.AIarXiv:2501.11425v32025
  10. A User Simulator for Task-Completion Dialogues

    Xiujun Li, Zachary C. Lipton, Bhuwan Dhingra +3

    cs.LGcs.AIcs.CLarXiv:1612.05688v32016
  11. DeepStory: Video Story QA by Deep Embedded Memory Networks

    Kyung-Min Kim, Min-Oh Heo, Seong-Ho Choi +1

    cs.CVcs.AIcs.CLarXiv:1707.00836v12017
  12. RAGDiffusion++: From Macro-Retrieval to Micro-Fidelity Alignment for Garment Generation

    Yuhan Li, Xianfeng Tan, Fangao Zeng +6

    cs.CVcs.AIarXiv:2608.29280v12026
  13. APIFlow-Bench: Measuring Whether Agents Survive Long, Dependent API Workflows

    Zelin Wan, Arash Nourian, Xiaoxiao Li +2

    cs.AIcs.LGcs.SEarXiv:2608.29128v12026
  14. LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?

    Kexian Tang, Junyao Gao, Yanhong Zeng +6

    cs.AIarXiv:2503.19990v42025
  15. D2A: A Dataset Built for AI-Based Vulnerability Detection Methods Using Differential Analysis

    Yunhui Zheng, Saurabh Pujar, Burn Lewis +6

    cs.SEcs.AIcs.LGarXiv:2102.07995v12021
  16. VideoRAG: Retrieval-Augmented Generation over Video Corpus

    Soyeong Jeong, Kangsan Kim, Jinheon Baek +1

    cs.CVcs.AIcs.CLarXiv:2501.05874v32025
  17. A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models

    Dong Shu, Xuansheng Wu, Haiyan Zhao +4

    cs.LGcs.AIcs.CLarXiv:2503.05613v32025
  18. Wavelet-Assisted Multi-Frequency Attention Network for Pansharpening

    Jie Huang, Rui Huang, Jinghao Xu +3

    eess.IVcs.AIcs.CVarXiv:2502.04903v12025
  19. A Survey on Large Language Models for Mathematical Reasoning

    Peng-Yuan Wang, Tian-Shuo Liu, Chenyang Wang +8

    cs.AIcs.CLarXiv:2506.08446v12025
  20. Integrating LLMs with ITS: Recent Advances, Potentials, Challenges, and Future Directions

    Doaa Mahmud, Hadeel Hajmohamed, Shamma Almentheri +4

    eess.SYcs.AIcs.ETarXiv:2501.04437v12025
  21. Language Models Use Trigonometry to Do Addition

    Subhash Kantamneni, Max Tegmark

    cs.AIcs.CLcs.LGarXiv:2502.00873v12025
  22. AI Literacy in K-12 and Higher Education in the Wake of Generative AI: An Integrative Review

    Xingjian Gu, Barbara J. Ericson

    cs.CYcs.AIarXiv:2503.00079v32025
  23. Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities

    Alexander Nikitin, Jannik Kossen, Yarin Gal +1

    cs.LGcs.AIcs.CLarXiv:2405.20003v12024
  24. Formal Policy Enforcement for Real-World Agentic Systems

    Nils Palumbo, Sarthak Choudhary, Jihye Choi +3

    cs.CRcs.AIcs.MAarXiv:2602.16708v32026
  25. Neural Language Modeling by Jointly Learning Syntax and Lexicon

    Yikang Shen, Zhouhan Lin, Chin-Wei Huang +1

    cs.CLcs.AIarXiv:1711.02013v22017
  26. MinTL: Minimalist Transfer Learning for Task-Oriented Dialogue Systems

    Zhaojiang Lin, Andrea Madotto, Genta Indra Winata +1

    cs.CLcs.AIarXiv:2009.12005v22020
  27. MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

    Shrey Pandit, Jiawei Xu, Junyuan Hong +4

    cs.CLcs.AIcs.LGarXiv:2502.14302v12025
  28. Meta-Learning by Adjusting Priors Based on Extended PAC-Bayes Theory

    Ron Amit, Ron Meir

    stat.MLcs.AIcs.LGarXiv:1711.01244v82017
  29. A Survey on the Optimization of Large Language Model-based Agents

    Shangheng Du, Jiabao Zhao, Jinxin Shi +4

    cs.AIarXiv:2503.12434v22025
  30. R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO

    Huanjin Yao, Qixiang Yin, Jingyi Zhang +8

    cs.CVcs.AIcs.CLarXiv:2505.16673v12025
  31. Perceptive Humanoid Parkour: Chaining Dynamic Human Skills via Motion Matching

    Zhen Wu, Xiaoyu Huang, Lujie Yang +8

    cs.ROcs.AIcs.LGarXiv:2602.15827v22026
  32. The Assistant's Ideal Self

    Mert Yazan

    cs.AIarXiv:2609.00304v12026
  33. When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail

    Xiaoxiao Li

    cs.AIcs.MAarXiv:2601.04748v22026
  34. BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval

    Hongjin Su, Howard Yen, Mengzhou Xia +12

    cs.CLcs.AIcs.IRarXiv:2407.12883v42024
  35. Advancing Multi-Agent Systems Through Model Context Protocol: Architecture, Implementation, and Applications

    Naveen Krishnan

    cs.MAcs.AIarXiv:2504.21030v12025
  36. STAMP: Scalable Task And Model-agnostic Collaborative Perception

    Xiangbo Gao, Runsheng Xu, Jiachen Li +3

    cs.CVcs.AIcs.ROarXiv:2501.18616v12025
  37. Beyond Token Positions: Safety Alignment Across Denoising Steps in Diffusion Language Models

    Guoli Wang, Haonan Shi, Tu Ouyang +1

    cs.CLcs.AIarXiv:2609.00495v12026
  38. Operational Regimes in Non-Convex Optimization: A Multiplier-Based Taxonomy

    Seyed Mohsen Kazemi, Ali Movaghar, Shaahin hessabi

    math.OCcs.AIeess.SParXiv:2609.00471v12026
  39. Securing Agentic AI: A Comprehensive Threat Model and Mitigation Framework for Generative AI Agents

    Vineeth Sai Narajala, Om Narayan

    cs.CRcs.AIarXiv:2504.19956v22025
  40. Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict

    Yihang Chen, Pin Qian, Su Wang +4

    cs.CLcs.AIarXiv:2605.14473v42026
  41. Are Language Models Models?

    Philip Resnik

    cs.CLcs.AIarXiv:2601.10421v12026
  42. The Impact of Reasoning Step Length on Large Language Models

    Mingyu Jin, Qinkai Yu, Dong Shu +5

    cs.CLcs.AIarXiv:2401.04925v42024
  43. Life-Cycle Emissions of AI Hardware: A Cradle-To-Grave Approach and Generational Trends

    Ian Schneider, Hui Xu, Stephan Benecke +4

    cs.ARcs.AIarXiv:2502.01671v12025
  44. Centering before Pruning: Lightweight Geometry Correction for Diversity-Based Visual Token Pruning in LVLMs

    Shunjie Wen, Jaeyeon Lee, Dong-Wan Choi

    cs.CVcs.AIarXiv:2608.30263v12026
  45. Parameter-Efficient Fine-Tuning for Foundation Models

    Dan Zhang, Tao Feng, Lilong Xue +3

    cs.CLcs.AIcs.LGarXiv:2501.13787v12025
  46. NVIDIA FLARE: Federated Learning from Simulation to Real-World

    Holger R. Roth, Yan Cheng, Yuhong Wen +20

    cs.LGcs.AIcs.CVarXiv:2210.13291v32022
  47. EpaCache: Error-Propagation-Aware Caching for Accelerating Diffusion-Based Visual Generation

    Yuhan Liu, Zongwei Hong, Jinglun Li +3

    cs.AIarXiv:2608.29264v12026
  48. TIPS: Turn-Level Information-Potential Reward Shaping for Search-Augmented LLMs

    Yutao Xie, Nathaniel Thomas, Nicklas Hansen +3

    cs.CLcs.AIcs.LGarXiv:2603.22293v12026
  49. Augmenting Human Performance with an XR Agent Learning from Online Behavior and BCI Evidence

    Ziheng Li, Xichen He, Haoyan Chen +10

    cs.AIcs.HCarXiv:2608.30369v12026
  50. Incompressible Knowledge Probes: Estimating Black-Box LLM Parameter Counts via Factual Capacity

    Bojie Li

    cs.LGcs.AIarXiv:2604.24827v22026
  51. School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

    Mia Taylor, James Chua, Jan Betley +2

    cs.AIarXiv:2508.17511v12025
  52. Thinking with Generated Images

    Ethan Chern, Zhulin Hu, Steffi Chern +5

    cs.CVcs.AIcs.CLarXiv:2505.22525v12025
  53. Astra: A Multi-Agent System for GPU Kernel Performance Optimization

    Anjiang Wei, Tianran Sun, Yogesh Seenichamy +5

    cs.DCcs.AIcs.CLarXiv:2509.07506v22025
  54. General In-Hand Object Rotation with Vision and Touch

    Haozhi Qi, Brent Yi, Sudharshan Suresh +4

    cs.ROcs.AIcs.CVarXiv:2309.09979v22023
  55. Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents

    Tianshi Xu, Huifeng Wen, Meng Li

    cs.AIarXiv:2605.22166v22026
  56. Position: Graph Learning Will Lose Relevance Due To Poor Benchmarks

    Maya Bechler-Speicher, Ben Finkelshtein, Fabrizio Frasca +9

    cs.LGcs.AIcs.NEarXiv:2502.14546v12025
  57. Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization

    Jiantong Jiang, Peiyu Yang, Rui Zhang +1

    cs.LGcs.AIcs.CLarXiv:2607.08057v12026
  58. A Generative Deep Learning Approach to Stochastic Downscaling of Precipitation Forecasts

    Lucy Harris, Andrew T. T. McRae, Matthew Chantry +2

    physics.ao-phcs.AIcs.CVarXiv:2204.02028v22022
  59. More Perspectives, Stronger Signals: Multi-Perspective Enhancement and Progressive Fusion for Multimodal Entity Representation Learning

    Chenyi Xiong, Yan Zhang, Jing Hu +5

    cs.AIarXiv:2608.29139v12026
  60. Agent Behavioral Contracts: Formal Specification and Runtime Enforcement for Reliable Autonomous AI Agents

    Varun Pratap Bhardwaj

    cs.AIcs.MAcs.SEarXiv:2602.22302v12026