Artificial Intelligence

Papers filed under cs.AI on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

14,641 to 14,700 of 15,326

  1. Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis

    Alona Strugatski, Licol Zeinfeld, Giora Alexandron

    cs.HCcs.AIcs.CLarXiv:2608.15630v12026
  2. NanoResearch: Co-Evolving Skills, Memory, and Policy for Personalized Research Automation

    Jinhang Xu, Qiyuan Zhu, Yujun Wu +11

    cs.AIarXiv:2605.10813v22026
  3. Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria

    Juanxi Tian, Fengyuan Liu, Jiaming Han +6

    cs.AIarXiv:2605.08354v12026
  4. AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment

    Yuchen Yuan, Zhenghuang Wu, Yuangan Li +2

    cs.AIarXiv:2608.16349v12026
  5. Baseline-Relative Counterfactual Refinement for Bit-Aware Visual Token Communication

    Jia Guo, Xiaohan Zhao, Changwang Liu +4

    cs.AIarXiv:2608.16192v12026
  6. OpenSTBench: Beyond Semantic Evaluation for Speech Translation

    Yanjie An, Yuxiang Zhao, Yichi Zhang +5

    eess.AScs.AIarXiv:2605.30792v12026
  7. Comprehensive Benchmarking of Deep Learning Architectures for Lung Cancer Histopathology

    Hadi Hasan, Safaa Salman, Lama Sleem +2

    cs.CVcs.AIarXiv:2608.15915v12026
  8. Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory

    Bingxin Xu, Yuzhang Shang, Emilio Ferrara

    cs.ROcs.AIcs.CVarXiv:2608.16889v12026
  9. Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs

    Siyuan Huang, Xiaoye Qu, Yafu Li +6

    cs.CVcs.AIarXiv:2605.00814v22026
  10. Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction

    Zhong Guan, Yongjian Guo, Haoran Sun +5

    cs.LGcs.AIarXiv:2605.12070v22026
  11. SkCC: Portable and Secure Skill Compilation for Cross-Framework LLM Agents

    Yipeng Ouyang, Yi Xiao, Yuhao Gu +1

    cs.CRcs.AIarXiv:2605.03353v42026
  12. HAGE: Harnessing Agentic Memory via RL-Driven Weighted Graph Evolution

    Dongming Jiang, Yi Li, Guanpeng Li +2

    cs.AIarXiv:2605.09942v12026
  13. Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models

    Yuanzhi Xu, Qian Gao, Jun Fan +4

    cs.CVcs.AIarXiv:2608.16805v12026
  14. JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

    Lin Song, Wenbo Li, Guoqing Ma +16

    cs.GRcs.AIcs.CLarXiv:2605.04128v22026
  15. Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State Uncertainty

    Joykirat Singh, Zaid Khan, Archiki Prasad +5

    cs.CLcs.AIarXiv:2605.11436v12026
  16. Steering the Flow: Inverting Face Recognition Models via Gradient-Guided Flow Matching

    Ye Lu, Shen Wang, Zhaoyang Zhang +4

    cs.CVcs.AIcs.CRarXiv:2608.16791v12026
  17. MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference

    Ruijie Zhou, Fanxu Meng, Yufei Xu +4

    cs.LGcs.AIarXiv:2605.07363v12026
  18. HarmTrace: Anchor-Calibrated Decoupled Optimization for Fine-Grained Target Identification in Harmful Memes

    Yujia Li, Yiqun Zhang, Zihan Cheng +7

    cs.CVcs.AIarXiv:2608.16622v12026
  19. Graph Machine Learning: An Opportunity for Power Systems

    Martin Sadric, Sebastian Pütz, Christian Nauck +4

    cs.LGcs.AIcs.CEarXiv:2608.16494v12026
  20. RagGAD: Rationale-Aware Conditional Gaussian Mixture Normalizing Flow for Unsupervised Graph Anomaly Detection

    Junxin Lu, Jing Zhao, Shiliang Sun

    cs.LGcs.AIarXiv:2608.16018v12026
  21. ATLAS: Scaffold-Free Algorithm Synthesis by LLMs via Embedding-Guided Quality-Diversity Search

    Danial Yazdani, Mohammad Nabi Omidvar, Yuan Sun +2

    cs.AIcs.NEarXiv:2608.15546v22026
  22. Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems

    Shihao Qi, Jie Ma, Rui Xing +15

    cs.AIarXiv:2605.14892v22026
  23. Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use

    Yize Cheng, Chenrui Fan, Mahdi JafariRaviz +2

    cs.AIarXiv:2605.14038v22026
  24. AgentFugue: Agent Scaling for Long-Horizon Tasks through Collective Reasoning

    Yuyang Hu, Hongjin Qian, Shuting Wang +5

    cs.AIcs.CLarXiv:2605.24486v12026
  25. Contrastive Energy Fields for Inference-Time Procedure Planning in Instructional Videos

    Mohamed Afham, Christoph Reich, Oliver Hahn +2

    cs.CVcs.AIarXiv:2608.16457v12026
  26. Graph Neural Assisted Actor-Critic for Latency-Efficient Edge Vision System

    Alam Noor, Luis Almeida, Kai Li +3

    cs.CVcs.AIarXiv:2608.16142v12026
  27. MUSE: An Interactive Meta-Agent for Understanding and Steering LLM-powered Data Science Systems

    Wei-Hao Chen, Weixi Tong, Yuan Tian +2

    cs.HCcs.AIarXiv:2608.16181v12026
  28. AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?

    Zhangchen Xu, Junda Chen, Yue Huang +16

    cs.AIcs.LGarXiv:2606.05080v12026
  29. Online Skill Learning for Web Agents via State-Grounded Dynamic Retrieval

    Jiaxi Li, Ke Deng, Yun Wang +5

    cs.AIarXiv:2606.04391v12026
  30. A Cognitively Motivated Multidimensional Framework for Evaluating Metaphor Explanations

    Ana Naveriani, Jakob Suchan, Stefano Zoia +3

    cs.CLcs.AIarXiv:2608.15828v12026
  31. Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation

    Che Liu, Lichao Ma, Xiangyu Tony Zhang +4

    cs.MMcs.AIcs.CVarXiv:2605.12034v22026
  32. KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving

    Zedong Liu, Xinyang Ma, Dejun Luo +9

    cs.DCcs.AIcs.NIarXiv:2605.13734v12026
  33. Toward AI-Friendly Cartography: Understanding How Color Design Influences Foundation Model Spatial Reasoning on Sequential Choropleth Maps

    Yonghe Sun, Zhenjia Liu, Hua Liao +4

    cs.AIarXiv:2608.15736v12026
  34. Look Before You Leap: Autonomous Exploration for LLM Agents

    Ziang Ye, Wentao Shi, Yuxin Liu +6

    cs.AIcs.CLarXiv:2605.16143v12026
  35. RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably

    Yufeng Du, Phillip Harris, Minyang Tian +5

    cs.CLcs.AIcs.LGarXiv:2605.15514v12026
  36. Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps

    Yanke Zhou, Yiduo Li, Hanlin Tang +6

    cs.CLcs.AIarXiv:2605.16928v22026
  37. ALKEMIE Agent: an autonomous platform for computational materials design

    Hongfu Huang, Yuzhe Li, Ao Xu +14

    cond-mat.mtrl-scics.AIarXiv:2608.15776v12026
  38. HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents

    Woongyeng Yeo, Yumin Choi, Taekyung Ki +1

    cs.LGcs.AIcs.CLarXiv:2605.17873v12026
  39. Generative Recursive Reasoning

    Junyeob Baek, Mingyu Jo, Minsu Kim +3

    cs.AIarXiv:2605.19376v22026
  40. Visualizing Uncertainty-to-Action Composition for Human Oversight

    Chisom Anyabolu, Akshat Dubey, Georges Hattab

    cs.HCcs.AIarXiv:2608.16428v12026
  41. HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction

    Jiahao Ji, Ji Ma, Runhan Zhang +12

    cs.ROcs.AIarXiv:2608.16222v12026
  42. A Policy Algebra for Trust-Preserving Agentic AI Execution

    Bhaskar Tripathi, Anurag Kumar, Ramendra Kumar +1

    cs.AIarXiv:2608.16402v12026
  43. Assessing LLMs' mathematical abilities requires understanding the various mechanisms of mathematical creativity

    Silvère Gangloff

    cs.AImath.HOarXiv:2608.16118v12026
  44. QUMem: Personalized Memory for Query-Conditioned User-State Inference in LLM Agents

    Heng Wang, Yifei Li, Lingling Zhang +4

    cs.CLcs.AIarXiv:2608.16168v12026
  45. From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

    Xingjian Wang, Zhao Wang, Taihang Hu +14

    cs.CVcs.AIarXiv:2608.18076v12026
  46. Counting Documents Is Not Counting Text: Unit Bias in Web-PDF Corpus Statistics

    Luca Foppiano

    cs.CLcs.AIarXiv:2608.16390v12026
  47. Dynamic Multi-Byte Prediction With Hierarchical Language Models

    Abraham Toluwase Owodunni, Chibuzor Okocha, Christan Grant +2

    cs.AIarXiv:2608.15454v12026
  48. Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems

    Jianing Zhu, Yeonju Ro, John Robertson +5

    cs.AIcs.CLcs.MAarXiv:2605.26302v12026
  49. VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions

    Yuxin Chen, Yi Zhang, Zhengzhou Cai +11

    cs.AIarXiv:2605.27141v12026
  50. PhoneWorld: Scaling Phone-Use Agent Environments

    Yuxuan Liu, Xin Lai, Junyi Li +21

    cs.CLcs.AIcs.LGarXiv:2605.29486v22026
  51. COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation

    Tianyi Zhou, Dongrui Liu, Leitao Yuan +2

    cs.AIcs.CLcs.LGarXiv:2605.31264v12026
  52. TRON: Targeted Rule-Verifiable Online Environments for Visual Reasoning RL

    Tianze Yang, Yucheng Shi, Ruitong Sun +3

    cs.AIarXiv:2606.01599v12026
  53. Walk Before You Run: The Importance of Data Exploration for Data Analysis Agents

    Yike Yuan, Virum Ranka, Tina Lasisi +1

    cs.DBcs.AIarXiv:2608.16045v12026
  54. When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Agents

    Jiawei Liu, Jiacheng Guo, Tian Zhang +7

    cs.ROcs.AIarXiv:2608.16806v12026
  55. NICE: Scale-Stable Perturbations for Graph Neural Network Explanations via Noise Corruption

    Ziluowen Luo, Jun Yin, Ruochen Liu +4

    cs.LGcs.AIarXiv:2608.16038v22026
  56. Software Engineering for AI-driven Building Operation

    Philipp Zech, Sascha Hammes, Johannes Weninger +2

    cs.SEcs.AIarXiv:2608.16237v12026
  57. Learning to Unlearn: Machine Unlearning via Learning the Unlearning Behaviors

    Hang Zhang, Kaifeng Zhang, Yixiao Ma +3

    cs.LGcs.AIarXiv:2608.16700v12026
  58. Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans

    Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani +2

    cs.CVcs.AIcs.CLarXiv:2608.16514v12026
  59. Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement

    Shenao Chen, Yidan Xu, Xiangmin Han +5

    cs.AIarXiv:2608.16628v12026
  60. StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

    Liya Zhu, Xin Ma, Tao Liu +35

    cs.AIarXiv:2608.17800v12026