Artificial Intelligence

Papers filed under cs.AI on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,541 to 3,600 of 15,439

  1. Wink: Recovering from Misbehaviors in Coding Agents

    Rahul Nanda, Chandra Maddila, Smriti Jha +3

    cs.SEcs.AIcs.HCarXiv:2602.17037v22026
  2. Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect Judges

    Chen Feng, Minghe Shen, Ananth Balashankar +2

    cs.LGcs.AIcs.CVarXiv:2601.20913v12026
  3. Label Efficient Semi-Supervised Learning via Graph Filtering

    Qimai Li, Xiao-Ming Wu, Han Liu +2

    cs.LGcs.AIstat.MLarXiv:1901.09993v32019
  4. A Systematic Review for Transformer-based Long-term Series Forecasting

    Liyilei Su, Xumin Zuo, Rui Li +3

    cs.LGcs.AIarXiv:2310.20218v12023
  5. Trajectory-Informed Memory Generation for Self-Improving Agent Systems

    Gaodan Fang, Vatche Isahagian, K. R. Jayaram +4

    cs.AIcs.DBcs.IRarXiv:2603.10600v12026
  6. Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents

    Zehong Wang, Fang Wu, Hongru Wang +8

    cs.AIcs.CLcs.LGarXiv:2601.22311v12026
  7. GUI-GENESIS: Automated Synthesis of Efficient Environments with Verifiable Rewards for GUI Agent Post-Training

    Yuan Cao, Dezhi Ran, Mengzhou Wu +9

    cs.AIcs.LGarXiv:2602.14093v12026
  8. M2A: Multimodal Memory Agent with Dual-Layer Hybrid Memory for Long-Term Personalized Interactions

    Junyu Feng, Binxiao Xu, Jiayi Chen +8

    cs.AIarXiv:2602.07624v12026
  9. AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition

    Ruipeng Wang, Yuxin Chen, Yukai Wang +9

    cs.AIarXiv:2602.11348v22026
  10. Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning

    Yihong Huang, Fei Ma, Yihua Shao +4

    cs.CVcs.AIcs.CLarXiv:2602.02951v12026
  11. Free Speech and Artificial Intelligence

    Etienne Brown

    cs.CYcs.AIarXiv:2608.28973v12026
  12. Not the Same Protector: Deployment-Dependent Protective Intervention in LLMs

    Eunna Lee, Soomyoung Lee, Jungpyo Nam +6

    cs.CRcs.AIarXiv:2608.29136v12026
  13. Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks

    Ang Li, Yin Zhou, Vethavikashini Chithrra Raghuram +2

    cs.LGcs.AIarXiv:2502.08586v12025
  14. Limits of trust in medical AI

    Joshua Hatherley

    cs.LGcs.AIcs.CYarXiv:2503.16692v22025
  15. Development of an Autonomous AI Coding Agent using Monte Carlo Tree Search (MCTS) and Gemini LLM Frameworks

    Pravin Game, Vipin Ramakrishnan, Prathamesh Wagh

    cs.LGcs.AIarXiv:2608.29096v12026
  16. An effective algorithm for hyperparameter optimization of neural networks

    Gonzalo Diaz, Achille Fokoue, Giacomo Nannicini +1

    cs.AIcs.LGcs.NEarXiv:1705.08520v12017
  17. Automatic Conversion of NICE Guidelines to an Executable Computational Model Using Large Language Models

    Ashvin Gupta, Denys Prociuk, Alessandra Russo +1

    cs.AIarXiv:2608.30022v12026
  18. MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks

    Zixuan Ke, Yifei Ming, Austin Xu +7

    cs.AIcs.CLcs.MAarXiv:2601.14652v52026
  19. 6G-Bench: An Open Benchmark for Semantic Communication and Network-Level Reasoning with Foundation Models in AI-Native 6G Networks

    Mohamed Amine Ferrag, Abderrahmane Lakas, Merouane Debbah

    cs.NIcs.AIarXiv:2602.08675v12026
  20. A Generalized Optimization Engine (GOE) for Edge AI Inference Acceleration

    Venkat R. Dasari, Jakob A. Adams, Vinod K. Mishra +1

    cs.AIarXiv:2608.28652v12026
  21. Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts

    Saman Rahbar, Xiliang Zhu, Irvin Cardoza +1

    cs.CLcs.AIcs.LGarXiv:2609.00330v12026
  22. Large-Scale Optimization Model Auto-Formulation: Harnessing LLM Flexibility via Structured Workflow

    Kuo Liang, Yuhang Lu, Jianming Mao +7

    cs.AIcs.LGarXiv:2601.09635v32026
  23. From Flat Logs to Causal Graphs: Hierarchical Failure Attribution for LLM-based Multi-Agent Systems

    Yawen Wang, Wenjie Wu, Junjie Wang +1

    cs.AIcs.SEarXiv:2602.23701v12026
  24. MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning

    Xiaoliang Fu, Jiaye Lin, Yangyi Fang +7

    cs.LGcs.AIarXiv:2602.17550v32026
  25. Interpretable All-Type Audio Deepfake Detection with Audio LLMs via Frequency-Time Reinforcement Learning

    Yuankun Xie, Xiaoxuan Guo, Jiayi Zhou +6

    cs.SDcs.AIarXiv:2601.02983v12026
  26. How Identity and Opinion Shape Political Sycophancy in LLMs

    Li-Ni Fu, Chang-Chih Meng, Chien-Hua Chen +2

    cs.AIcs.CLcs.CYarXiv:2608.29198v12026
  27. Attention Sensitivity Is Not Enough: Dissociating Attention-Level and Behavioural In-Context Learning under Fine-Tuning

    Jinyuan Zhang, Peng He, He Hu +2

    cs.LGcs.AIcs.CLarXiv:2609.00064v12026
  28. GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models

    Zuyao Xu, Yuqi Qiu, Lu Sun +14

    cs.CRcs.AIarXiv:2602.06718v22026
  29. BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks

    Xinming Tu, Tianze Wang, Yingzhou +4

    cs.CLcs.AIcs.SEarXiv:2604.24955v12026
  30. Modern Views of Machine Learning for Precision Psychiatry

    Zhe Sage Chen, Prathamesh, Kulkarni +4

    cs.LGcs.AIq-bio.NCarXiv:2204.01607v22022
  31. Review Before Trust: Source-Grounded Integrity Gates for AI-Assisted Personal Health Records

    Nora Girda, Adrian Groza

    cs.AIarXiv:2608.29965v12026
  32. FineCIR: Explicit Parsing of Fine-Grained Modification Semantics for Composed Image Retrieval

    Zixu Li, Zhiheng Fu, Yupeng Hu +3

    cs.CVcs.AIarXiv:2503.21309v12025
  33. Quantifying Frontier LLM Capabilities for Container Sandbox Escape

    Rahul Marchand, Art O Cathain, Jerome Wynne +8

    cs.CRcs.AIarXiv:2603.02277v32026
  34. PerturbDiff: Functional Diffusion for Single-Cell Perturbation Modeling

    Xinyu Yuan, Xixian Liu, Ya Shi Zhang +3

    cs.LGcs.AIarXiv:2602.19685v12026
  35. GAM: Hierarchical Graph-based Agentic Memory for LLM Agents

    Zhaofen Wu, Hanrong Zhang, Fulin Lin +9

    cs.AIarXiv:2604.12285v12026
  36. scDFM: Distributional Flow Matching Model for Robust Single-Cell Perturbation Prediction

    Chenglei Yu, Chuanrui Wang, Bangyan Liao +1

    q-bio.QMcs.AIarXiv:2602.07103v12026
  37. How Your Credentials Are Leaked by LLM Agent Skills: An Empirical Study

    Zhihao Chen, Ying Zhang, Yi Liu +7

    cs.CRcs.AIarXiv:2604.03070v22026
  38. PersistBench: When Should Long-Term Memories Be Forgotten by LLMs?

    Sidharth Pulipaka, Oliver Chen, Manas Sharma +3

    cs.AIarXiv:2602.01146v22026
  39. SynCrash: A Multi-Stage Pipeline for Zero-Shot Accident Detection and Localization in Traffic Surveillance Video

    Arkya Jyoti Bagchi, Ritul Jangir, Varun Raskar

    cs.CVcs.AIarXiv:2608.29759v12026
  40. PromptCharm: Text-to-Image Generation through Multi-modal Prompting and Refinement

    Zhijie Wang, Yuheng Huang, Da Song +2

    cs.HCcs.AIarXiv:2403.04014v12024
  41. Sustainable Hybrid Document-Routed Retrieval for Financial RAG: Resolving the Robustness-Precision Trade-off

    Zhiyuan Cheng, Longying Lai, Yue Liu

    cs.CLcs.AIcs.IRarXiv:2603.26815v32026
  42. Medically Aware GPT-3 as a Data Generator for Medical Dialogue Summarization

    Bharath Chintagunta, Namit Katariya, Xavier Amatriain +1

    cs.CLcs.AIcs.LGarXiv:2110.07356v12021
  43. A Systematic Security Evaluation of OpenClaw and Its Variants

    Yuhang Wang, Haichang Gao, Zhenxing Niu +4

    cs.CRcs.AIarXiv:2604.03131v12026
  44. Towards Secure Agent Skills: Architecture, Threat Taxonomy, and Security Analysis

    Zhiyuan Li, Jingzheng Wu, Xiang Ling +2

    cs.CRcs.AIarXiv:2604.02837v12026
  45. Cost-Effective Repository Exploration for Agentic Issue Localization

    Mohammad Nour Al Awad, Sergey Ivanov

    cs.SEcs.AIarXiv:2608.29675v12026
  46. From Spark to Fire: Modeling and Mitigating Error Cascades in LLM-Based Multi-Agent Collaboration

    Yizhe Xie, Congcong Zhu, Xinyue Zhang +5

    cs.MAcs.AIarXiv:2603.04474v22026
  47. Resilient Routing: Risk-Aware Dynamic Routing in Smart Logistics via Spatiotemporal Graph Learning

    Zhiming Xue, Sichen Zhao, Yalun Qi +2

    cs.AIarXiv:2601.13632v22026
  48. Forget or Fine-tune? A Comparative Study of Machine Unlearning Strategies for Noisy Label Correction

    João L. P. Santana, Filipe R. Cordeiro

    cs.LGcs.AIarXiv:2608.30046v12026
  49. IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs

    Chuan Guo, Juan Felipe Ceron Uribe, Sicheng Zhu +10

    cs.AIcs.CLcs.CRarXiv:2603.10521v12026
  50. HorizonBench: Long-Horizon Personalization with Evolving Preferences

    Shuyue Stella Li, Bhargavi Paranjape, Kerem Oktar +9

    cs.CLcs.AIarXiv:2604.17283v12026
  51. Change in Abstract Argumentation Frameworks: Adding an Argument

    Claudette Cayrol, Florence Dupin de Saint-Cyr, Marie-Christine Lagasquie-Schiex

    cs.AIarXiv:1401.3838v12014
  52. Understanding Agent Scaling in LLM-Based Multi-Agent Systems via Diversity

    Yingxuan Yang, Chengrui Qu, Muning Wen +5

    cs.AIcs.LGarXiv:2602.03794v12026
  53. Measuring Progress Toward AGI: A Cognitive Framework

    Ryan Burnell, Yumeya Yamamori, Orhan Firat +10

    cs.AIarXiv:2605.28405v12026
  54. Theory of Mind for Multi-Agent Collaboration via Large Language Models

    Huao Li, Yu Quan Chong, Simon Stepputtis +4

    cs.CLcs.AIarXiv:2310.10701v32023
  55. Delegating Before Learning: Where Generative AI Sits in Students' Professional Communication

    Jared Ren, Soobin Cho

    cs.HCcs.AIarXiv:2608.28837v12026
  56. Defending Wearable VLMs Against Private Attribute Inference

    Zhimin Li, Pan Wang, Jingxian Chen +4

    cs.CVcs.AIarXiv:2608.28691v12026
  57. MemMachine: A Ground-Truth-Preserving Memory System for Personalized AI Agents

    Shu Wang, Edwin Yu, Oscar Love +4

    cs.AIarXiv:2604.04853v12026
  58. PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing

    Yiwen Song, Yale Song, Tomas Pfister +1

    cs.AIcs.LGcs.MAarXiv:2604.05018v12026
  59. Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents

    Wenkai Yang, Xiaohan Bi, Yankai Lin +3

    cs.CRcs.AIcs.CLarXiv:2402.11208v22024
  60. OAT: Ordered Action Tokenization

    Chaoqi Liu, Xiaoshen Han, Jiawei Gao +3

    cs.ROcs.AIcs.LGarXiv:2602.04215v22026