Artificial Intelligence
Papers filed under cs.AI on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
3,541 to 3,600 of 15,439
Wink: Recovering from Misbehaviors in Coding Agents
Rahul Nanda, Chandra Maddila, Smriti Jha +3
cs.SEcs.AIcs.HCarXiv:2602.17037v22026Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect Judges
Chen Feng, Minghe Shen, Ananth Balashankar +2
cs.LGcs.AIcs.CVarXiv:2601.20913v12026Label Efficient Semi-Supervised Learning via Graph Filtering
Qimai Li, Xiao-Ming Wu, Han Liu +2
cs.LGcs.AIstat.MLarXiv:1901.09993v32019A Systematic Review for Transformer-based Long-term Series Forecasting
Liyilei Su, Xumin Zuo, Rui Li +3
cs.LGcs.AIarXiv:2310.20218v12023Trajectory-Informed Memory Generation for Self-Improving Agent Systems
Gaodan Fang, Vatche Isahagian, K. R. Jayaram +4
cs.AIcs.DBcs.IRarXiv:2603.10600v12026Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents
Zehong Wang, Fang Wu, Hongru Wang +8
cs.AIcs.CLcs.LGarXiv:2601.22311v12026GUI-GENESIS: Automated Synthesis of Efficient Environments with Verifiable Rewards for GUI Agent Post-Training
Yuan Cao, Dezhi Ran, Mengzhou Wu +9
cs.AIcs.LGarXiv:2602.14093v12026M2A: Multimodal Memory Agent with Dual-Layer Hybrid Memory for Long-Term Personalized Interactions
Junyu Feng, Binxiao Xu, Jiayi Chen +8
cs.AIarXiv:2602.07624v12026AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition
Ruipeng Wang, Yuxin Chen, Yukai Wang +9
cs.AIarXiv:2602.11348v22026Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning
Yihong Huang, Fei Ma, Yihua Shao +4
cs.CVcs.AIcs.CLarXiv:2602.02951v12026Free Speech and Artificial Intelligence
Etienne Brown
cs.CYcs.AIarXiv:2608.28973v12026Not the Same Protector: Deployment-Dependent Protective Intervention in LLMs
Eunna Lee, Soomyoung Lee, Jungpyo Nam +6
cs.CRcs.AIarXiv:2608.29136v12026Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks
Ang Li, Yin Zhou, Vethavikashini Chithrra Raghuram +2
cs.LGcs.AIarXiv:2502.08586v12025Limits of trust in medical AI
Joshua Hatherley
cs.LGcs.AIcs.CYarXiv:2503.16692v22025Development of an Autonomous AI Coding Agent using Monte Carlo Tree Search (MCTS) and Gemini LLM Frameworks
Pravin Game, Vipin Ramakrishnan, Prathamesh Wagh
cs.LGcs.AIarXiv:2608.29096v12026An effective algorithm for hyperparameter optimization of neural networks
Gonzalo Diaz, Achille Fokoue, Giacomo Nannicini +1
cs.AIcs.LGcs.NEarXiv:1705.08520v12017Automatic Conversion of NICE Guidelines to an Executable Computational Model Using Large Language Models
Ashvin Gupta, Denys Prociuk, Alessandra Russo +1
cs.AIarXiv:2608.30022v12026MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks
Zixuan Ke, Yifei Ming, Austin Xu +7
cs.AIcs.CLcs.MAarXiv:2601.14652v520266G-Bench: An Open Benchmark for Semantic Communication and Network-Level Reasoning with Foundation Models in AI-Native 6G Networks
Mohamed Amine Ferrag, Abderrahmane Lakas, Merouane Debbah
cs.NIcs.AIarXiv:2602.08675v12026A Generalized Optimization Engine (GOE) for Edge AI Inference Acceleration
Venkat R. Dasari, Jakob A. Adams, Vinod K. Mishra +1
cs.AIarXiv:2608.28652v12026Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts
Saman Rahbar, Xiliang Zhu, Irvin Cardoza +1
cs.CLcs.AIcs.LGarXiv:2609.00330v12026Large-Scale Optimization Model Auto-Formulation: Harnessing LLM Flexibility via Structured Workflow
Kuo Liang, Yuhang Lu, Jianming Mao +7
cs.AIcs.LGarXiv:2601.09635v32026From Flat Logs to Causal Graphs: Hierarchical Failure Attribution for LLM-based Multi-Agent Systems
Yawen Wang, Wenjie Wu, Junjie Wang +1
cs.AIcs.SEarXiv:2602.23701v12026MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning
Xiaoliang Fu, Jiaye Lin, Yangyi Fang +7
cs.LGcs.AIarXiv:2602.17550v32026Interpretable All-Type Audio Deepfake Detection with Audio LLMs via Frequency-Time Reinforcement Learning
Yuankun Xie, Xiaoxuan Guo, Jiayi Zhou +6
cs.SDcs.AIarXiv:2601.02983v12026How Identity and Opinion Shape Political Sycophancy in LLMs
Li-Ni Fu, Chang-Chih Meng, Chien-Hua Chen +2
cs.AIcs.CLcs.CYarXiv:2608.29198v12026Attention Sensitivity Is Not Enough: Dissociating Attention-Level and Behavioural In-Context Learning under Fine-Tuning
Jinyuan Zhang, Peng He, He Hu +2
cs.LGcs.AIcs.CLarXiv:2609.00064v12026GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models
Zuyao Xu, Yuqi Qiu, Lu Sun +14
cs.CRcs.AIarXiv:2602.06718v22026BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
Xinming Tu, Tianze Wang, Yingzhou +4
cs.CLcs.AIcs.SEarXiv:2604.24955v12026Modern Views of Machine Learning for Precision Psychiatry
Zhe Sage Chen, Prathamesh, Kulkarni +4
cs.LGcs.AIq-bio.NCarXiv:2204.01607v22022Review Before Trust: Source-Grounded Integrity Gates for AI-Assisted Personal Health Records
Nora Girda, Adrian Groza
cs.AIarXiv:2608.29965v12026FineCIR: Explicit Parsing of Fine-Grained Modification Semantics for Composed Image Retrieval
Zixu Li, Zhiheng Fu, Yupeng Hu +3
cs.CVcs.AIarXiv:2503.21309v12025Quantifying Frontier LLM Capabilities for Container Sandbox Escape
Rahul Marchand, Art O Cathain, Jerome Wynne +8
cs.CRcs.AIarXiv:2603.02277v32026PerturbDiff: Functional Diffusion for Single-Cell Perturbation Modeling
Xinyu Yuan, Xixian Liu, Ya Shi Zhang +3
cs.LGcs.AIarXiv:2602.19685v12026GAM: Hierarchical Graph-based Agentic Memory for LLM Agents
Zhaofen Wu, Hanrong Zhang, Fulin Lin +9
cs.AIarXiv:2604.12285v12026scDFM: Distributional Flow Matching Model for Robust Single-Cell Perturbation Prediction
Chenglei Yu, Chuanrui Wang, Bangyan Liao +1
q-bio.QMcs.AIarXiv:2602.07103v12026How Your Credentials Are Leaked by LLM Agent Skills: An Empirical Study
Zhihao Chen, Ying Zhang, Yi Liu +7
cs.CRcs.AIarXiv:2604.03070v22026PersistBench: When Should Long-Term Memories Be Forgotten by LLMs?
Sidharth Pulipaka, Oliver Chen, Manas Sharma +3
cs.AIarXiv:2602.01146v22026SynCrash: A Multi-Stage Pipeline for Zero-Shot Accident Detection and Localization in Traffic Surveillance Video
Arkya Jyoti Bagchi, Ritul Jangir, Varun Raskar
cs.CVcs.AIarXiv:2608.29759v12026PromptCharm: Text-to-Image Generation through Multi-modal Prompting and Refinement
Zhijie Wang, Yuheng Huang, Da Song +2
cs.HCcs.AIarXiv:2403.04014v12024Sustainable Hybrid Document-Routed Retrieval for Financial RAG: Resolving the Robustness-Precision Trade-off
Zhiyuan Cheng, Longying Lai, Yue Liu
cs.CLcs.AIcs.IRarXiv:2603.26815v32026Medically Aware GPT-3 as a Data Generator for Medical Dialogue Summarization
Bharath Chintagunta, Namit Katariya, Xavier Amatriain +1
cs.CLcs.AIcs.LGarXiv:2110.07356v12021A Systematic Security Evaluation of OpenClaw and Its Variants
Yuhang Wang, Haichang Gao, Zhenxing Niu +4
cs.CRcs.AIarXiv:2604.03131v12026Towards Secure Agent Skills: Architecture, Threat Taxonomy, and Security Analysis
Zhiyuan Li, Jingzheng Wu, Xiang Ling +2
cs.CRcs.AIarXiv:2604.02837v12026Cost-Effective Repository Exploration for Agentic Issue Localization
Mohammad Nour Al Awad, Sergey Ivanov
cs.SEcs.AIarXiv:2608.29675v12026From Spark to Fire: Modeling and Mitigating Error Cascades in LLM-Based Multi-Agent Collaboration
Yizhe Xie, Congcong Zhu, Xinyue Zhang +5
cs.MAcs.AIarXiv:2603.04474v22026Resilient Routing: Risk-Aware Dynamic Routing in Smart Logistics via Spatiotemporal Graph Learning
Zhiming Xue, Sichen Zhao, Yalun Qi +2
cs.AIarXiv:2601.13632v22026Forget or Fine-tune? A Comparative Study of Machine Unlearning Strategies for Noisy Label Correction
João L. P. Santana, Filipe R. Cordeiro
cs.LGcs.AIarXiv:2608.30046v12026IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
Chuan Guo, Juan Felipe Ceron Uribe, Sicheng Zhu +10
cs.AIcs.CLcs.CRarXiv:2603.10521v12026HorizonBench: Long-Horizon Personalization with Evolving Preferences
Shuyue Stella Li, Bhargavi Paranjape, Kerem Oktar +9
cs.CLcs.AIarXiv:2604.17283v12026Change in Abstract Argumentation Frameworks: Adding an Argument
Claudette Cayrol, Florence Dupin de Saint-Cyr, Marie-Christine Lagasquie-Schiex
cs.AIarXiv:1401.3838v12014Understanding Agent Scaling in LLM-Based Multi-Agent Systems via Diversity
Yingxuan Yang, Chengrui Qu, Muning Wen +5
cs.AIcs.LGarXiv:2602.03794v12026Measuring Progress Toward AGI: A Cognitive Framework
Ryan Burnell, Yumeya Yamamori, Orhan Firat +10
cs.AIarXiv:2605.28405v12026Theory of Mind for Multi-Agent Collaboration via Large Language Models
Huao Li, Yu Quan Chong, Simon Stepputtis +4
cs.CLcs.AIarXiv:2310.10701v32023Delegating Before Learning: Where Generative AI Sits in Students' Professional Communication
Jared Ren, Soobin Cho
cs.HCcs.AIarXiv:2608.28837v12026Defending Wearable VLMs Against Private Attribute Inference
Zhimin Li, Pan Wang, Jingxian Chen +4
cs.CVcs.AIarXiv:2608.28691v12026MemMachine: A Ground-Truth-Preserving Memory System for Personalized AI Agents
Shu Wang, Edwin Yu, Oscar Love +4
cs.AIarXiv:2604.04853v12026PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing
Yiwen Song, Yale Song, Tomas Pfister +1
cs.AIcs.LGcs.MAarXiv:2604.05018v12026Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
Wenkai Yang, Xiaohan Bi, Yankai Lin +3
cs.CRcs.AIcs.CLarXiv:2402.11208v22024OAT: Ordered Action Tokenization
Chaoqi Liu, Xiaoshen Han, Jiawei Gao +3
cs.ROcs.AIcs.LGarXiv:2602.04215v22026