Artificial Intelligence

Papers filed under cs.AI on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,241 to 3,300 of 15,416

  1. Hybrid Self-evolving Structured Memory for GUI Agents

    Sibo Zhu, Wenyi Wu, Kun Zhou +2

    cs.AIcs.LGarXiv:2603.10291v12026
  2. Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle

    Happy Bhati

    cs.SEcs.AIarXiv:2609.04681v12026
  3. In Line with Context: Repository-Level Code Generation via Context Inlining

    Chao Hu, Wenhao Zeng, Yuling Shi +2

    cs.SEcs.AIarXiv:2601.00376v32026
  4. SCAPES: Semantically Conditioned Autoregressive Prior for Environmental Sounds

    Esteban Gutiérrez, Lonce Wyse, Frederic Font +1

    cs.SDcs.AIcs.LGarXiv:2609.04634v12026
  5. Dynamic Adaptation of the LLM Context for Generating Routines with Coupled Semantics

    Gnaneswar Villuri, Hashmath Shaik, Alex Doboli

    cs.SEcs.AIarXiv:2609.04570v12026
  6. Persistent Teacher Anchoring for Tool-Using Agents

    Hyun Bin Park, Kyungho Song, Sangmin Lee +1

    cs.LGcs.AIcs.CLarXiv:2609.04773v12026
  7. Systematic Evaluation of Single-Cell Foundation Model Interpretability Reveals Attention Captures Co-Expression Rather Than Unique Regulatory Signal

    Ihor Kendiukhov

    q-bio.GNcs.AIarXiv:2602.17532v12026
  8. $α^3$-Bench: A Unified Benchmark of Safety, Robustness, and Efficiency for LLM-Based UAV Agents over 6G Networks

    Mohamed Amine Ferrag, Abderrahmane Lakas, Merouane Debbah

    eess.SYcs.AIarXiv:2601.03281v12026
  9. When Does an Interpretation Count as Established? The Formation, Evaluation, and Responsibility of Interpretation in Generative AI

    Deyu Jing

    cs.CYcs.AIarXiv:2609.04766v12026
  10. MultiDocFusion: Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documents

    Joongmin Shin, Chanjun Park, Jeongbae Park +2

    cs.AIcs.CLarXiv:2604.12352v12026
  11. Convolutional Kolmogorov-Arnold Networks

    Alexander Dylan Bodner, Antonio Santiago Tepsich, Jack Natan Spolski +1

    cs.CVcs.AIarXiv:2406.13155v32024
  12. Counting Belief Propagation

    Kristian Kersting, Babak Ahmadi, Sriraam Natarajan

    cs.AIarXiv:1205.2637v12012
  13. Building a research-software catalog with a coding agent: from hackathon prototype to public deployment

    Kazuyoshi Yoshimi, Satoshi Terasaki, Gotai Yamada

    cs.SEcs.AIcs.CYarXiv:2609.04711v12026
  14. Simulation-free Unbalanced Dynamic Optimal Transport with General Growth Penalty

    Junda Ying, Yuxuan Wang, Bowen Yang +2

    cs.LGcs.AIq-bio.QMarXiv:2609.04710v12026
  15. Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation

    Zhanghao Hu, Qinglin Zhu, Runcong Zhao +4

    cs.CLcs.AIarXiv:2602.02007v42026
  16. Wireless Foundation Models: State-of-the-Art and Open Challenges

    Alonso M. Pacheco Huachaca, Juan J. Rodriguez Rodriguez, Ahmed Aboulfotouh +3

    eess.SPcs.AIarXiv:2609.04707v12026
  17. WAXAL: A Large-Scale Multilingual African Language Speech Corpus

    Abdoulaye Diack, Perry Nelson, Kwaku Agbesi +40

    eess.AScs.AIcs.CLarXiv:2602.02734v32026
  18. Neural-Guided Deductive Search for Real-Time Program Synthesis from Examples

    Ashwin Kalyan, Abhishek Mohta, Oleksandr Polozov +3

    cs.AIcs.LGcs.PLarXiv:1804.01186v22018
  19. The Need for a Socially-Grounded Persona Framework for User Simulation

    Pranav Narayanan Venkit, Yu Li, Yada Pruksachatkun +1

    cs.CLcs.AIcs.CYarXiv:2601.07110v22026
  20. KVzap: Fast, Adaptive, and Faithful KV Cache Pruning

    Simon Jegou, Maximilian Jeblick

    cs.LGcs.AIcs.CLarXiv:2601.07891v22026
  21. Tracing Audio Grounding and Answer Selection in Audio LLMs

    Hyebin Cho, Suho Yoo, Jihoo Jung +1

    cs.CLcs.AIcs.LGarXiv:2609.04637v12026
  22. Subliminal Effects in Your Data: A General Mechanism via Log-Linearity

    Ishaq Aden-Ali, Noah Golowich, Allen Liu +3

    cs.LGcs.AIcs.CLarXiv:2602.04863v12026
  23. D-Former: A U-shaped Dilated Transformer for 3D Medical Image Segmentation

    Yixuan Wu, Kuanlun Liao, Jintai Chen +4

    cs.CVcs.AIarXiv:2201.00462v22022
  24. Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards

    Tengjun Jin, Yoojin Choi, Yuxuan Zhu +1

    cs.AIcs.DBarXiv:2601.08778v32026
  25. PetQA: Benchmarking Veterinary Knowledge and Clinical Reasoning

    Taegyun Kim, Youngwook Ham, Jungwook Rhim +3

    cs.CLcs.AIcs.CVarXiv:2609.04598v12026
  26. Dual-Part Multi-Lateral Branched Network for Multi-Class Segmentation in Cardiovascular Catheterization Angiograms

    Olatunji Omisore, Ahmed Elazab, Ali Shahidinejad +1

    cs.CVcs.AIcs.ROarXiv:2609.04590v12026
  27. When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models

    Gnaneswar Villuri, Hashmath Shaik, Alex Doboli

    cs.CLcs.AIarXiv:2609.04582v12026
  28. DeltaEvolve: Accelerating Scientific Discovery through Momentum-Driven Evolution

    Jiachen Jiang, Tianyu Ding, Zhihui Zhu

    cs.AIcs.LGarXiv:2602.02919v12026
  29. SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass

    Yewei Liu, Xiyuan Wang, Yansheng Mao +3

    cs.CLcs.AIarXiv:2602.06358v32026
  30. DVGT-2: Vision-Geometry-Action Model for Autonomous Driving at Scale

    Sicheng Zuo, Zixun Xie, Wenzhao Zheng +6

    cs.CVcs.AIcs.ROarXiv:2604.00813v32026
  31. Language Model Circuits Are Sparse in the Neuron Basis

    Aryaman Arora, Zhengxuan Wu, Jacob Steinhardt +1

    cs.CLcs.AIarXiv:2601.22594v22026
  32. Lying with Truths: Open-Channel Multi-Agent Collusion for Belief Manipulation via Generative Montage

    Jinwei Hu, Xinmiao Huang, Youcheng Sun +2

    cs.CLcs.AIcs.MAarXiv:2601.01685v22026
  33. Training-Free Halving of Activated Experts in Fine-Grained Mixture-of-Experts Models

    Xing Chen, Hengshuai Yao

    cs.LGcs.AIarXiv:2609.04575v12026
  34. Quantifying Uncertainties in Natural Language Processing Tasks

    Yijun Xiao, William Yang Wang

    cs.CLcs.AIcs.LGarXiv:1811.07253v12018
  35. Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe

    Dain Kim, Eungi Cho, Kyumin Kim +2

    cs.AIcs.CLarXiv:2609.05395v12026
  36. Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence

    Urja Pawar, Rajitha Ramanayake, Nabeel Kemal +4

    cs.AIarXiv:2609.05385v12026
  37. PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice

    Yuzhen Shi, Huanghai Liu, Yiran Hu +27

    cs.CLcs.AIcs.CYarXiv:2601.16669v22026
  38. CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents

    Haoting Shi, Wenhao Wang, Weicheng Fang +6

    cs.AIarXiv:2609.05374v12026
  39. A Deep Generative Model for Synthesizing Labeled Wireless Signals

    Yuxiao Li, Keke Hu, Santiago Mazuelas +1

    cs.AIarXiv:2609.05396v12026
  40. LLM-42: Enabling Determinism in LLM Inference with Verified Speculation

    Raja Gond, Aditya K Kamath, Ramachandran Ramjee +1

    cs.LGcs.AIcs.DCarXiv:2601.17768v22026
  41. Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education

    Rayed AlGhamdi

    cs.AIarXiv:2609.05346v12026
  42. A Minimal Agent for Automated Theorem Proving

    Borja Requena, Austin Letson, Krystian Nowakowski +2

    cs.AIarXiv:2602.24273v32026
  43. Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

    Matthias Busch, Marius Tacke, Sviatlana V. Lamaka +4

    cs.AIarXiv:2609.05381v12026
  44. AI for Computational Design Science: A Responsible Human-AI Framework and Case Study on Short-Form Video Safety Surveillance

    Wenli Zhang, Jiaheng Xie, Zhihe Pan +3

    cs.AIarXiv:2609.05270v12026
  45. GUT: Quantifying and Optimizing the Reasoning Uncertainty of LLMs via Graph Complexity

    Shuang Liang, Xin-Yu Hu, Xiang-Jun Ou +1

    cs.AIarXiv:2609.05284v12026
  46. Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits

    Jinyi Ye, Lei Cao, Ding Chen +1

    physics.soc-phcs.AIcs.CYarXiv:2605.18890v12026
  47. Multi-Domain Collaborative Filtering

    Yu Zhang, Bin Cao, Dit-Yan Yeung

    cs.IRcs.AIarXiv:1203.3535v12012
  48. Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness

    Alexander Neubauer, Tianzhen Hong, Han Li +4

    cs.AIcs.CLeess.SYarXiv:2609.05314v12026
  49. PIArena: A Platform for Prompt Injection Evaluation

    Runpeng Geng, Chenlong Yin, Yanting Wang +2

    cs.CRcs.AIcs.CLarXiv:2604.08499v12026
  50. DGPO: Distribution Guided Policy Optimization for Fine Grained Credit Assignment

    Hongbo Jin, Rongpeng Zhu, Zhongjing Du +4

    cs.LGcs.AIarXiv:2605.03327v22026
  51. RubricRAG: Towards Interpretable and Reliable LLM Evaluation via Domain Knowledge Retrieval for Rubric Generation

    Kaustubh D. Dhole, Eugene Agichtein

    cs.IRcs.AIcs.CLarXiv:2603.20882v12026
  52. FadeMem: Biologically-Inspired Forgetting for Efficient Agent Memory

    Lei Wei, Xiao Peng, Xu Dong +2

    cs.AIcs.CLarXiv:2601.18642v22026
  53. Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents

    Jiazheng Sun, Boyu Yang, Binhao Yuan +2

    cs.AIcs.SEarXiv:2609.05261v12026
  54. MCP-ITP: An Automated Framework for Implicit Tool Poisoning in MCP

    Ruiqi Li, Zhiqiang Wang, Yunhao Yao +1

    cs.CRcs.AIarXiv:2601.07395v12026
  55. Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability

    Ankit Goyal, Jaideep Ray

    cs.AIcs.CLcs.IRarXiv:2609.05339v12026
  56. Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models

    José Luciano Verçosa Marques, Frederico Jorge Heitmann, Daniel Omar Perez +2

    cs.AIcs.CLarXiv:2609.05333v12026
  57. LLM-Driven Algorithm Design for Quantum Circuit Synthesis based on Binary Decision Diagrams

    Yoonju Sim, Federico Berto, Chuanbo Hua +2

    cs.AIcs.ARarXiv:2609.05327v12026
  58. Post-Hoc Reasoning in Chain of Thought: Decoding and Steering Pre-Committed Answers

    Kyle Cox, Darius Kianersi, Adrià Garriga-Alonso

    cs.AIarXiv:2603.01437v22026
  59. Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods

    Maria Mahbub, Ashley Rice, Michael R. Munroe +2

    cs.AIarXiv:2609.05289v12026
  60. HiMem: Hierarchical Long-Term Memory for LLM Long-Horizon Agents

    Ningning Zhang, Xingxing Yang, Zhizhong Tan +2

    cs.AIarXiv:2601.06377v12026