Artificial Intelligence

Papers filed under cs.AI on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

12,181 to 12,240 of 15,266

  1. Multi-Agent Pathfinding: Definitions, Variants, and Benchmarks

    Roni Stern, Nathan Sturtevant, Ariel Felner +9

    cs.AIcs.MAcs.ROarXiv:1906.08291v12019
  2. FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol

    Jie Zhu, Yimin Tian, Boyang Li +8

    cs.AIcs.CLarXiv:2603.24943v12026
  3. Machine Learning Testing: Survey, Landscapes and Horizons

    Jie M. Zhang, Mark Harman, Lei Ma +1

    cs.LGcs.AIcs.SEarXiv:1906.10742v22019
  4. Overcoming Exploration in Reinforcement Learning with Demonstrations

    Ashvin Nair, Bob McGrew, Marcin Andrychowicz +2

    cs.LGcs.AIcs.NEarXiv:1709.10089v22017
  5. TIES-Merging: Resolving Interference When Merging Models

    Prateek Yadav, Derek Tam, Leshem Choshen +2

    cs.LGcs.AIcs.CLarXiv:2306.01708v22023
  6. CCTU: A Benchmark for Tool Use under Complex Constraints

    Junjie Ye, Guoqiang Zhang, Wenjie Fu +3

    cs.CLcs.AIarXiv:2603.15309v12026
  7. Demystifying Video Reasoning

    Ruisi Wang, Zhongang Cai, Fanyi Pu +11

    cs.CVcs.AIarXiv:2603.16870v32026
  8. EvolVE: Evolutionary Search for LLM-based Verilog Generation and Optimization

    Wei-Po Hsin, Ren-Hao Deng, Yao-Ting Hsieh +2

    cs.AIcs.NEcs.PLarXiv:2601.18067v12026
  9. LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation

    Jiazheng Xing, Fei Du, Hangjie Yuan +7

    cs.CVcs.AIarXiv:2603.20192v12026
  10. DiffusionCLIP: Text-Guided Diffusion Models for Robust Image Manipulation

    Gwanghyun Kim, Taesung Kwon, Jong Chul Ye

    cs.CVcs.AIcs.LGarXiv:2110.02711v62021
  11. Agentic AI and the next intelligence explosion

    James Evans, Benjamin Bratton, Blaise Agüera y Arcas

    cs.AIarXiv:2603.20639v12026
  12. Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

    Melanie Sclar, Yejin Choi, Yulia Tsvetkov +1

    cs.CLcs.AIcs.LGarXiv:2310.11324v22023
  13. RAISE: Requirement-Adaptive Evolutionary Refinement for Training-Free Text-to-Image Alignment

    Liyao Jiang, Ruichen Chen, Chao Gao +1

    cs.CVcs.AIarXiv:2603.00483v12026
  14. Manifold-Aware Exploration for Reinforcement Learning in Video Generation

    Mingzhe Zheng, Weijie Kong, Yue Wu +9

    cs.CVcs.AIarXiv:2603.21872v12026
  15. Reward-free Alignment for Conflicting Objectives

    Peter Chen, Xiaopeng Li, Xi Chen +1

    cs.CLcs.AIcs.LGarXiv:2602.02495v32026
  16. A Benchmark for Interpretability Methods in Deep Neural Networks

    Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans +1

    cs.LGcs.AIstat.MLarXiv:1806.10758v32018
  17. Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image Analysis

    Yucheng Tang, Dong Yang, Wenqi Li +5

    cs.CVcs.AIcs.LGarXiv:2111.14791v22021
  18. PEARL: Personalized Streaming Video Understanding Model

    Yuanhong Zheng, Ruichuan An, Xiaopeng Lin +10

    cs.CVcs.AIcs.IRarXiv:2603.20422v12026
  19. Parseval Networks: Improving Robustness to Adversarial Examples

    Moustapha Cisse, Piotr Bojanowski, Edouard Grave +2

    stat.MLcs.AIcs.CRarXiv:1704.08847v22017
  20. Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning

    Chengzu Li, Zanyi Wang, Jiaang Li +9

    cs.LGcs.AIcs.CLarXiv:2601.21037v12026
  21. KAPSO: A Knowledge-grounded framework for Autonomous Program Synthesis and Optimization

    Alireza Nadafian, Alireza Mohammadshahi, Majid Yazdani

    cs.AIcs.CLcs.SEarXiv:2601.21526v22026
  22. Deep Reconstruction-Classification Networks for Unsupervised Domain Adaptation

    Muhammad Ghifary, W. Bastiaan Kleijn, Mengjie Zhang +2

    cs.CVcs.AIcs.LGarXiv:1607.03516v22016
  23. Residual Context Diffusion Language Models

    Yuezhou Hu, Harman Singh, Monishwaran Maheswaran +10

    cs.CLcs.AIarXiv:2601.22954v22026
  24. HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing

    Yizhao Gao, Jianyu Wei, Qihao Zhang +11

    cs.CLcs.AIarXiv:2602.03560v12026
  25. ReWorld: An Interactive World Model with Long-Horizon Memory

    Zhifei Chen, Luozhou Wang, Guibao Shen +8

    cs.AIarXiv:2608.23565v12026
  26. Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing

    Jiyuan Wang, Chunyu Lin, Lei Sun +8

    cs.CVcs.AIarXiv:2603.03143v22026
  27. Chain of World: World Model Thinking in Latent Motion

    Fuxiang Yang, Donglin Di, Lulu Tang +6

    cs.CVcs.AIcs.ROarXiv:2603.03195v12026
  28. AgentArk: Distilling Multi-Agent Intelligence into a Single LLM Agent

    Yinyi Luo, Yiqiao Jin, Weichen Yu +6

    cs.AIcs.MAarXiv:2602.03955v32026
  29. Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress

    Chen Yang, Haiyuan Wan, Rengrong Xiong +2

    cs.AIcs.LGarXiv:2608.19408v12026
  30. Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units

    Jianhui Chen, Yuzhang Luo, Liangming Pan

    cs.CLcs.AIcs.LGarXiv:2601.21996v22026
  31. Drug discovery with explainable artificial intelligence

    José Jiménez-Luna, Francesca Grisoni, Gisbert Schneider

    cs.AIcs.LGstat.MLarXiv:2007.00523v22020
  32. CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty

    Johannes Kirmayr, Lukas Stappen, Elisabeth André

    cs.AIarXiv:2601.22027v12026
  33. Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model

    Dongwon Kim, Gawon Seo, Jinsung Lee +2

    cs.CVcs.AIcs.ROarXiv:2603.05438v12026
  34. TAPAS: Weakly Supervised Table Parsing via Pre-training

    Jonathan Herzig, Paweł Krzysztof Nowak, Thomas Müller +2

    cs.IRcs.AIcs.CLarXiv:2004.02349v22020
  35. A Neural Network Approach to Context-Sensitive Generation of Conversational Responses

    Alessandro Sordoni, Michel Galley, Michael Auli +6

    cs.CLcs.AIcs.LGarXiv:1506.06714v12015
  36. Lost in the Noise: How Reasoning Models Fail with Contextual Distractors

    Seongyun Lee, Yongrae Jo, Minju Seo +2

    cs.AIcs.CLarXiv:2601.07226v12026
  37. Action-Conditional Video Prediction using Deep Networks in Atari Games

    Junhyuk Oh, Xiaoxiao Guo, Honglak Lee +2

    cs.LGcs.AIcs.CVarXiv:1507.08750v22015
  38. OpenAssistant Conversations -- Democratizing Large Language Model Alignment

    Andreas Köpf, Yannic Kilcher, Dimitri von Rütte +15

    cs.CLcs.AIarXiv:2304.07327v22023
  39. Cross-Task Generalization via Natural Language Crowdsourcing Instructions

    Swaroop Mishra, Daniel Khashabi, Chitta Baral +1

    cs.CLcs.AIcs.CVarXiv:2104.08773v42021
  40. MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens

    Yu Chen, Runkai Chen, Sheng Yi +9

    cs.CLcs.AIcs.IRarXiv:2603.23516v22026
  41. F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare

    Daniil Plyusov, Alexey Gorbatovski, Boris Shaposhnikov +4

    cs.LGcs.AIarXiv:2602.06717v22026
  42. Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

    Bradley Brown, Jordan Juravsky, Ryan Ehrlich +4

    cs.LGcs.AIarXiv:2407.21787v32024
  43. A Theoretical Analysis of Contrastive Unsupervised Representation Learning

    Sanjeev Arora, Hrishikesh Khandeparkar, Mikhail Khodak +2

    cs.LGcs.AIstat.MLarXiv:1902.09229v12019
  44. MemoBrain: Executive Memory as an Agentic Brain for Reasoning

    Hongjin Qian, Zhao Cao, Zheng Liu

    cs.AIcs.CLcs.IRarXiv:2601.08079v12026
  45. Prime Agent: A Self-Improving RLM Harness

    Seth Karten, Alex L. Zhang, Kevin Thomas +8

    cs.AIcs.CLcs.SEarXiv:2608.23552v12026
  46. ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks

    Samin Mahdizadeh Sani, Max Ku, Nima Jamali +23

    cs.GRcs.AIcs.CVarXiv:2603.27862v12026
  47. Transfer Learning in Deep Reinforcement Learning: A Survey

    Zhuangdi Zhu, Kaixiang Lin, Anil K. Jain +1

    cs.LGcs.AIstat.MLarXiv:2009.07888v72020
  48. FlowAct-R1: Towards Interactive Humanoid Video Generation

    Lizhen Wang, Yongming Zhu, Zhipeng Ge +15

    cs.CVcs.AIarXiv:2601.10103v12026
  49. G-LNS: Generative Large Neighborhood Search for LLM-Based Automatic Heuristic Design

    Baoyun Zhao, He Wang, Liang Zeng

    cs.AIarXiv:2602.08253v12026
  50. Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications

    Tzafrir Rehan

    cs.SEcs.AIarXiv:2603.08806v12026
  51. Toward Efficient Agents: Memory, Tool learning, and Planning

    Xiaofang Yang, Lijun Li, Heng Zhou +12

    cs.AIcs.CLarXiv:2601.14192v22026
  52. Decoupled Knowledge Distillation

    Borui Zhao, Quan Cui, Renjie Song +2

    cs.CVcs.AIarXiv:2203.08679v22022
  53. TabTransformer: Tabular Data Modeling Using Contextual Embeddings

    Xin Huang, Ashish Khetan, Milan Cvitkovic +1

    cs.LGcs.AIarXiv:2012.06678v12020
  54. Self-Improving World Modelling with Latent Actions

    Yifu Qiu, Zheng Zhao, Waylon Li +4

    cs.LGcs.AIcs.CLarXiv:2602.06130v22026
  55. Survey of the State of the Art in Natural Language Generation: Core tasks, applications and evaluation

    Albert Gatt, Emiel Krahmer

    cs.CLcs.AIcs.NEarXiv:1703.09902v42017
  56. LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations

    William Lugoloobi, Thomas Foster, William Bankes +1

    cs.CLcs.AIcs.LGarXiv:2602.09924v42026
  57. WideSeek-R1: Exploring Width Scaling for Broad Information Seeking via Multi-Agent Reinforcement Learning

    Zelai Xu, Zhexuan Xu, Ruize Zhang +7

    cs.AIcs.LGcs.MAarXiv:2602.04634v32026
  58. Crowdsourcing Multiple Choice Science Questions

    Johannes Welbl, Nelson F. Liu, Matt Gardner

    cs.HCcs.AIcs.CLarXiv:1707.06209v12017
  59. Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

    Yann Dubois, Balázs Galambosi, Percy Liang +1

    cs.LGcs.AIcs.CLarXiv:2404.04475v22024
  60. A White Paper on Neural Network Quantization

    Markus Nagel, Marios Fournarakis, Rana Ali Amjad +3

    cs.LGcs.AIcs.CVarXiv:2106.08295v12021