Artificial Intelligence

Papers filed under cs.AI on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

4,981 to 5,040 of 15,271

  1. Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks

    William F. Shen, Xinchi Qiu, Chenxi Whitehouse +6

    cs.LGcs.AIarXiv:2602.05125v12026
  2. An Agentic System for Rare Disease Diagnosis with Traceable Reasoning

    Weike Zhao, Chaoyi Wu, Yanjie Fan +10

    cs.CLcs.AIcs.CVarXiv:2506.20430v32025
  3. Generative Retrieval for E-commerce: Jointly Learning Embedding and Codebook with Same Product Cluster

    Songtao Fang, Zihao Xu, Shaowei Wei +2

    cs.IRcs.AIarXiv:2608.30606v12026
  4. Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems

    Christian Walder, Deep Karkhanis

    cs.LGcs.AIcs.CLarXiv:2505.15201v52025
  5. LightThinker: Thinking Step-by-Step Compression

    Jintian Zhang, Yuqi Zhu, Mengshu Sun +6

    cs.CLcs.AIcs.IRarXiv:2502.15589v22025
  6. Think Before Recommend: Unleashing the Latent Reasoning Power for Sequential Recommendation

    Jiakai Tang, Sunhao Dai, Teng Shi +5

    cs.IRcs.AIcs.CLarXiv:2503.22675v32025
  7. M-LLM Based Video Frame Selection for Efficient Video Understanding

    Kai Hu, Feng Gao, Xiaohan Nie +8

    cs.CVcs.AIarXiv:2502.19680v22025
  8. Learning to Prove Theorems via Interacting with Proof Assistants

    Kaiyu Yang, Jia Deng

    cs.LOcs.AIcs.LGarXiv:1905.09381v12019
  9. On The Reasons Behind Decisions

    Adnan Darwiche, Auguste Hirth

    cs.AIcs.LGarXiv:2002.09284v12020
  10. WiseSpec: Requirements-Driven Agents for Code Generation

    Zhao Tian

    cs.SEcs.AIarXiv:2609.00568v12026
  11. Neural 3D Morphable Models: Spiral Convolutional Networks for 3D Shape Representation Learning and Generation

    Giorgos Bouritsas, Sergiy Bokhnyak, Stylianos Ploumpis +2

    cs.CVcs.AIcs.GRarXiv:1905.02876v32019
  12. Polished but Unresolved: Identifying Late-Stage Pressure States in Long-Horizon Tool-Use Agents

    Haoyang Chen, Yi Liu, Jianzhi Shao +3

    cs.AIcs.CLarXiv:2609.00823v12026
  13. Hilbert: Recursively Building Formal Proofs with Informal Reasoning

    Sumanth Varambally, Thomas Voice, Yanchao Sun +3

    cs.AIcs.FLcs.LGarXiv:2509.22819v22025
  14. From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding

    Raul Ortega, José Manuel Gómez-Pérez

    cs.CVcs.AIcs.CLarXiv:2609.00948v12026
  15. RVT-2: Learning Precise Manipulation from Few Demonstrations

    Ankit Goyal, Valts Blukis, Jie Xu +3

    cs.ROcs.AIcs.CVarXiv:2406.08545v12024
  16. StreamingVLM: Real-Time Understanding for Infinite Video Streams

    Ruyi Xu, Guangxuan Xiao, Yukang Chen +3

    cs.CVcs.AIcs.CLarXiv:2510.09608v22025
  17. Automated Conjecture Resolution with Formal Verification

    Haocheng Ju, Guoxiong Gao, Jiedong Jiang +13

    cs.LGcs.AIarXiv:2604.03789v22026
  18. FinLifeBench: Exhaustive Life-Event History and Financial-State Reconstruction from Longitudinal Banking Dialogue

    Hangyeul Lee, Juyoung Oh, Jaeyong Ko +5

    cs.AIcs.CLarXiv:2609.01198v22026
  19. CHIP: CHannel Independence-based Pruning for Compact Neural Networks

    Yang Sui, Miao Yin, Yi Xie +3

    cs.CVcs.AIcs.LGarXiv:2110.13981v32021
  20. Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective

    Zhoujun Cheng, Shibo Hao, Tianyang Liu +21

    cs.LGcs.AIcs.CLarXiv:2506.14965v12025
  21. Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding

    Seil Kang, Jinyeong Kim, Junhyeok Kim +1

    cs.CVcs.AIarXiv:2503.06287v12025
  22. AgentFold: Long-Horizon Web Agents with Proactive Context Management

    Rui Ye, Zhongwang Zhang, Kuan Li +12

    cs.CLcs.AIcs.LGarXiv:2510.24699v12025
  23. LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights

    Ze Sheng, Zhicheng Chen, Shuning Gu +3

    cs.CRcs.AIarXiv:2502.07049v22025
  24. (V)LMs generalize beyond surface co-occurrence: Evidence from cross-modal number agreement

    Zach Studdiford, Kanishka Misra

    cs.CLcs.AIarXiv:2609.00443v12026
  25. MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

    Dongzhi Jiang, Renrui Zhang, Ziyu Guo +11

    cs.CVcs.AIcs.CLarXiv:2502.09621v12025
  26. Large Language Models Often Know When They Are Being Evaluated

    Joe Needham, Giles Edkins, Govind Pimpale +2

    cs.CLcs.AIarXiv:2505.23836v32025
  27. Lagged Coupling: Internal Representations Become Readable Before They Become Causal

    Xining Xun

    cs.CLcs.AIarXiv:2609.01048v12026
  28. In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents

    Zhen Tan, Jun Yan, I-Hung Hsu +12

    cs.CLcs.AIarXiv:2503.08026v22025
  29. Does Fault Localization Beat a Fresh Attempt? A Placebo-Controlled Study of Test-Guided Code Repair

    Anik Jha

    cs.SEcs.AIcs.LGarXiv:2609.00854v12026
  30. REVISE: Validity-Guided Recovery for Online Revisions in Agent Workflows

    Ruoling Qi, Xuaner Wu, Penghang Liu +2

    cs.AIarXiv:2609.00643v12026
  31. Spawn Freely, Act Sparingly: Progressive Risk Vesting for Recursive LLM-Agent Trees

    Molly Wang

    cs.AIcs.LGmath.PRarXiv:2609.01035v12026
  32. Unveiling Privacy Risks in LLM Agent Memory

    Bo Wang, Weiyi He, Shenglai Zeng +4

    cs.CRcs.AIarXiv:2502.13172v22025
  33. Automated data processing and feature engineering for deep learning and big data applications: a survey

    Alhassan Mumuni, Fuseini Mumuni

    cs.LGcs.AIcs.DBarXiv:2403.11395v22024
  34. The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

    Nitay Calderon, Roi Reichart, Rotem Dror

    cs.CLcs.AIcs.HCarXiv:2501.10970v42025
  35. Learning-Assisted Congestion-Aware Route Scheduling for Semiconductor Fab Material Control Systems

    Hao Yin, Meiqi Tu, Anbang Liu +2

    cs.AImath.OCarXiv:2608.30520v12026
  36. Competitive Programming with Large Reasoning Models

    OpenAI, :, Ahmed El-Kishky +23

    cs.LGcs.AIcs.CLarXiv:2502.06807v22025
  37. Blended RAG: Improving RAG (Retriever-Augmented Generation) Accuracy with Semantic Search and Hybrid Query-Based Retrievers

    Kunal Sawarkar, Abhilasha Mangal, Shivam Raj Solanki

    cs.IRcs.AIcs.CLarXiv:2404.07220v22024
  38. Scaling RL to Long Videos

    Yukang Chen, Wei Huang, Baifeng Shi +11

    cs.CVcs.AIcs.CLarXiv:2507.07966v42025
  39. Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models

    Chuer Chen, Zichen Wang, Yi He +2

    cs.CVcs.AIarXiv:2609.02502v12026
  40. MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

    Qian Chen, Yafeng Chen, Yanni Chen +33

    cs.CLcs.AIcs.HCarXiv:2501.06282v12025
  41. 20 years of network community detection

    Santo Fortunato, M. E. J. Newman

    physics.soc-phcs.AIcs.SIarXiv:2208.00111v22022
  42. Accelerating Diffusion LLMs via Adaptive Parallel Decoding

    Daniel Israel, Guy Van den Broeck, Aditya Grover

    cs.CLcs.AIcs.LGarXiv:2506.00413v22025
  43. CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

    Yuxuan Zhu, Antony Kellermann, Dylan Bowman +13

    cs.CRcs.AIarXiv:2503.17332v42025
  44. Preference Leakage: A Contamination Problem in LLM-as-a-judge

    Dawei Li, Renliang Sun, Yue Huang +6

    cs.LGcs.AIcs.CLarXiv:2502.01534v32025
  45. On the Prospects of Dynamic LLM Conversations in Software Development

    Annemarie Wittig, Alina Mailach, Janet Siegmund +1

    cs.SEcs.AIarXiv:2608.30756v12026
  46. Inductive Moment Matching

    Linqi Zhou, Stefano Ermon, Jiaming Song

    cs.LGcs.AIstat.MLarXiv:2503.07565v72025
  47. GR-3 Technical Report

    Chilam Cheang, Sijin Chen, Zhongren Cui +18

    cs.ROcs.AIcs.CVarXiv:2507.15493v22025
  48. Capabilities of GPT-5 on Multimodal Medical Reasoning

    Shansong Wang, Mingzhe Hu, Qiang Li +2

    cs.CLcs.AIarXiv:2508.08224v22025
  49. NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model

    NVIDIA, :, Aarti Basant +214

    cs.CLcs.AIcs.LGarXiv:2508.14444v42025
  50. PyG 2.0: Scalable Learning on Real World Graphs

    Matthias Fey, Jinu Sunil, Akihiro Nitta +10

    cs.LGcs.AIarXiv:2507.16991v22025
  51. DriveSuprim: Towards Precise Trajectory Selection for End-to-End Planning

    Wenhao Yao, Zhenxin Li, Shiyi Lan +4

    cs.ROcs.AIcs.CVarXiv:2506.06659v32025
  52. I Know What You Trained Last Summer: A Survey on Stealing Machine Learning Models and Defences

    Daryna Oliynyk, Rudolf Mayer, Andreas Rauber

    cs.LGcs.AIcs.CRarXiv:2206.08451v22022
  53. DLIME: A Deterministic Local Interpretable Model-Agnostic Explanations Approach for Computer-Aided Diagnosis Systems

    Muhammad Rehman Zafar, Naimul Mefraz Khan

    cs.LGcs.AIstat.MLarXiv:1906.10263v12019
  54. Challenges and Countermeasures for Adversarial Attacks on Deep Reinforcement Learning

    Inaam Ilahi, Muhammad Usama, Junaid Qadir +4

    cs.LGcs.AIcs.CRarXiv:2001.09684v22020
  55. Few-Shot Object Detection with Fully Cross-Transformer

    Guangxing Han, Jiawei Ma, Shiyuan Huang +2

    cs.CVcs.AIcs.MMarXiv:2203.15021v22022
  56. Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer Memory

    Yuqi Wu, Wenzhao Zheng, Jie Zhou +1

    cs.CVcs.AIcs.LGarXiv:2507.02863v22025
  57. Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?

    Haoang Chi, He Li, Wenjing Yang +5

    cs.AIcs.CLcs.LGarXiv:2506.21215v12025
  58. Latent Recurrent Thoughts: Recurrent Refinement of Proposed Latents for Reasoning with Frozen LLMs

    Zhaoliang Chen, Jie Fu

    cs.AIcs.CLarXiv:2609.01117v12026
  59. Designing an Auditable LLM-Supported Workflow for Qualitative Thematic Analysis

    Nadia Jul Jeldtoft, Tariq Yousef

    cs.AIcs.SEarXiv:2608.30543v12026
  60. Contrastive Explanation: A Structural-Model Approach

    Tim Miller

    cs.AIarXiv:1811.03163v22018