Artificial Intelligence

Papers filed under cs.AI on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

781 to 840 of 15,236

  1. SPARQA: Skeleton-based Semantic Parsing for Complex Questions over Knowledge Bases

    Yawei Sun, Lingling Zhang, Gong Cheng +1

    cs.CLcs.AIarXiv:2003.13956v12020
  2. Multi-Agent Deep Research: Training Multi-Agent Systems with M-GRPO

    Haoyang Hong, Jiajun Yin, Yuan Wang +14

    cs.AIarXiv:2511.13288v22025
  3. BIRD-INTERACT: Re-imagining Text-to-SQL Evaluation for Large Language Models via Lens of Dynamic Interactions

    Nan Huo, Xiaohan Xu, Jinyang Li +21

    cs.AIarXiv:2510.05318v32025
  4. UserRL: Training Interactive User-Centric Agent via Reinforcement Learning

    Cheng Qian, Zuxin Liu, Akshara Prabhakar +10

    cs.AIcs.CLcs.LGarXiv:2509.19736v12025
  5. AI Harness Engineering: A Runtime Substrate for Foundation-Model Software Agents

    Hailin Zhong, Shengxin Zhu

    cs.SEcs.AIarXiv:2605.13357v12026
  6. Single-stream Policy Optimization

    Zhongwen Xu, Zihan Ding

    cs.LGcs.AIstat.MLarXiv:2509.13232v22025
  7. DensePhysNet: Learning Dense Physical Object Representations via Multi-step Dynamic Interactions

    Zhenjia Xu, Jiajun Wu, Andy Zeng +2

    cs.ROcs.AIcs.CVarXiv:1906.03853v22019
  8. A Survey on Efficient Vision-Language-Action Models

    Zhaoshu Yu, Bo Wang, Pengpeng Zeng +7

    cs.CVcs.AIcs.LGarXiv:2510.24795v22025
  9. RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning

    Sicheng Feng, Kaiwen Tuo, Song Wang +3

    cs.CVcs.AIarXiv:2510.02240v22025
  10. Equilibrium Matching: Generative Modeling with Implicit Energy-Based Models

    Runqian Wang, Yilun Du

    cs.LGcs.AIcs.CVarXiv:2510.02300v32025
  11. Adam-mini: Use Fewer Learning Rates To Gain More

    Yushun Zhang, Congliang Chen, Ziniu Li +6

    cs.LGcs.AIarXiv:2406.16793v72024
  12. Improving Human-AI Partnerships in Child Welfare: Understanding Worker Practices, Challenges, and Desires for Algorithmic Decision Support

    Anna Kawakami, Venkatesh Sivaraman, Hao-Fei Cheng +7

    cs.HCcs.AIcs.CYarXiv:2204.02310v12022
  13. NORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference Rewards

    Chia-Yu Hung, Navonil Majumder, Haoyuan Deng +7

    cs.ROcs.AIarXiv:2511.14659v12025
  14. Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward

    Peter Chen, Xiaopeng Li, Ziniu Li +3

    cs.LGcs.AIcs.CLarXiv:2512.16912v32025
  15. WorldGen: From Text to Traversable and Interactive 3D Worlds

    Dilin Wang, Hyunyoung Jung, Tom Monnier +22

    cs.CVcs.AIarXiv:2511.16825v12025
  16. Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks

    Cheng Yang, Haiyuan Wan, Yiran Peng +8

    cs.CVcs.AIarXiv:2511.15065v22025
  17. From Next-Token to Next-Block: A Principled Adaptation Path for Diffusion LLMs

    Yuchuan Tian, Yuchen Liang, Shuo Zhang +10

    cs.CLcs.AIarXiv:2512.06776v22025
  18. Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training

    Luming Yang, Haoxian Liu, Siqing Li +4

    cs.MAcs.AIcs.HCarXiv:2609.10939v12026
  19. Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge

    Yuanchen Bai, Zijian Ding, Angelique Taylor

    cs.AIcs.HCcs.MAarXiv:2609.10724v12026
  20. NVIDIA Nemotron 3: Efficient and Open Intelligence

    NVIDIA, :, Aaron Blakeman +356

    cs.CLcs.AIcs.LGarXiv:2512.20856v12025
  21. Understanding Operator Attitudes Toward AI-Supported Decision Making in Maritime Operations

    Doreen Jirak, Armeen Saroukanoff, Dirk van Rooy

    cs.HCcs.AIarXiv:2609.11805v12026
  22. AI Soccer Analyst: Stage-Aware and Verifiable Human-AI Collaboration for Soccer Data Analysis

    Calvin Yeung, Keisuke Fujii

    cs.HCcs.AIarXiv:2609.11224v12026
  23. Thought Communication in Multiagent Collaboration

    Yujia Zheng, Zhuokai Zhao, Zijian Li +4

    cs.LGcs.AIcs.MAarXiv:2510.20733v12025
  24. Mixture of Horizons in Action Chunking

    Dong Jing, Gang Wang, Jiaqi Liu +7

    cs.ROcs.AIcs.CVarXiv:2511.19433v22025
  25. How AI Coders Discuss, Disagree, and Reach Consensus: Challenges and Opportunities for LLM-Based Qualitative Coding

    Jeongyeon Kim, John Mitchell

    cs.HCcs.AIarXiv:2609.11109v12026
  26. Code2Video: A Code-centric Paradigm for Educational Video Generation

    Yanzhe Chen, Kevin Qinghong Lin, Mike Zheng Shou

    cs.CVcs.AIcs.CLarXiv:2510.01174v12025
  27. VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models

    Xinlei Yu, Chengming Xu, Guibin Zhang +7

    cs.CVcs.AIcs.LGarXiv:2511.11007v22025
  28. DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning

    Shih-Yang Liu, Xin Dong, Ximing Lu +9

    cs.LGcs.AIcs.CLarXiv:2510.15110v12025
  29. Understanding the Effect of Out-of-distribution Examples and Interactive Explanations on Human-AI Decision Making

    Han Liu, Vivian Lai, Chenhao Tan

    cs.AIcs.CYcs.HCarXiv:2101.05303v42021
  30. Multilingual Routing in Mixture-of-Experts

    Lucas Bandarkar, Chenyuan Yang, Mohsen Fayyaz +2

    cs.CLcs.AIcs.LGarXiv:2510.04694v22025
  31. Demystifying the Privacy-Utility Trade-off in LLM Interactions

    Zhenhua Liu, Zhanxu Xie, Junjie Yu +3

    cs.AIcs.CRarXiv:2609.10992v12026
  32. A Survey of Data Agents: Emerging Paradigm or Overstated Hype?

    Yizhang Zhu, Liangwei Wang, Chenyu Yang +22

    cs.DBcs.AIarXiv:2510.23587v22025
  33. ToolUniverse: An open platform for democratizing AI scientists

    Shanghua Gao, Richard Zhu, Pengwei Sui +8

    cs.AIcs.LGarXiv:2509.23426v32025
  34. terms.txt: A Consent and Compensation Protocol for Agentic Web Access

    Rajarshi Chowdhury

    cs.NIcs.AIcs.CRarXiv:2609.11152v12026
  35. LITA: Language Instructed Temporal-Localization Assistant

    De-An Huang, Shijia Liao, Subhashree Radhakrishnan +4

    cs.CVcs.AIarXiv:2403.19046v12024
  36. Distinguishing cause from effect using observational data: methods and benchmarks

    Joris M. Mooij, Jonas Peters, Dominik Janzing +2

    cs.LGcs.AIstat.MLarXiv:1412.3773v32014
  37. QuantiPhy: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models

    Li Puyin, Tiange Xiang, Ella Mao +5

    cs.AIarXiv:2512.19526v12025
  38. Muon Outperforms Adam in Tail-End Associative Memory Learning

    Shuche Wang, Fengzhuo Zhang, Jiaxiang Li +6

    cs.LGcs.AImath.OCarXiv:2509.26030v22025
  39. On the Pitfalls of Measuring Emergent Communication

    Ryan Lowe, Jakob Foerster, Y-Lan Boureau +2

    cs.LGcs.AIcs.CLarXiv:1903.05168v12019
  40. Meta-RL Induces Exploration in Language Agents

    Yulun Jiang, Liangze Jiang, Damien Teney +2

    cs.LGcs.AIarXiv:2512.16848v22025
  41. Paper2Video: Automatic Video Generation from Scientific Papers

    Zeyu Zhu, Kevin Qinghong Lin, Mike Zheng Shou

    cs.CVcs.AIcs.CLarXiv:2510.05096v22025
  42. MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes

    Yu Ying Chiu, Michael S. Lee, Rachel Calcott +17

    cs.CLcs.AIcs.CYarXiv:2510.16380v22025
  43. LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild

    Jiayu Wang, Yifei Ming, Riya Dulepet +7

    cs.AIarXiv:2510.14240v52025
  44. Remote Labor Index: Measuring AI Automation of Remote Work

    Mantas Mazeika, Alice Gatti, Cristina Menghini +44

    cs.LGcs.AIcs.CLarXiv:2510.26787v12025
  45. CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards

    Xiangyuan Xue, Yifan Zhou, Guibin Zhang +7

    cs.CLcs.AIarXiv:2510.08529v22025
  46. A theory of multiclass boosting

    Indraneel Mukherjee, Robert E. Schapire

    stat.MLcs.AIarXiv:1108.2989v12011
  47. pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation

    Hansheng Chen, Kai Zhang, Hao Tan +3

    cs.LGcs.AIcs.CVarXiv:2510.14974v32025
  48. Deep-Fake CAPTCHA: Mitigating Next-Generation Social Engineering Attacks

    Guy Frankovits, Lior Yasur, Fred M. Grabovski +1

    cs.CRcs.AIarXiv:2609.11404v12026
  49. ARE: Scaling Up Agent Environments and Evaluations

    Romain Froger, Pierre Andrews, Matteo Bettini +21

    cs.AIcs.CLarXiv:2509.17158v22025
  50. Building a Foundational Guardrail for General Agentic Systems via Synthetic Data

    Yue Huang, Hang Hua, Yujun Zhou +11

    cs.LGcs.AIcs.CLarXiv:2510.09781v12025
  51. SynthID-Image: Image watermarking at internet scale

    Sven Gowal, Rudy Bunel, Florian Stimberg +23

    cs.CRcs.AIarXiv:2510.09263v12025
  52. No-Box Vulnerability Analysis: Description-only Detection of Indirect Prompt Injection Vulnerabilities in MCP Servers

    Zehua Zhang, Jie Hu, Pratham Hegde +13

    cs.CRcs.AIarXiv:2609.10854v12026
  53. Are We on the Right Way to Assessing LLM-as-a-Judge?

    Yuanning Feng, Sinan Wang, Zhengxiang Cheng +2

    cs.CLcs.AIarXiv:2512.16041v12025
  54. A Simple but Tough-to-Beat Data Augmentation Approach for Natural Language Understanding and Generation

    Dinghan Shen, Mingzhi Zheng, Yelong Shen +2

    cs.CLcs.AIarXiv:2009.13818v22020
  55. Architecting the Secure AI-SOC: A Neurosymbolic Framework for Pipeline Integrity and Threat Mitigation

    Anna Gazani, Spyridon Kounoupidis, Panagiotis Katsaros +3

    cs.CRcs.AIarXiv:2609.10707v12026
  56. TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models

    Zheng Ding, Weirui Ye

    cs.LGcs.AIcs.CVarXiv:2512.08153v12025
  57. Responsible Autonomy

    Virginia Dignum

    cs.AIarXiv:1706.02513v12017
  58. SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

    Qibai Chen, Zeming Liu

    cs.AIcs.SEarXiv:2609.11180v12026
  59. BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure

    Shenghan Zheng, Zonglin Di, Yimin Liu +19

    cs.CRcs.AIcs.SEarXiv:2609.11028v12026
    Summaries:한국어
  60. Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents

    Susheel Suresh, Hazel Mak, Sahil Bhatnagar +2

    cs.AIcs.SEarXiv:2609.11060v12026