Artificial Intelligence

Papers filed under cs.AI on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

5,281 to 5,340 of 15,328

  1. Second-order Non-local Attention Networks for Person Re-identification

    Bryan, Xia, Yuan Gong +2

    cs.CVcs.AIcs.LGarXiv:1909.00295v12019
  2. Modeling What Changes: Sparse, Residual World Models for Object-Centric Manipulation

    Param Thakkar, Parsika Paresh Shah, Manisha Sushant Gote

    cs.ROcs.AIarXiv:2609.02046v12026
  3. Technology Readiness Levels for AI & ML

    Alexander Lavin, Gregory Renard

    cs.SEcs.AIcs.LGarXiv:2006.12497v32020
  4. Measurement-Driven Sub-Network Selection for On-Premise Retrieval-Augmented Factory Agents

    Vasileios Rizeakos, Georgios Paisios, Alexandros Machairas +2

    cs.AIarXiv:2609.02760v12026
  5. More Agents Is All You Need

    Junyou Li, Qin Zhang, Yangbin Yu +2

    cs.CLcs.AIcs.LGarXiv:2402.05120v22024
  6. SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment

    Qingyu Meng, Yiwei Zha, Jiahuan Pei +3

    cs.LGcs.AIcs.CRarXiv:2609.02293v12026
  7. Seed1.8 Model Card: Towards Generalized Real-World Agency

    Bytedance Seed

    cs.AIarXiv:2603.20633v32026
  8. Untangling the Mechanisms of Misleading Context in Medical Question Answering

    Robin Linzmayer, Noémie Elhadad

    cs.CLcs.AIcs.LGarXiv:2609.02754v12026
  9. Towards One-for-All Robustness Across a Continuum of Threat Levels

    Zhichao Hou, Xiaorui Liu

    cs.LGcs.AIarXiv:2609.02440v12026
  10. Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models

    Zhifei Xie, Mingbao Lin, Zihang Liu +3

    cs.SDcs.AIcs.CLarXiv:2503.02318v22025
  11. CALIP: Zero-Shot Enhancement of CLIP with Parameter-free Attention

    Ziyu Guo, Renrui Zhang, Longtian Qiu +4

    cs.CVcs.AIcs.MMarXiv:2209.14169v22022
  12. ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents

    Manasi Sharma, Chen Bo Calvin Zhang, Chaithanya Bandi +13

    cs.AIcs.CLcs.LGarXiv:2511.07685v12025
  13. VoRTeC: Taming Foundation Flow for One-step Real time Video Compression

    Yichong Xia, Qinhong Wu, Bin Chen +3

    cs.CVcs.AIarXiv:2609.02291v22026
  14. Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving

    Luoxin Chen, Jinming Gu, Liankai Huang +33

    cs.AIcs.CLarXiv:2507.23726v22025
  15. OmegaUse-SOP: SOP Engineering for Professional Computer Use from Human Demonstrations

    Yixiong Xiao, Lang An, Hucheng Yang +9

    cs.HCcs.AIarXiv:2609.02149v22026
  16. What matters for Representation Alignment: Global Information or Spatial Structure?

    Jaskirat Singh, Xingjian Leng, Zongze Wu +4

    cs.CVcs.AIcs.GRarXiv:2512.10794v12025
  17. R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

    Hengguang Zhou, Xirui Li, Ruochen Wang +3

    cs.AIcs.CVcs.LGarXiv:2503.05132v22025
  18. KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding

    Zhangchen Xu, Yang Liu, Yueqin Yin +2

    cs.LGcs.AIcs.CLarXiv:2503.02951v22025
  19. HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models

    Renjie Xie, Juncheng Yang, Aoting Hu +4

    cs.AIarXiv:2609.02029v12026
  20. Describe Anything: Detailed Localized Image and Video Captioning

    Long Lian, Yifan Ding, Yunhao Ge +8

    cs.CVcs.AIarXiv:2504.16072v12025
  21. The Art of Scaling Reinforcement Learning Compute for LLMs

    Devvrit Khatri, Lovish Madaan, Rishabh Tiwari +6

    cs.LGcs.AIarXiv:2510.13786v12025
  22. OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models

    Mengdi Jia, Zekun Qi, Shaochen Zhang +5

    cs.CVcs.AIcs.CLarXiv:2506.03135v32025
  23. Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

    Ailin Huang, Boyong Wu, Bruce Wang +142

    cs.CLcs.AIcs.HCarXiv:2502.11946v22025
  24. Domain and Function: A Dual-Space Model of Semantic Relations and Compositions

    Peter D. Turney

    cs.CLcs.AIcs.LGarXiv:1309.4035v12013
  25. Higher Structures in Deep Learning

    Michael L. Roberts, Carlos Zapata Carratalá. Nicholas J. Cooper, Lijun Chen +2

    cs.LGcs.AIarXiv:2609.00472v12026
  26. TimeMixer++: A General Time Series Pattern Machine for Universal Predictive Analysis

    Shiyu Wang, Jiawei Li, Xiaoming Shi +6

    cs.LGcs.AIarXiv:2410.16032v52024
  27. Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models

    Linhai Ma, Rita El Hachem, Mahatab El Hajj +2

    cs.CLcs.AIarXiv:2609.00191v12026
  28. Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable

    Tiansheng Huang, Sihao Hu, Fatih Ilhan +4

    cs.CRcs.AIcs.LGarXiv:2503.00555v22025
  29. Online Difficulty Filtering for Reasoning Oriented Reinforcement Learning

    Sanghwan Bae, Jiwoo Hong, Min Young Lee +3

    cs.CLcs.AIarXiv:2504.03380v22025
  30. Induction and Inquiry via Probabilistic Reasoning over Language and Code

    Wasu Top Piriyakulkij, Sam Acquaviva, Cassidy Langenfeld +2

    cs.AIarXiv:2609.01815v12026
  31. Mathematical exploration and discovery at scale

    Bogdan Georgiev, Javier Gómez-Serrano, Terence Tao +1

    cs.NEcs.AImath.CAarXiv:2511.02864v32025
  32. Colorization Transformer

    Manoj Kumar, Dirk Weissenborn, Nal Kalchbrenner

    cs.CVcs.AIcs.LGarXiv:2102.04432v22021
  33. Towards Robust Mathematical Reasoning

    Thang Luong, Dawsen Hwang, Hoang H. Nguyen +17

    cs.CLcs.AIarXiv:2511.01846v12025
  34. The Do-Calculus Revisited

    Judea Pearl

    cs.AIstat.MEarXiv:1210.4852v12012
  35. ProofNet: Autoformalizing and Formally Proving Undergraduate-Level Mathematics

    Zhangir Azerbayev, Bartosz Piotrowski, Hailey Schoelkopf +3

    cs.CLcs.AIcs.LOarXiv:2302.12433v12023
  36. Quantifying social organization and political polarization in online platforms

    Isaac Waller, Ashton Anderson

    cs.SIcs.AIcs.CYarXiv:2010.00590v32020
  37. Question and Answer Test-Train Overlap in Open-Domain Question Answering Datasets

    Patrick Lewis, Pontus Stenetorp, Sebastian Riedel

    cs.CLcs.AIarXiv:2008.02637v12020
  38. Neural Collapse Under MSE Loss: Proximity to and Dynamics on the Central Path

    X. Y. Han, Vardan Papyan, David L. Donoho

    cs.LGcs.AImath.DGarXiv:2106.02073v42021
  39. VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild

    Puyuan Peng, Po-Yao Huang, Shang-Wen Li +2

    eess.AScs.AIcs.CLarXiv:2403.16973v32024
  40. TerraMind: Large-Scale Generative Multimodality for Earth Observation

    Johannes Jakubik, Felix Yang, Benedikt Blumenstiel +13

    cs.CVcs.AIarXiv:2504.11171v52025
  41. Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons

    Anthony Liang, Yigit Korkmaz, Jiahui Zhang +14

    cs.ROcs.AIcs.LGarXiv:2603.02115v22026
  42. Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

    NVIDIA, :, Aaron Blakeman +311

    cs.CLcs.AIcs.LGarXiv:2512.20848v12025
  43. A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?

    Qiyuan Zhang, Fuyuan Lyu, Zexu Sun +10

    cs.CLcs.AIarXiv:2503.24235v32025
  44. Training Agents Inside of Scalable World Models

    Danijar Hafner, Wilson Yan, Timothy Lillicrap

    cs.AIcs.LGcs.ROarXiv:2509.24527v12025
  45. The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm

    Noah Amsel, David Persson, Christopher Musco +1

    cs.LGcs.AIcs.CLarXiv:2505.16932v52025
  46. Autonomous discovery of new structure-plausibility laws for explainable and rapid crystal diagnosis and screening

    Zhilong Song, Lixue Cheng

    cond-mat.mtrl-scics.AIarXiv:2609.01209v12026
  47. SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems

    Rui Yang, Junjie Xu, Zhengyu Liu +4

    cs.CRcs.AIarXiv:2609.00595v12026
  48. Ultralytics YOLO Evolution: An Overview of YOLO26, YOLO11, YOLOv8 and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition

    Ranjan Sapkota, Manoj Karkee

    cs.CVcs.AIarXiv:2510.09653v32025
  49. Programming Refusal with Conditional Activation Steering

    Bruce W. Lee, Inkit Padhi, Karthikeyan Natesan Ramamurthy +4

    cs.LGcs.AIcs.CLarXiv:2409.05907v32024
  50. Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation

    Qianhao Yuan, Jie Lou, Xing Yu +4

    cs.CVcs.AIcs.CLarXiv:2605.18740v42026
  51. GOOD: A Graph Out-of-Distribution Benchmark

    Shurui Gui, Xiner Li, Limei Wang +1

    cs.LGcs.AIarXiv:2206.08452v22022
  52. Deep Exemplar-based Video Colorization

    Bo Zhang, Mingming He, Jing Liao +4

    cs.CVcs.AIcs.LGarXiv:1906.09909v12019
  53. Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

    Haizhong Zheng, Yang Zhou, Brian R. Bartoldson +4

    cs.AIcs.LGarXiv:2506.02177v12025
  54. TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate

    Amir Zandieh, Majid Daliri, Majid Hadian +1

    cs.LGcs.AIcs.DBarXiv:2504.19874v12025
  55. GTA1: GUI Test-time Scaling Agent

    Yan Yang, Dongxu Li, Yutong Dai +12

    cs.AIarXiv:2507.05791v52025
  56. Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling

    Kangjia Zhao, Jiajun Li, Haozhan Shen +8

    cs.CLcs.AIarXiv:2609.00949v12026
  57. In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access?

    Koshiro Aoki, Ryota Takatsuki, Gouki Minegishi +2

    cs.AIarXiv:2609.00904v12026
  58. Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

    Zeyuan Yang, Xueyang Yu, Delin Chen +2

    cs.CVcs.AIarXiv:2506.17218v12025
  59. A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

    Wei Xiong, Jiarui Yao, Yuhui Xu +8

    cs.LGcs.AIcs.CLarXiv:2504.11343v22025
  60. dLLM: Simple Diffusion Language Modeling

    Zhanhui Zhou, Lingjie Chen, Hanghang Tong +1

    cs.CLcs.AIcs.LGarXiv:2602.22661v12026