Artificial Intelligence
Papers filed under cs.AI on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
5,281 to 5,340 of 15,328
Second-order Non-local Attention Networks for Person Re-identification
Bryan, Xia, Yuan Gong +2
cs.CVcs.AIcs.LGarXiv:1909.00295v12019Modeling What Changes: Sparse, Residual World Models for Object-Centric Manipulation
Param Thakkar, Parsika Paresh Shah, Manisha Sushant Gote
cs.ROcs.AIarXiv:2609.02046v12026Technology Readiness Levels for AI & ML
Alexander Lavin, Gregory Renard
cs.SEcs.AIcs.LGarXiv:2006.12497v32020Measurement-Driven Sub-Network Selection for On-Premise Retrieval-Augmented Factory Agents
Vasileios Rizeakos, Georgios Paisios, Alexandros Machairas +2
cs.AIarXiv:2609.02760v12026More Agents Is All You Need
Junyou Li, Qin Zhang, Yangbin Yu +2
cs.CLcs.AIcs.LGarXiv:2402.05120v22024SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment
Qingyu Meng, Yiwei Zha, Jiahuan Pei +3
cs.LGcs.AIcs.CRarXiv:2609.02293v12026Seed1.8 Model Card: Towards Generalized Real-World Agency
Bytedance Seed
cs.AIarXiv:2603.20633v32026Untangling the Mechanisms of Misleading Context in Medical Question Answering
Robin Linzmayer, Noémie Elhadad
cs.CLcs.AIcs.LGarXiv:2609.02754v12026Towards One-for-All Robustness Across a Continuum of Threat Levels
Zhichao Hou, Xiaorui Liu
cs.LGcs.AIarXiv:2609.02440v12026Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models
Zhifei Xie, Mingbao Lin, Zihang Liu +3
cs.SDcs.AIcs.CLarXiv:2503.02318v22025CALIP: Zero-Shot Enhancement of CLIP with Parameter-free Attention
Ziyu Guo, Renrui Zhang, Longtian Qiu +4
cs.CVcs.AIcs.MMarXiv:2209.14169v22022ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
Manasi Sharma, Chen Bo Calvin Zhang, Chaithanya Bandi +13
cs.AIcs.CLcs.LGarXiv:2511.07685v12025VoRTeC: Taming Foundation Flow for One-step Real time Video Compression
Yichong Xia, Qinhong Wu, Bin Chen +3
cs.CVcs.AIarXiv:2609.02291v22026Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving
Luoxin Chen, Jinming Gu, Liankai Huang +33
cs.AIcs.CLarXiv:2507.23726v22025OmegaUse-SOP: SOP Engineering for Professional Computer Use from Human Demonstrations
Yixiong Xiao, Lang An, Hucheng Yang +9
cs.HCcs.AIarXiv:2609.02149v22026What matters for Representation Alignment: Global Information or Spatial Structure?
Jaskirat Singh, Xingjian Leng, Zongze Wu +4
cs.CVcs.AIcs.GRarXiv:2512.10794v12025R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
Hengguang Zhou, Xirui Li, Ruochen Wang +3
cs.AIcs.CVcs.LGarXiv:2503.05132v22025KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
Zhangchen Xu, Yang Liu, Yueqin Yin +2
cs.LGcs.AIcs.CLarXiv:2503.02951v22025HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models
Renjie Xie, Juncheng Yang, Aoting Hu +4
cs.AIarXiv:2609.02029v12026Describe Anything: Detailed Localized Image and Video Captioning
Long Lian, Yifan Ding, Yunhao Ge +8
cs.CVcs.AIarXiv:2504.16072v12025The Art of Scaling Reinforcement Learning Compute for LLMs
Devvrit Khatri, Lovish Madaan, Rishabh Tiwari +6
cs.LGcs.AIarXiv:2510.13786v12025OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
Mengdi Jia, Zekun Qi, Shaochen Zhang +5
cs.CVcs.AIcs.CLarXiv:2506.03135v32025Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Ailin Huang, Boyong Wu, Bruce Wang +142
cs.CLcs.AIcs.HCarXiv:2502.11946v22025Domain and Function: A Dual-Space Model of Semantic Relations and Compositions
Peter D. Turney
cs.CLcs.AIcs.LGarXiv:1309.4035v12013Higher Structures in Deep Learning
Michael L. Roberts, Carlos Zapata Carratalá. Nicholas J. Cooper, Lijun Chen +2
cs.LGcs.AIarXiv:2609.00472v12026TimeMixer++: A General Time Series Pattern Machine for Universal Predictive Analysis
Shiyu Wang, Jiawei Li, Xiaoming Shi +6
cs.LGcs.AIarXiv:2410.16032v52024Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models
Linhai Ma, Rita El Hachem, Mahatab El Hajj +2
cs.CLcs.AIarXiv:2609.00191v12026Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
Tiansheng Huang, Sihao Hu, Fatih Ilhan +4
cs.CRcs.AIcs.LGarXiv:2503.00555v22025Online Difficulty Filtering for Reasoning Oriented Reinforcement Learning
Sanghwan Bae, Jiwoo Hong, Min Young Lee +3
cs.CLcs.AIarXiv:2504.03380v22025Induction and Inquiry via Probabilistic Reasoning over Language and Code
Wasu Top Piriyakulkij, Sam Acquaviva, Cassidy Langenfeld +2
cs.AIarXiv:2609.01815v12026Mathematical exploration and discovery at scale
Bogdan Georgiev, Javier Gómez-Serrano, Terence Tao +1
cs.NEcs.AImath.CAarXiv:2511.02864v32025Colorization Transformer
Manoj Kumar, Dirk Weissenborn, Nal Kalchbrenner
cs.CVcs.AIcs.LGarXiv:2102.04432v22021Towards Robust Mathematical Reasoning
Thang Luong, Dawsen Hwang, Hoang H. Nguyen +17
cs.CLcs.AIarXiv:2511.01846v12025The Do-Calculus Revisited
Judea Pearl
cs.AIstat.MEarXiv:1210.4852v12012ProofNet: Autoformalizing and Formally Proving Undergraduate-Level Mathematics
Zhangir Azerbayev, Bartosz Piotrowski, Hailey Schoelkopf +3
cs.CLcs.AIcs.LOarXiv:2302.12433v12023Quantifying social organization and political polarization in online platforms
Isaac Waller, Ashton Anderson
cs.SIcs.AIcs.CYarXiv:2010.00590v32020Question and Answer Test-Train Overlap in Open-Domain Question Answering Datasets
Patrick Lewis, Pontus Stenetorp, Sebastian Riedel
cs.CLcs.AIarXiv:2008.02637v12020Neural Collapse Under MSE Loss: Proximity to and Dynamics on the Central Path
X. Y. Han, Vardan Papyan, David L. Donoho
cs.LGcs.AImath.DGarXiv:2106.02073v42021VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
Puyuan Peng, Po-Yao Huang, Shang-Wen Li +2
eess.AScs.AIcs.CLarXiv:2403.16973v32024TerraMind: Large-Scale Generative Multimodality for Earth Observation
Johannes Jakubik, Felix Yang, Benedikt Blumenstiel +13
cs.CVcs.AIarXiv:2504.11171v52025Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons
Anthony Liang, Yigit Korkmaz, Jiahui Zhang +14
cs.ROcs.AIcs.LGarXiv:2603.02115v22026Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
NVIDIA, :, Aaron Blakeman +311
cs.CLcs.AIcs.LGarXiv:2512.20848v12025A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
Qiyuan Zhang, Fuyuan Lyu, Zexu Sun +10
cs.CLcs.AIarXiv:2503.24235v32025Training Agents Inside of Scalable World Models
Danijar Hafner, Wilson Yan, Timothy Lillicrap
cs.AIcs.LGcs.ROarXiv:2509.24527v12025The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm
Noah Amsel, David Persson, Christopher Musco +1
cs.LGcs.AIcs.CLarXiv:2505.16932v52025Autonomous discovery of new structure-plausibility laws for explainable and rapid crystal diagnosis and screening
Zhilong Song, Lixue Cheng
cond-mat.mtrl-scics.AIarXiv:2609.01209v12026SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems
Rui Yang, Junjie Xu, Zhengyu Liu +4
cs.CRcs.AIarXiv:2609.00595v12026Ultralytics YOLO Evolution: An Overview of YOLO26, YOLO11, YOLOv8 and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition
Ranjan Sapkota, Manoj Karkee
cs.CVcs.AIarXiv:2510.09653v32025Programming Refusal with Conditional Activation Steering
Bruce W. Lee, Inkit Padhi, Karthikeyan Natesan Ramamurthy +4
cs.LGcs.AIcs.CLarXiv:2409.05907v32024Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
Qianhao Yuan, Jie Lou, Xing Yu +4
cs.CVcs.AIcs.CLarXiv:2605.18740v42026GOOD: A Graph Out-of-Distribution Benchmark
Shurui Gui, Xiner Li, Limei Wang +1
cs.LGcs.AIarXiv:2206.08452v22022Deep Exemplar-based Video Colorization
Bo Zhang, Mingming He, Jing Liao +4
cs.CVcs.AIcs.LGarXiv:1906.09909v12019Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Haizhong Zheng, Yang Zhou, Brian R. Bartoldson +4
cs.AIcs.LGarXiv:2506.02177v12025TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
Amir Zandieh, Majid Daliri, Majid Hadian +1
cs.LGcs.AIcs.DBarXiv:2504.19874v12025GTA1: GUI Test-time Scaling Agent
Yan Yang, Dongxu Li, Yutong Dai +12
cs.AIarXiv:2507.05791v52025Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling
Kangjia Zhao, Jiajun Li, Haozhan Shen +8
cs.CLcs.AIarXiv:2609.00949v12026In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access?
Koshiro Aoki, Ryota Takatsuki, Gouki Minegishi +2
cs.AIarXiv:2609.00904v12026Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
Zeyuan Yang, Xueyang Yu, Delin Chen +2
cs.CVcs.AIarXiv:2506.17218v12025A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Wei Xiong, Jiarui Yao, Yuhui Xu +8
cs.LGcs.AIcs.CLarXiv:2504.11343v22025dLLM: Simple Diffusion Language Modeling
Zhanhui Zhou, Lingjie Chen, Hanghang Tong +1
cs.CLcs.AIcs.LGarXiv:2602.22661v12026