Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
20,041 to 20,100 of 20,205
GPT-Red: Automated Red Teaming via Self-Play at Scale
Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal +15
cs.CRcs.AIcs.CLarXiv:2607.26115v12026Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning
Junyao Yang, Yucheng Shi, Zongxia Li +6
cs.LGcs.CLarXiv:2607.18722v32026MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis
Yihao Chen, Shi Chang, Khaled Chawa +4
cs.SEcs.CLcs.LGarXiv:2607.27146v12026QQWorld: Quantile-Quantile Matching for World Model Regularization
Zhoushun Yu, Xiaoyu Hu, Xiangyu Xu
cs.LGcs.AIcs.CVarXiv:2607.28415v12026You Only Pass Once: Answering and Abstaining Together in a Single Forward Pass of a Frozen Language Model
Ziyang Luo, Zhongyao Chu, Xinjie He +4
cs.CLcs.LGarXiv:2608.14465v12026More Correct Mass, Worse Answers: Why Power Sampling Can Fail and How to Fix It
Haohui Yang, Jiaxing Sun, Xiujun Ma
cs.LGarXiv:2608.14420v12026Non-Shattering at and Above the Dynamical Temperature in the Spherical Pure p-Spin Model
Taegyun Kim
math.PRcond-mat.dis-nncs.LGarXiv:2608.14369v12026Style or Signature? Artist-Disjoint Evaluation of Style Classification in Frozen Vision Embeddings
Rory Ashton
cs.CVcs.LGarXiv:2608.14435v12026Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization
Zhicheng Cai, Xinyuan Guo, Hanlin Wu +4
cs.LGcs.AIarXiv:2607.10169v12026$π\mathbf{R}^2$: Reactive Real-time Flow Policies
Sungjae Park, Shubham Tulsiani
cs.ROcs.AIcs.LGarXiv:2607.26055v12026Can AI agents conduct open-ended AI research? Early evidence from two case studies
Peter Kirgis, Sayash Kapoor, Andrew Schwartz +21
cs.AIcs.CYcs.LGarXiv:2607.27191v22026NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs
Jiarong Zhao, Zhikai Lei, Zhiheng Xi +5
cs.SEcs.AIcs.LGarXiv:2607.14186v62026CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition
Lai Wei, Chengqi Li, Jiapeng Li +3
cs.CVcs.AIcs.CLarXiv:2607.25294v12026Flux-OPD: On-Policy Distillation with Evolving Contexts
Yuran Wang, Zekun Wang, Bohan Zeng +10
cs.LGcs.AIarXiv:2607.28022v12026Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning
Jian Hu, Huiying Li, Hao Zhang +8
cs.LGcs.CLcs.DCarXiv:2607.21653v12026ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
Yao Xiao, Reuben Tan, Zhen Zhu +3
cs.CVcs.AIcs.LGarXiv:2607.28627v12026HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
Simple AI, :, Yuteng Wei +16
cs.ROcs.CVcs.LGarXiv:2607.25895v12026Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale
Yash Pandya, Sahil Gupta, Sarthak Harne +10
cs.AIcs.LGarXiv:2607.28074v12026EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents
Luigi Sigillo, Matteo Silvestri, Francesco Tabaro +9
cs.CLcs.AIcs.IRarXiv:2607.28229v12026SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution
Zhiyuan Yao, Yuxin Chen, Zhengxi Lu +13
cs.LGcs.AIarXiv:2607.26784v12026MedMix: Specialization-Consistent Federated Sparse MoEs under Modality Heterogeneity
Adiba Orzikulova, Dong Min Kim, Jaehong Yoon +1
cs.LGarXiv:2608.13911v12026Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
Alexi Gladstone, Heng Ji, Yilun Du
cs.LGcs.AIcs.CLarXiv:2607.27372v12026Robust Dual-Model Collaborative Random Vector Functional Link Network
A. Quadir, A. Rahaman, Mushir Akhtar +1
cs.LGarXiv:2608.13628v12026Catching the Imposter: Self-Supervised Learning of Physical Coherence with Cross-Entity Feature Permutations
Aleksei Rozanov, Arvind Renganathan, Vipin Kumar
cs.LGarXiv:2608.14372v12026Weak-to-Strong On-Policy Distillation
Fangxu Yu, Weijia Xu, Michael Xu +2
cs.LGarXiv:2607.26246v22026A Vocabulary for Multi-Agent Automated Research Systems
Bardiya Akhbari
cs.AIcs.LGcs.MAarXiv:2607.22682v12026ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow
Dongxiu Liu, Haoyi Niu, Peng Cheng +5
cs.LGcs.CVcs.ROarXiv:2607.27924v22026When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation
Jiabin Shen, Guang Chen, Chengjun Mao
cs.CLcs.LGarXiv:2607.07050v42026AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
Bing Yan, Gregory Wolfe, Stefano Martiniani +1
cs.CLcs.AIcs.IRarXiv:2607.28618v12026An Exam for Active Observers
Jiarui Zhang, Muzi Tao, Shangshang Wang +3
cs.CVcs.AIcs.CLarXiv:2607.16165v12026Interactive Training 2: Auditable Control Plane for Live Model Training
Wentao Zhang, Xuanhe Pan, Han Zhou +2
cs.LGarXiv:2607.18314v12026Metis: Memory Foundation Model
Zeyu Zhang, Ziliang Guo, Yihang Sun +14
cs.CLcs.LGarXiv:2607.26760v22026CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation
Satyam Kumar, Saurabh Jha
cs.LGcs.AIarXiv:2607.16955v12026Multi-Turn On-Policy Distillation with Prefix Replay
Baohao Liao, Hanze Dong, Christof Monz +3
cs.LGcs.AIcs.CLarXiv:2607.04763v32026Agentic Transaction: Towards ACID-Compliant Agent Systems
Zhaoyan Sun, Xiaoxiao Wang, Guoliang Li
cs.DBcs.AIcs.CLarXiv:2608.13900v12026Engineering Signals of Human-AI Collaboration in the Agentic Coding Era: A Longitudinal Analysis of 33,228 Pull Requests from vLLM and SGLang with Implications for Biomedical AI Agents and Bioinformatics Pipeline Developmen
Jiada Li, Xuesong Ye, Olamide Olowoniyi
cs.SEcs.AIcs.ETarXiv:2608.13884v12026Split the Labor: Separating Evidence Interpretation from Decision Aggregation
Zhelun Wu
cs.AIcs.CLcs.LGarXiv:2608.14509v12026Attributing Preprocessing Invariance in Spectral Foundation Models
Dongjun Wei, Hongyi Wu, Yinuo Zou
cs.AIcs.CEcs.LGarXiv:2608.14227v12026Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis
Aryan Luthra, Kshitij Jain, Siddharth Arya +2
cs.AIcs.CRcs.LGarXiv:2608.13608v12026The Query Knows What to Forget: A Second Erase Direction for Linear Attention
Dhruman Gupta, Aritra Das, Debayan Gupta
cs.LGarXiv:2608.13668v12026Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation
Scott H. Hawley
cs.SDcs.LGeess.ASarXiv:2608.04378v12026FreeBalance: Pre-Routing Online Moe Load Balancing via Residual Workload Prediction
Pengfei Chen, Yize Wu, Shouxu Kuang +2
cs.AIcs.LGarXiv:2608.14205v12026Removing Temporal Note Redundancy Improves Multimodal Reinforcement Learning for Medicine
Chenran Weng, Joo Seung Lee, Malini Mahendra +1
cs.AIcs.LGarXiv:2608.14157v12026DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack
Hoseong Tae, Jong-Seok Lee
cs.CVcs.LGarXiv:2608.03207v12026Adjacency-Based Spectral Proxy Control of Mobile Communication Agents
Mariana del Castillo, Federico Larroca
cs.ROcs.LGcs.MAarXiv:2608.13616v12026Omega-S: A Functional Resilience Index for LLM Fine-Tuning
Alberto Acedo
cs.LGcs.NEq-bio.MNarXiv:2608.03887v12026FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds
Kapil Wanaskar, Gaytri Jena, Aman Chadha +3
cs.AIcs.CVcs.LGarXiv:2608.01049v12026To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing
Amir M. Ebrahimi, Mohammed Mehedi Hasan, Aaditya Bhatia +2
cs.SEcs.AIcs.LGarXiv:2607.28887v12026Towards Interpretable Foundation Models for Retinal Fundus Images
Samuel Ofosu Mensah, Camila Roa, Kerol Djoumessi +1
cs.CVcs.LGstat.COarXiv:2603.18846v42026EEG-PRISM: Physiologically-Grounded Interpretability of Predictions by EEG Foundation Models
Deeksha M Shama, Punnisa Amornsirikul, Archana Venkataraman
cs.LGarXiv:2608.13676v12026Sequence prediction under a lying oracle
Puspabeethi Samanta, Nikhil Karamchandani, Jayakrishnan Nair
cs.LGcs.ITarXiv:2608.14102v12026Model-agnostic Retrieval-Augmented Extended Forecasting for time series
Juan Pablo Villa Serna, Rohan Asthana, Vasileios Belagiannis
cs.LGarXiv:2608.14054v12026Geometric Filtering of LLM-Generated Samples for Few-Shot Text Classification
Benjamín Schindler, Gonzalo A. Ruz
cs.LGcs.CLarXiv:2608.13866v12026Structure-Guided Spatiotemporal Attention Graph Neural Network for Traffic Flow Prediction
Xuanmian He, Can Li, Wanjing Ma
cs.LGcs.AIarXiv:2608.14177v12026RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction
Changwoo Baek, Seungjun Shin, Kyeongbo Kong
cs.CLcs.LGarXiv:2608.01247v12026ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads
Şuayp Talha Kocabay, Talha Rüzgar Akkuş, Kamer Ali Yuksel
cs.CLcs.LGarXiv:2608.02703v12026Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors
Alexander Scheinker
stat.MLcs.LGphysics.comp-pharXiv:2608.00675v22026SAGE: Surrogate-gradient Adaptation via Attention-Guided Entropy for Spiking Transformers
Kiran Nair, Rodrigue Rizk, KC Santosh
cs.LGcs.AIcs.CVarXiv:2608.13702v12026AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
Dong Yan, Jian Liang, Dapeng Hu +4
cs.AIcs.LGarXiv:2608.00155v12026TRUE-Colon: Exposing a Consistent Transfer Asymmetry in Real-Time Polyp Detection
Sebastian Doerrich, Andreas Franz Schwab, Francesco Di Salvo +3
eess.IVcs.CVcs.LGarXiv:2608.13711v12026