Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

20,041 to 20,100 of 20,205

  1. GPT-Red: Automated Red Teaming via Self-Play at Scale

    Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal +15

    cs.CRcs.AIcs.CLarXiv:2607.26115v12026
  2. Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning

    Junyao Yang, Yucheng Shi, Zongxia Li +6

    cs.LGcs.CLarXiv:2607.18722v32026
  3. MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis

    Yihao Chen, Shi Chang, Khaled Chawa +4

    cs.SEcs.CLcs.LGarXiv:2607.27146v12026
  4. QQWorld: Quantile-Quantile Matching for World Model Regularization

    Zhoushun Yu, Xiaoyu Hu, Xiangyu Xu

    cs.LGcs.AIcs.CVarXiv:2607.28415v12026
  5. You Only Pass Once: Answering and Abstaining Together in a Single Forward Pass of a Frozen Language Model

    Ziyang Luo, Zhongyao Chu, Xinjie He +4

    cs.CLcs.LGarXiv:2608.14465v12026
  6. More Correct Mass, Worse Answers: Why Power Sampling Can Fail and How to Fix It

    Haohui Yang, Jiaxing Sun, Xiujun Ma

    cs.LGarXiv:2608.14420v12026
  7. Non-Shattering at and Above the Dynamical Temperature in the Spherical Pure p-Spin Model

    Taegyun Kim

    math.PRcond-mat.dis-nncs.LGarXiv:2608.14369v12026
  8. Style or Signature? Artist-Disjoint Evaluation of Style Classification in Frozen Vision Embeddings

    Rory Ashton

    cs.CVcs.LGarXiv:2608.14435v12026
  9. Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization

    Zhicheng Cai, Xinyuan Guo, Hanlin Wu +4

    cs.LGcs.AIarXiv:2607.10169v12026
  10. $π\mathbf{R}^2$: Reactive Real-time Flow Policies

    Sungjae Park, Shubham Tulsiani

    cs.ROcs.AIcs.LGarXiv:2607.26055v12026
  11. Can AI agents conduct open-ended AI research? Early evidence from two case studies

    Peter Kirgis, Sayash Kapoor, Andrew Schwartz +21

    cs.AIcs.CYcs.LGarXiv:2607.27191v22026
  12. NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs

    Jiarong Zhao, Zhikai Lei, Zhiheng Xi +5

    cs.SEcs.AIcs.LGarXiv:2607.14186v62026
  13. CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

    Lai Wei, Chengqi Li, Jiapeng Li +3

    cs.CVcs.AIcs.CLarXiv:2607.25294v12026
  14. Flux-OPD: On-Policy Distillation with Evolving Contexts

    Yuran Wang, Zekun Wang, Bohan Zeng +10

    cs.LGcs.AIarXiv:2607.28022v12026
  15. Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning

    Jian Hu, Huiying Li, Hao Zhang +8

    cs.LGcs.CLcs.DCarXiv:2607.21653v12026
  16. ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

    Yao Xiao, Reuben Tan, Zhen Zhu +3

    cs.CVcs.AIcs.LGarXiv:2607.28627v12026
  17. HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

    Simple AI, :, Yuteng Wei +16

    cs.ROcs.CVcs.LGarXiv:2607.25895v12026
  18. Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale

    Yash Pandya, Sahil Gupta, Sarthak Harne +10

    cs.AIcs.LGarXiv:2607.28074v12026
  19. EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents

    Luigi Sigillo, Matteo Silvestri, Francesco Tabaro +9

    cs.CLcs.AIcs.IRarXiv:2607.28229v12026
  20. SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

    Zhiyuan Yao, Yuxin Chen, Zhengxi Lu +13

    cs.LGcs.AIarXiv:2607.26784v12026
  21. MedMix: Specialization-Consistent Federated Sparse MoEs under Modality Heterogeneity

    Adiba Orzikulova, Dong Min Kim, Jaehong Yoon +1

    cs.LGarXiv:2608.13911v12026
  22. Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation

    Alexi Gladstone, Heng Ji, Yilun Du

    cs.LGcs.AIcs.CLarXiv:2607.27372v12026
  23. Robust Dual-Model Collaborative Random Vector Functional Link Network

    A. Quadir, A. Rahaman, Mushir Akhtar +1

    cs.LGarXiv:2608.13628v12026
  24. Catching the Imposter: Self-Supervised Learning of Physical Coherence with Cross-Entity Feature Permutations

    Aleksei Rozanov, Arvind Renganathan, Vipin Kumar

    cs.LGarXiv:2608.14372v12026
  25. Weak-to-Strong On-Policy Distillation

    Fangxu Yu, Weijia Xu, Michael Xu +2

    cs.LGarXiv:2607.26246v22026
  26. A Vocabulary for Multi-Agent Automated Research Systems

    Bardiya Akhbari

    cs.AIcs.LGcs.MAarXiv:2607.22682v12026
  27. ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow

    Dongxiu Liu, Haoyi Niu, Peng Cheng +5

    cs.LGcs.CVcs.ROarXiv:2607.27924v22026
  28. When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation

    Jiabin Shen, Guang Chen, Chengjun Mao

    cs.CLcs.LGarXiv:2607.07050v42026
  29. AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis

    Bing Yan, Gregory Wolfe, Stefano Martiniani +1

    cs.CLcs.AIcs.IRarXiv:2607.28618v12026
  30. An Exam for Active Observers

    Jiarui Zhang, Muzi Tao, Shangshang Wang +3

    cs.CVcs.AIcs.CLarXiv:2607.16165v12026
  31. Interactive Training 2: Auditable Control Plane for Live Model Training

    Wentao Zhang, Xuanhe Pan, Han Zhou +2

    cs.LGarXiv:2607.18314v12026
  32. Metis: Memory Foundation Model

    Zeyu Zhang, Ziliang Guo, Yihang Sun +14

    cs.CLcs.LGarXiv:2607.26760v22026
  33. CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation

    Satyam Kumar, Saurabh Jha

    cs.LGcs.AIarXiv:2607.16955v12026
  34. Multi-Turn On-Policy Distillation with Prefix Replay

    Baohao Liao, Hanze Dong, Christof Monz +3

    cs.LGcs.AIcs.CLarXiv:2607.04763v32026
  35. Agentic Transaction: Towards ACID-Compliant Agent Systems

    Zhaoyan Sun, Xiaoxiao Wang, Guoliang Li

    cs.DBcs.AIcs.CLarXiv:2608.13900v12026
  36. Engineering Signals of Human-AI Collaboration in the Agentic Coding Era: A Longitudinal Analysis of 33,228 Pull Requests from vLLM and SGLang with Implications for Biomedical AI Agents and Bioinformatics Pipeline Developmen

    Jiada Li, Xuesong Ye, Olamide Olowoniyi

    cs.SEcs.AIcs.ETarXiv:2608.13884v12026
  37. Split the Labor: Separating Evidence Interpretation from Decision Aggregation

    Zhelun Wu

    cs.AIcs.CLcs.LGarXiv:2608.14509v12026
  38. Attributing Preprocessing Invariance in Spectral Foundation Models

    Dongjun Wei, Hongyi Wu, Yinuo Zou

    cs.AIcs.CEcs.LGarXiv:2608.14227v12026
  39. Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis

    Aryan Luthra, Kshitij Jain, Siddharth Arya +2

    cs.AIcs.CRcs.LGarXiv:2608.13608v12026
  40. The Query Knows What to Forget: A Second Erase Direction for Linear Attention

    Dhruman Gupta, Aritra Das, Debayan Gupta

    cs.LGarXiv:2608.13668v12026
  41. Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation

    Scott H. Hawley

    cs.SDcs.LGeess.ASarXiv:2608.04378v12026
  42. FreeBalance: Pre-Routing Online Moe Load Balancing via Residual Workload Prediction

    Pengfei Chen, Yize Wu, Shouxu Kuang +2

    cs.AIcs.LGarXiv:2608.14205v12026
  43. Removing Temporal Note Redundancy Improves Multimodal Reinforcement Learning for Medicine

    Chenran Weng, Joo Seung Lee, Malini Mahendra +1

    cs.AIcs.LGarXiv:2608.14157v12026
  44. DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack

    Hoseong Tae, Jong-Seok Lee

    cs.CVcs.LGarXiv:2608.03207v12026
  45. Adjacency-Based Spectral Proxy Control of Mobile Communication Agents

    Mariana del Castillo, Federico Larroca

    cs.ROcs.LGcs.MAarXiv:2608.13616v12026
  46. Omega-S: A Functional Resilience Index for LLM Fine-Tuning

    Alberto Acedo

    cs.LGcs.NEq-bio.MNarXiv:2608.03887v12026
  47. FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds

    Kapil Wanaskar, Gaytri Jena, Aman Chadha +3

    cs.AIcs.CVcs.LGarXiv:2608.01049v12026
  48. To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing

    Amir M. Ebrahimi, Mohammed Mehedi Hasan, Aaditya Bhatia +2

    cs.SEcs.AIcs.LGarXiv:2607.28887v12026
  49. Towards Interpretable Foundation Models for Retinal Fundus Images

    Samuel Ofosu Mensah, Camila Roa, Kerol Djoumessi +1

    cs.CVcs.LGstat.COarXiv:2603.18846v42026
  50. EEG-PRISM: Physiologically-Grounded Interpretability of Predictions by EEG Foundation Models

    Deeksha M Shama, Punnisa Amornsirikul, Archana Venkataraman

    cs.LGarXiv:2608.13676v12026
  51. Sequence prediction under a lying oracle

    Puspabeethi Samanta, Nikhil Karamchandani, Jayakrishnan Nair

    cs.LGcs.ITarXiv:2608.14102v12026
  52. Model-agnostic Retrieval-Augmented Extended Forecasting for time series

    Juan Pablo Villa Serna, Rohan Asthana, Vasileios Belagiannis

    cs.LGarXiv:2608.14054v12026
  53. Geometric Filtering of LLM-Generated Samples for Few-Shot Text Classification

    Benjamín Schindler, Gonzalo A. Ruz

    cs.LGcs.CLarXiv:2608.13866v12026
  54. Structure-Guided Spatiotemporal Attention Graph Neural Network for Traffic Flow Prediction

    Xuanmian He, Can Li, Wanjing Ma

    cs.LGcs.AIarXiv:2608.14177v12026
  55. RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction

    Changwoo Baek, Seungjun Shin, Kyeongbo Kong

    cs.CLcs.LGarXiv:2608.01247v12026
  56. ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads

    Şuayp Talha Kocabay, Talha Rüzgar Akkuş, Kamer Ali Yuksel

    cs.CLcs.LGarXiv:2608.02703v12026
  57. Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors

    Alexander Scheinker

    stat.MLcs.LGphysics.comp-pharXiv:2608.00675v22026
  58. SAGE: Surrogate-gradient Adaptation via Attention-Guided Entropy for Spiking Transformers

    Kiran Nair, Rodrigue Rizk, KC Santosh

    cs.LGcs.AIcs.CVarXiv:2608.13702v12026
  59. AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

    Dong Yan, Jian Liang, Dapeng Hu +4

    cs.AIcs.LGarXiv:2608.00155v12026
  60. TRUE-Colon: Exposing a Consistent Transfer Asymmetry in Real-Time Polyp Detection

    Sebastian Doerrich, Andreas Franz Schwab, Francesco Di Salvo +3

    eess.IVcs.CVcs.LGarXiv:2608.13711v12026