Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

7,861 to 7,920 of 20,199

  1. MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning

    Zirui Cheng, Xun Xu, Tiankai Chen +7

    cs.LGarXiv:2608.12724v12026
  2. Small-Scale Experiments: Are We There Yet?

    Nicholas Lourie, Kyunghyun Cho, Karen Ullrich +1

    cs.LGarXiv:2608.11859v12026
  3. Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing

    Tianci Liu, Zihan Dong, Tianchun Li +8

    cs.CLcs.AIcs.LGarXiv:2608.11660v12026
  4. UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations

    Dvir Samuel, Guy Bar-Shalom, Fabrizio Frasca +4

    cs.CVcs.LGarXiv:2608.10835v12026
  5. Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference

    Burc Gokden

    cs.LGcs.CLarXiv:2608.10288v12026
  6. Multimodal Model Diffing for Feature Discovery and Control

    Hunar Batra, Lachin Naghashyar, Ashkan Khakzar +4

    cs.CVcs.AIcs.CLarXiv:2608.09928v12026
  7. Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA

    Mind Lab, :, Vin Bo +74

    cs.LGcs.CLarXiv:2608.09819v12026
  8. RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance

    Dongchi Huang, Hongyin Zhang, Bohan Hou +12

    cs.ROcs.CVcs.LGarXiv:2608.09853v12026
  9. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning

    Björn Engdahl, Adrian Kosowski, Jan Chorowski +6

    cs.NEcs.AIcs.LGarXiv:2608.09888v12026
  10. Learning Locomotion Skills Using DeepRL: Does the Choice of Action Space Matter?

    Xue Bin Peng, Michiel van de Panne

    cs.LGcs.GRcs.ROarXiv:1611.01055v12016
  11. Parameter Exploration for RLVR via Variational Learning

    Vatsal Venkatkrishna, Nico Daheim, Iryna Gurevych

    cs.LGcs.AIcs.CLarXiv:2608.09805v12026
  12. Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail

    NVIDIA, :, Yan Wang +41

    cs.ROcs.AIcs.LGarXiv:2511.00088v22025
  13. OpenThoughts: Data Recipes for Reasoning Models

    Etash Guha, Ryan Marten, Sedrick Keh +47

    cs.LGarXiv:2506.04178v22025
  14. Direct Optimization of a 3D Finite-Source Reflector via Neural-Network Parameterization

    Roel Hacking, Lisa Kusch, Martijn Anthonissen +1

    physics.opticscs.LGarXiv:2609.00899v12026
  15. Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory

    Taeil Kim, Kangsan Kim, Sung Ju Hwang

    cs.AIcs.LGarXiv:2608.07169v12026
  16. CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

    Fanzhe Meng, Guoxin Chen, Jiale Zhao +6

    cs.LGcs.CLarXiv:2608.06352v12026
  17. Kimi K3: Open Frontier Intelligence

    Kimi Team, Tongtong Bai, Yifan Bai +399

    cs.CLcs.LGarXiv:2607.24653v22026
  18. EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video

    Ryan Hoque, Peide Huang, David J. Yoon +2

    cs.CVcs.LGcs.ROarXiv:2505.11709v32025
  19. From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models

    Jiale Han, Xiang Li, Jing Qian +7

    cs.AIcs.LGarXiv:2608.06020v12026
  20. AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

    Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao +10

    cs.AIcs.LGarXiv:2608.05987v12026
  21. Recursive Synthesis for Long-Horizon Terminal Tasks

    Zhongzhi Li, Yucheng Shi, Zongxia Li +8

    cs.AIcs.LGarXiv:2608.05466v32026
  22. Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

    Junlin Han, Shengbang Tong, David Fan +4

    cs.CVcs.LGcs.MMarXiv:2608.05000v22026
  23. Gradient-Update Mismatch: Rethinking Conflict-Free Training of Physics-Informed Neural Networks

    Jing Xiao, Xinhai Chen, Qinglin Wang +5

    cs.LGarXiv:2609.01558v12026
  24. Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

    Yinghui He, Ling Yang, Jiarui Liu +6

    cs.CLcs.LGarXiv:2608.05139v12026
  25. Data Movement Is All You Need: A Case Study on Optimizing Transformers

    Andrei Ivanov, Nikoli Dryden, Tal Ben-Nun +2

    cs.LGstat.MLarXiv:2007.00072v32020
  26. RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States

    Yi Yang, Zhennan Chen, Yihong Zhuang +5

    cs.LGcs.CLarXiv:2608.02508v32026
  27. Verbalizable Representations Form a Global Workspace in Language Models

    Wes Gurnee, Nicholas Sofroniew, Adam Pearce +13

    cs.CLcs.AIcs.LGarXiv:2607.15495v12026
    Summaries:한국어
  28. RoboTTT: Context Scaling for Robot Policies

    Yunfan Jiang, Yevgen Chebotar, Ruijie Zheng +8

    cs.ROcs.AIcs.LGarXiv:2607.15275v12026
  29. BadWAM: When World-Action Models Dream Right but Act Wrong

    Qi Li, Xingyi Yang, Xinchao Wang

    cs.LGcs.ROarXiv:2607.15207v12026
    Summaries:한국어
  30. LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget

    Changhai Zhou, Kieran Liu, Yuhua Zhou +17

    cs.LGcs.DCarXiv:2607.14952v12026
  31. Self-Improvements in Modern Agentic Systems: A Survey

    Zhe Ren, Yimeng Chen, Dandan Guo +9

    cs.AIcs.CLcs.LGarXiv:2607.13104v12026
  32. Spurious Rewards: Rethinking Training Signals in RLVR

    Rulin Shao, Shuyue Stella Li, Rui Xin +11

    cs.AIcs.LGarXiv:2506.10947v22025
  33. xHC: Expanded Hyper-Connections

    Xiangdong Zhang, Xiaohan Qin, Sunan Zou +10

    cs.LGcs.CLarXiv:2607.14530v12026
    Summaries:한국어
  34. DeepLoop: Depth Scaling for Looped Transformers

    Shuzhen Li, Yifan Zhang, Jiacheng Guo +2

    cs.LGcs.AIarXiv:2607.13491v22026
  35. Contribution-Aware Bandwidth Allocation for Multimodal Split Learning

    Iason Ofeidis, Leandros Tassiulas

    cs.LGcs.DCcs.NIarXiv:2609.01406v12026
  36. EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

    Deyao Zhu, Xin Zhou, Shengling Qin +44

    cs.CLcs.LGarXiv:2607.05155v12026
  37. Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity

    Loïc Cabannes, Pierre-Emmanuel Mazaré, Gergely Szilvasy +6

    cs.LGarXiv:2607.07386v12026
  38. Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

    Zhenyu Hou, Yujiang Li, Jie Tang +1

    cs.LGcs.AIarXiv:2607.07508v12026
  39. advertorch v0.1: An Adversarial Robustness Toolbox based on PyTorch

    Gavin Weiguang Ding, Luyu Wang, Xiaomeng Jin

    cs.LGcs.CRcs.CVarXiv:1902.07623v12019
  40. Denser $\neq$ Better: Limits of On-Policy Self-Distillation for Continual Post-Training

    Meng Wang, Haohan Zhao, Wenzhuo Liu +7

    cs.LGcs.CLarXiv:2607.01763v12026
  41. Physics-Informed Machine Learning: A Survey on Problems, Methods and Applications

    Zhongkai Hao, Songming Liu, Yichi Zhang +4

    cs.LGcs.AIcs.CVarXiv:2211.08064v22022
  42. Training Language Models to Reason Efficiently

    Daman Arora, Andrea Zanette

    cs.LGcs.CLarXiv:2502.04463v42025
  43. Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs

    Fahd Seddik, Fatemeh Fard

    cs.CLcs.LGarXiv:2606.27378v12026
    Summaries:한국어
  44. The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

    Jing Liang, Hongyao Tang, Yi Ma +9

    cs.LGarXiv:2606.29526v12026
  45. Multi-Block Diffusion Language Models

    Yijie Jin, Jiajun Xu, Yuxuan Liu +8

    cs.LGcs.CLarXiv:2606.29215v22026
    Summaries:한국어
  46. Delayed Verification Destabilizes Multi-Agent LLM Belief: Instability Thresholds and Optimal Corrector Placement

    Igor Itkin

    cs.MAcs.CLcs.LGarXiv:2606.27409v12026
  47. When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models

    Josef Chen

    cs.AIcs.LGarXiv:2606.27288v12026
  48. Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It

    Yupu Hao, Zhuoran Jin, Huanxuan Liao +2

    cs.CLcs.LGarXiv:2606.26027v12026
  49. Autodata: An agentic data scientist to create high quality synthetic data

    Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu +12

    cs.AIcs.CLcs.LGarXiv:2606.25996v32026
    Summaries:한국어
  50. Autodata: An agentic data scientist to create high quality synthetic data

    Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu +12

    cs.AIcs.CLcs.LGarXiv:2606.25996v22026
    Summaries:한국어
  51. SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verification

    Swapnil Bhattacharyya, Mayank Baranwal

    cs.AIcs.LGmath.OCarXiv:2609.00728v12026
  52. Discretizing Reward Models

    Vijay Viswanathan, Shiqi Wang, Devamanyu Hazarika +4

    cs.LGarXiv:2606.21795v12026
  53. When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning?

    Xuanfei Ren, Tengyang Xie

    stat.MLcs.LGarXiv:2606.18531v12026
  54. Rethinking the Role of Efficient Attention in Hybrid Architectures

    Ziqing Qiao, Yinuo Xu, Chaojun Xiao +6

    cs.CLcs.LGarXiv:2606.15378v12026
  55. A Stationary (and Therefore Compatible) Representation is All You Need

    Niccolò Biondi, Federico Pernici, Simone Ricci +1

    cs.LGarXiv:2606.12488v12026
  56. AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

    Zhengxuan Wu, Aryaman Arora, Atticus Geiger +5

    cs.CLcs.AIcs.LGarXiv:2501.17148v32025
  57. Rethinking the Divergence Regularization in LLM RL

    Jiarui Yao, Xiangxin Zhou, Penghui Qi +3

    cs.LGarXiv:2606.09821v12026
  58. Evolving Agents in the Dark: Retrospective Harness Optimization via Self-Preference

    Wenbo Pan, Shujie Liu, Chin-Yew Lin +5

    cs.AIcs.CLcs.LGarXiv:2606.05922v22026
  59. Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning

    Atoosa Chegini, Soheil Feizi

    cs.CLcs.AIcs.LGarXiv:2606.01682v12026
  60. Discovering Cooperative Pipelines: Autoresearch for Sequential Social Dilemmas

    Víctor Gallego

    cs.MAcs.AIcs.LGarXiv:2605.30003v12026