Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
7,861 to 7,920 of 20,199
MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning
Zirui Cheng, Xun Xu, Tiankai Chen +7
cs.LGarXiv:2608.12724v12026Small-Scale Experiments: Are We There Yet?
Nicholas Lourie, Kyunghyun Cho, Karen Ullrich +1
cs.LGarXiv:2608.11859v12026Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing
Tianci Liu, Zihan Dong, Tianchun Li +8
cs.CLcs.AIcs.LGarXiv:2608.11660v12026UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations
Dvir Samuel, Guy Bar-Shalom, Fabrizio Frasca +4
cs.CVcs.LGarXiv:2608.10835v12026Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference
Burc Gokden
cs.LGcs.CLarXiv:2608.10288v12026Multimodal Model Diffing for Feature Discovery and Control
Hunar Batra, Lachin Naghashyar, Ashkan Khakzar +4
cs.CVcs.AIcs.CLarXiv:2608.09928v12026Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
Mind Lab, :, Vin Bo +74
cs.LGcs.CLarXiv:2608.09819v12026RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance
Dongchi Huang, Hongyin Zhang, Bohan Hou +12
cs.ROcs.CVcs.LGarXiv:2608.09853v12026BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
Björn Engdahl, Adrian Kosowski, Jan Chorowski +6
cs.NEcs.AIcs.LGarXiv:2608.09888v12026Learning Locomotion Skills Using DeepRL: Does the Choice of Action Space Matter?
Xue Bin Peng, Michiel van de Panne
cs.LGcs.GRcs.ROarXiv:1611.01055v12016Parameter Exploration for RLVR via Variational Learning
Vatsal Venkatkrishna, Nico Daheim, Iryna Gurevych
cs.LGcs.AIcs.CLarXiv:2608.09805v12026Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
NVIDIA, :, Yan Wang +41
cs.ROcs.AIcs.LGarXiv:2511.00088v22025OpenThoughts: Data Recipes for Reasoning Models
Etash Guha, Ryan Marten, Sedrick Keh +47
cs.LGarXiv:2506.04178v22025Direct Optimization of a 3D Finite-Source Reflector via Neural-Network Parameterization
Roel Hacking, Lisa Kusch, Martijn Anthonissen +1
physics.opticscs.LGarXiv:2609.00899v12026Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
Taeil Kim, Kangsan Kim, Sung Ju Hwang
cs.AIcs.LGarXiv:2608.07169v12026CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks
Fanzhe Meng, Guoxin Chen, Jiale Zhao +6
cs.LGcs.CLarXiv:2608.06352v12026Kimi K3: Open Frontier Intelligence
Kimi Team, Tongtong Bai, Yifan Bai +399
cs.CLcs.LGarXiv:2607.24653v22026EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
Ryan Hoque, Peide Huang, David J. Yoon +2
cs.CVcs.LGcs.ROarXiv:2505.11709v32025From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models
Jiale Han, Xiang Li, Jing Qian +7
cs.AIcs.LGarXiv:2608.06020v12026AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao +10
cs.AIcs.LGarXiv:2608.05987v12026Recursive Synthesis for Long-Horizon Terminal Tasks
Zhongzhi Li, Yucheng Shi, Zongxia Li +8
cs.AIcs.LGarXiv:2608.05466v32026Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
Junlin Han, Shengbang Tong, David Fan +4
cs.CVcs.LGcs.MMarXiv:2608.05000v22026Gradient-Update Mismatch: Rethinking Conflict-Free Training of Physics-Informed Neural Networks
Jing Xiao, Xinhai Chen, Qinglin Wang +5
cs.LGarXiv:2609.01558v12026Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning
Yinghui He, Ling Yang, Jiarui Liu +6
cs.CLcs.LGarXiv:2608.05139v12026Data Movement Is All You Need: A Case Study on Optimizing Transformers
Andrei Ivanov, Nikoli Dryden, Tal Ben-Nun +2
cs.LGstat.MLarXiv:2007.00072v32020RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States
Yi Yang, Zhennan Chen, Yihong Zhuang +5
cs.LGcs.CLarXiv:2608.02508v32026Verbalizable Representations Form a Global Workspace in Language Models
Wes Gurnee, Nicholas Sofroniew, Adam Pearce +13
cs.CLcs.AIcs.LGarXiv:2607.15495v12026Summaries:한국어RoboTTT: Context Scaling for Robot Policies
Yunfan Jiang, Yevgen Chebotar, Ruijie Zheng +8
cs.ROcs.AIcs.LGarXiv:2607.15275v12026BadWAM: When World-Action Models Dream Right but Act Wrong
Qi Li, Xingyi Yang, Xinchao Wang
cs.LGcs.ROarXiv:2607.15207v12026Summaries:한국어LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
Changhai Zhou, Kieran Liu, Yuhua Zhou +17
cs.LGcs.DCarXiv:2607.14952v12026Self-Improvements in Modern Agentic Systems: A Survey
Zhe Ren, Yimeng Chen, Dandan Guo +9
cs.AIcs.CLcs.LGarXiv:2607.13104v12026Spurious Rewards: Rethinking Training Signals in RLVR
Rulin Shao, Shuyue Stella Li, Rui Xin +11
cs.AIcs.LGarXiv:2506.10947v22025xHC: Expanded Hyper-Connections
Xiangdong Zhang, Xiaohan Qin, Sunan Zou +10
cs.LGcs.CLarXiv:2607.14530v12026Summaries:한국어DeepLoop: Depth Scaling for Looped Transformers
Shuzhen Li, Yifan Zhang, Jiacheng Guo +2
cs.LGcs.AIarXiv:2607.13491v22026Contribution-Aware Bandwidth Allocation for Multimodal Split Learning
Iason Ofeidis, Leandros Tassiulas
cs.LGcs.DCcs.NIarXiv:2609.01406v12026EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments
Deyao Zhu, Xin Zhou, Shengling Qin +44
cs.CLcs.LGarXiv:2607.05155v12026Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity
Loïc Cabannes, Pierre-Emmanuel Mazaré, Gergely Szilvasy +6
cs.LGarXiv:2607.07386v12026Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
Zhenyu Hou, Yujiang Li, Jie Tang +1
cs.LGcs.AIarXiv:2607.07508v12026advertorch v0.1: An Adversarial Robustness Toolbox based on PyTorch
Gavin Weiguang Ding, Luyu Wang, Xiaomeng Jin
cs.LGcs.CRcs.CVarXiv:1902.07623v12019Denser $\neq$ Better: Limits of On-Policy Self-Distillation for Continual Post-Training
Meng Wang, Haohan Zhao, Wenzhuo Liu +7
cs.LGcs.CLarXiv:2607.01763v12026Physics-Informed Machine Learning: A Survey on Problems, Methods and Applications
Zhongkai Hao, Songming Liu, Yichi Zhang +4
cs.LGcs.AIcs.CVarXiv:2211.08064v22022Training Language Models to Reason Efficiently
Daman Arora, Andrea Zanette
cs.LGcs.CLarXiv:2502.04463v42025Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs
Fahd Seddik, Fatemeh Fard
cs.CLcs.LGarXiv:2606.27378v12026Summaries:한국어The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning
Jing Liang, Hongyao Tang, Yi Ma +9
cs.LGarXiv:2606.29526v12026Multi-Block Diffusion Language Models
Yijie Jin, Jiajun Xu, Yuxuan Liu +8
cs.LGcs.CLarXiv:2606.29215v22026Summaries:한국어Delayed Verification Destabilizes Multi-Agent LLM Belief: Instability Thresholds and Optimal Corrector Placement
Igor Itkin
cs.MAcs.CLcs.LGarXiv:2606.27409v12026When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models
Josef Chen
cs.AIcs.LGarXiv:2606.27288v12026Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It
Yupu Hao, Zhuoran Jin, Huanxuan Liao +2
cs.CLcs.LGarXiv:2606.26027v12026Autodata: An agentic data scientist to create high quality synthetic data
Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu +12
cs.AIcs.CLcs.LGarXiv:2606.25996v32026Summaries:한국어Autodata: An agentic data scientist to create high quality synthetic data
Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu +12
cs.AIcs.CLcs.LGarXiv:2606.25996v22026Summaries:한국어SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verification
Swapnil Bhattacharyya, Mayank Baranwal
cs.AIcs.LGmath.OCarXiv:2609.00728v12026Discretizing Reward Models
Vijay Viswanathan, Shiqi Wang, Devamanyu Hazarika +4
cs.LGarXiv:2606.21795v12026When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning?
Xuanfei Ren, Tengyang Xie
stat.MLcs.LGarXiv:2606.18531v12026Rethinking the Role of Efficient Attention in Hybrid Architectures
Ziqing Qiao, Yinuo Xu, Chaojun Xiao +6
cs.CLcs.LGarXiv:2606.15378v12026A Stationary (and Therefore Compatible) Representation is All You Need
Niccolò Biondi, Federico Pernici, Simone Ricci +1
cs.LGarXiv:2606.12488v12026AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
Zhengxuan Wu, Aryaman Arora, Atticus Geiger +5
cs.CLcs.AIcs.LGarXiv:2501.17148v32025Rethinking the Divergence Regularization in LLM RL
Jiarui Yao, Xiangxin Zhou, Penghui Qi +3
cs.LGarXiv:2606.09821v12026Evolving Agents in the Dark: Retrospective Harness Optimization via Self-Preference
Wenbo Pan, Shujie Liu, Chin-Yew Lin +5
cs.AIcs.CLcs.LGarXiv:2606.05922v22026Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning
Atoosa Chegini, Soheil Feizi
cs.CLcs.AIcs.LGarXiv:2606.01682v12026Discovering Cooperative Pipelines: Autoresearch for Sequential Social Dilemmas
Víctor Gallego
cs.MAcs.AIcs.LGarXiv:2605.30003v12026