Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
7,321 to 7,380 of 20,044
Continual Evolution Strategies in Control Tasks
Nicola Pitzalis, Eleni Nisioti, Antonio Carta +2
cs.NEcs.LGarXiv:2608.13600v12026The Loss Does Not See the Basis, but Adam Does
Devender Singh
cs.LGmath.OCstat.MLarXiv:2608.05136v12026Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop
Igor Itkin
cs.AIcond-mat.stat-mechcs.CLarXiv:2608.11215v12026MemSFT: Mitigating Alignment Tax with an External Parametric Memory
Jiarui Wang, Xiang Shi, Jiaqi Cao +8
cs.LGcs.CLarXiv:2607.25614v12026Kimi Linear: An Expressive, Efficient Attention Architecture
Kimi Team, Yu Zhang, Zongyu Lin +57
cs.CLcs.LGarXiv:2510.26692v22025ISO: An RLVR-Native Optimization Stack
Hanqing Zhu, Wenyan Cong, Zhizhou Sha +8
cs.LGcs.AIarXiv:2607.19331v12026Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code
Niels Mündler-Sasahara, Hristo Venev, Dawn Song +2
cs.PLcs.AIcs.LGarXiv:2607.13921v22026Summaries:한국어KronQ: LLM Quantization via Kronecker-Factored Hessian
Donghyun Lee, Yuhang Li, Ruokai Yin +1
cs.LGarXiv:2607.07964v22026Partially Correlated Verifier Cascades in LLM Harnesses: Concave Log-Odds, Polynomial Reliability, and Blind-Spot Ceilings
Jiangang Han
math.STcs.AIcs.LGarXiv:2607.13918v12026UI-MOPD: Multi-Platform On-Policy Distillation for Unified GUI Agents
Niu Lian, Tongbo Chen, Zhehao Yu +8
cs.CLcs.AIcs.CVarXiv:2607.04425v22026Trust Region Policy Distillation
Zhengpeng Xie, Li Lyna Zhang, Zeke Xie +1
cs.LGcs.AIarXiv:2607.04751v12026MANCE: Manifold Aware Concept Erasure
Matan Avitan, Yoav Goldberg, Yanai Elazar
cs.LGarXiv:2607.03973v12026Go-with-the-Track: Video Compositing and Motion Control with Point Tracking
Koichi Namekata, Yash Kant, Zhizheng Liu +9
cs.CVcs.LGarXiv:2606.20891v12026APPO: Agentic Procedural Policy Optimization
Xucong Wang, Ziyu Ma, Yong Wang +5
cs.LGcs.AIarXiv:2606.12384v22026Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning
Zhiyuan Zhou, Andy Peng, Charles Xu +4
cs.LGcs.AIarXiv:2606.11087v12026Redesign Mixture-of-Experts Routers with Manifold Power Iteration
Songhao Wu, Ang Lv, Ruobing Xie +1
cs.LGcs.AIcs.CLarXiv:2606.12397v12026Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning
Renjie Mao, Xiangxin Zhou, Lvfang Tao +7
cs.LGcs.AIarXiv:2606.10968v22026Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests
Thanawat Lodkaew, Johannes Ackermann, Soichiro Nishimori +3
cs.LGcs.AIcs.CLarXiv:2606.07379v22026Beyond Holistic Models: Systematic Component-level Benchmarking of Deep Multivariate Time-Series Forecasting
Shuang Liang, Chaochuan Hou, Xu Yao +4
cs.LGarXiv:2605.26562v12026OPRD: On-Policy Representation Distillation
Shenzhi Yang, Guangcheng Zhu, Bowen Song +8
cs.LGcs.AIarXiv:2606.06021v42026Why Muon Outperforms Adam: A Curvature Perspective
Shuche Wang, Fengzhuo Zhang, Jiaxiang Li +2
cs.LGcs.AIarXiv:2606.04662v12026Two-Fidelity Best-Action Identification for Stochastic Minimax Tree
Peter Chen, Xi Chen
cs.LGcs.AIarXiv:2606.01708v12026OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification
Yuhang Zhou, Lizhu Zhang, Yifan Wu +5
cs.LGcs.CLarXiv:2606.01476v22026MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems
Xinle Deng, Ruobin Zhong, Hujin Peng +15
cs.CLcs.AIcs.LGarXiv:2605.28732v32026Parallax: Parameterized Local Linear Attention for Language Modeling
Yifei Zuo, Dhruv Pai, Zhichen Zeng +3
cs.LGcs.AIcs.CLarXiv:2605.29157v12026GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models
Xiaohang Tang, Keyue Jiang, Che Liu +4
cs.LGcs.AIarXiv:2605.29398v12026The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction
Shu Wan, Abhinav Gorantla, Huan Liu +1
cs.LGcs.AIstat.MEarXiv:2605.29411v22026Automatic Construction and Natural-Language Description of Nonparametric Regression Models
James Robert Lloyd, David Duvenaud, Roger Grosse +2
stat.MLcs.LGarXiv:1402.4304v32014RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains
Haoxiang Jiang, Zihan Dong, Tianci Liu +5
cs.LGcs.CLarXiv:2605.29156v12026Trust Region Q Adjoint Matching
Yonghoon Dong, Kyungmin Lee, Changyeon Kim +2
cs.LGcs.AIcs.ROarXiv:2605.27079v12026Pruning and Distilling Mixture-of-Experts into Dense Language Models
Junhyuck Kim, Jihun Yun, Haechan Kim +3
cs.CLcs.AIcs.LGarXiv:2605.28207v22026Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation
Yuanyi Wang, Su Lu, Yanggan Gu +6
cs.LGarXiv:2605.26844v12026Contrastive Distribution Matching for Amortized Sequential Monte Carlo in Discrete Diffusion
Jaihoon Kim, Taehoon Yoon, Prin Phunyaphibarn +3
cs.LGarXiv:2605.23346v12026When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs
Yifan Zeng, Yiran Wu, Yaolun Zhang +4
cs.AIcs.LGarXiv:2605.24202v22026Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
Tim Tsz-Kit Lau, Weijie Su
math.OCcs.AIcs.LGarXiv:2605.18106v42026Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws
Nandan Kumar Jha, Brandon Reagen
cs.LGarXiv:2605.21803v12026The Distillation Game: Adaptive Attacks & Efficient Defenses
Youssef Allouah, Mahdi Haghifam, Sanmi Koyejo +1
cs.LGcs.AIarXiv:2605.22737v32026OCTOPUS: Optimized KV Cache for Transformers via Octahedral Parametrization Under optimal Squared error quantization
Mark Boss, Vikram Voleti, Simon Donné +1
cs.LGcs.AIarXiv:2605.21226v12026optimize_anything: A Universal API for Optimizing any Text Parameter
Lakshya A Agrawal, Donghyun Lee, Shangyin Tan +11
cs.CLcs.AIcs.LGarXiv:2605.19633v12026CM-EVS: Sparse Panoramic RGB-D-Pose Data for Complete Scene Coverage
Jiale Liu, Jungang Li, Jieming Yu +13
cs.CVcs.GRcs.LGarXiv:2605.15597v12026Forgetting That Sticks: Quantization-Permanent Unlearning via Circuit Attribution
Saisab Sadhu, Pratinav Seth, Vinay Kumar Sankarapu
cs.LGcs.CLcs.ETarXiv:2605.15138v12026HodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-Experts
Tao Zhong, Dongzhe Zheng, Christine Allen-Blanchette
cs.LGcs.AIcs.CLarXiv:2605.13997v12026Holder Policy Optimisation
Yuxiang Chen, Dingli Liang, Yihang Chen +8
cs.LGcs.AIarXiv:2605.12058v22026Reliable Chain-of-Thought via Prefix Consistency
Naoto Iwase, Yuki Ichihara, Mohammad Atif Quamar +1
stat.MLcs.CLcs.LGarXiv:2605.07654v12026Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training
Yuanyi Wang, Yifan Yang, Su Lu +9
cs.LGcs.ITarXiv:2605.09608v12026G-Zero: Self-Play for Open-Ended Generation from Zero Data
Chengsong Huang, Haolin Liu, Tong Zheng +7
cs.LGcs.AIcs.CLarXiv:2605.09959v12026Multi-Step-Ahead Time Series Prediction using Multiple-Output Support Vector Regression
Yukun Bao, Tao Xiong, Zhongyi Hu
cs.LGstat.MLarXiv:1401.2504v12014Sub-JEPA: Subspace Gaussian Regularization for Stable End-to-End World Models
Kai Zhao, Dongliang Nie, Yuchen Lin +4
cs.LGcs.AIarXiv:2605.09241v12026Rethinking State Tracking in Recurrent Models Through Error Control Dynamics
Jiwan Chung, Heechan Choi, Seon Joo Kim
cs.LGcs.CLarXiv:2605.07755v12026Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers
Pengqi Lu
cs.LGcs.CVarXiv:2605.06169v12026Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control
Bolian Li, Yifan Wang, Yi Ding +3
cs.LGcs.CLstat.MLarXiv:2604.26326v22026Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis
Zhisong Qiu, Shuofei Qiao, Kewei Xu +4
cs.CLcs.AIcs.CEarXiv:2604.24198v22026Sample Selection Using Multi-Task Autoencoders in Federated Learning with Non-IID Data
Emre Ardıç, Yakup Genç
cs.CVcs.LGarXiv:2604.26116v12026Convergent Evolution: How Different Language Models Learn Similar Number Representations
Deqing Fu, Tianyi Zhou, Mikhail Belkin +2
cs.CLcs.AIcs.LGarXiv:2604.20817v22026DiPO: Disentangled Perplexity Policy Optimization for Fine-grained Exploration-Exploitation Trade-Off
Xiaofan Li, Ming Yang, Zhiyuan Ma +9
cs.LGarXiv:2604.13902v12026Reinforcement Learning via Value Gradient Flow
Haoran Xu, Kaiwen Hu, Somayeh Sojoudi +1
cs.LGcs.AIarXiv:2604.14265v12026Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
Yecheng Wu, Song Han, Hai Cai
cs.LGcs.AIarXiv:2604.13010v22026Multi-Domain Riemannian Graph Gluing for Building Graph Foundation Models
Li Sun, Zhenhao Huang, Silei Chen +4
cs.LGarXiv:2603.00618v12026Steered LLM Activations are Non-Surjective
Aayush Mishra, Daniel Khashabi, Anqi Liu
cs.AIcs.LGarXiv:2604.09839v22026Tunable Soft Equivariance with Guarantees
Md Ashiqur Rahman, Lim Jun Hao, Jeremiah Jiang +2
cs.CVcs.LGarXiv:2603.26657v12026