Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
7,081 to 7,140 of 20,199
GPG: A Simple and Strong Reinforcement Learning Baseline for Model Reasoning
Xiangxiang Chu, Hailang Huang, Xiao Zhang +2
cs.LGcs.AIarXiv:2504.02546v42025Beyond Local Power: Functional Connectivity Analysis for Subject-Independent Learning Style Recognition
Wiga Maulana Baihaqi, Indriana Hidayah, Sri Kusrohmaniah +1
q-bio.NCcs.LGeess.SParXiv:2608.12000v12026AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset
Ivan Moshkov, Darragh Hanley, Ivan Sorokin +5
cs.AIcs.CLcs.LGarXiv:2504.16891v12025ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
Bill Yuchen Lin, Ronan Le Bras, Kyle Richardson +4
cs.AIcs.CLcs.LGarXiv:2502.01100v22025Dion3: Full-Stack Orthogonal Updates
Noah Amsel, Jack Zhang, Kwangjun Ahn +5
cs.LGcs.AIarXiv:2608.11612v12026InSight-doc: Agentic Visual Perception for Long-Document Understanding
Kaican Li, Weiyan Xie, Lewei Yao +4
cs.CVcs.CLcs.LGarXiv:2608.10628v12026A deep learning model for estimating story points
Morakot Choetkiertikul, Hoa Khanh Dam, Truyen Tran +3
cs.SEcs.LGstat.MLarXiv:1609.00489v22016Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networks
Ziwei Ji, Matus Telgarsky
cs.LGmath.OCstat.MLarXiv:1909.12292v42019Network Randomization: A Simple Technique for Generalization in Deep Reinforcement Learning
Kimin Lee, Kibok Lee, Jinwoo Shin +1
cs.LGstat.MLarXiv:1910.05396v32019Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure
Víctor Gallego
cs.LGcs.AIarXiv:2608.08722v12026Summaries:한국어CEM-RL: Combining evolutionary and gradient-based methods for policy search
Aloïs Pourchot, Olivier Sigaud
cs.LGcs.NEstat.MLarXiv:1810.01222v32018A Hybrid Nested Harness for Decoupling Structure and Parameters in LLM-Driven Optimization
Víctor Gallego
cs.LGcs.NEarXiv:2608.08156v12026Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution
Changzhi Liu, Yilun Liu, Sikuan Yan +2
cs.AIcs.LGarXiv:2608.07645v12026Finite-Sample Metric Non-Collapse for Geometrically Supervised Latent World Models in Control
Alain Bensoussan, Minh-Nhat Phung, Minh-Binh Tran
math.OCcs.LGarXiv:2608.07265v22026End-to-End Lane Marker Detection via Row-wise Classification
Seungwoo Yoo, Heeseok Lee, Heesoo Myeong +4
cs.CVcs.LGarXiv:2005.08630v12020On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification
Yongliang Wu, Yizhou Zhou, Zhou Ziheng +7
cs.LGarXiv:2508.05629v32025Recursive Harness Self-Improvement
Hyunin Lee, Jinglue Xu, Jeffrey Seely +3
cs.LGcs.AIarXiv:2607.15524v12026Summaries:한국어ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation
Qingyu Zhang, Qianhao Yuan, Hongyu Lin +5
cs.LGcs.AIcs.CLarXiv:2607.13124v22026Summaries:한국어Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning
Chen Tang, Yizhou Wang, Jianyu Wu +26
cs.CLcs.AIcs.CEarXiv:2607.07708v12026Summaries:한국어AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation
Andrey Podivilov, Vadim Lomshakov, Sergey Savin +4
cs.AIcs.LGcs.SEarXiv:2607.06624v22026Multiplayer Interactive World Models with Representation Autoencoders
Anthony Hu, Václav Volhejn, Adrien Ramanana Rahary +24
cs.CVcs.AIcs.LGarXiv:2607.05352v22026TESSERA v2: Scaling Pixel-wise Earth Foundation Models
Zhengpeng Feng, Sadiq Jaffer, Ira Shokar +12
cs.CVcs.LGarXiv:2607.03949v22026OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers
Donghyun Lee, Jitesh Chavan, Duy Nguyen +5
cs.CVcs.AIcs.LGarXiv:2607.02461v12026Discrete Diffusion Language Models for Interactive Radiology Report Drafting
Max Van Puyvelde, Halil Ibrahim Gulluk, Wim Van Criekinge +1
cs.AIcs.LGarXiv:2607.01436v12026Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads
Aryo Pradipta Gema, Beatrice Alex, Pasquale Minervini
cs.CLcs.AIcs.LGarXiv:2607.01002v12026QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents
Sergio Hernández-Gutiérrez, Matteo Merler, Ilze Amanda Auzina +3
cs.LGcs.AIcs.CLarXiv:2606.32034v12026TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning
Yuanda Xu, Zhengze Zhou, Hejian Sang +6
cs.LGcs.AIarXiv:2606.32017v22026SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions
Mohit Raghavendra, Anisha Gunjal, Aakash Sabharwal +1
cs.LGarXiv:2606.30573v12026$μ_0$: A Scalable 3D Interaction-Trace World Model
Seungjae Lee, Yoonkyo Jung, Jusuk Lee +6
cs.ROcs.CVcs.LGarXiv:2606.13769v22026Entropy as a Structural Prior: How a Log-Barrier on DiT Belief Space Drives Musical Diversity and Development
Zixi Li, Youzhen Li
cs.SDcs.LGeess.ASarXiv:2606.07207v12026Regret Minimization with Adaptive Opponents in Repeated Games
Mingyang Liu, Asuman Ozdaglar, Tiancheng Yu +1
cs.LGcs.AIcs.GTarXiv:2606.06486v12026Reinforcement Learning from Rich Feedback with Distributional DAgger
Rishabh Agrawal, Jacob Fein-Ashley, Paria Rashidinejad
cs.LGcs.AIcs.CLarXiv:2606.05152v22026Unlocking Feature Learning in Gated Delta Networks at Scale
Yifeng Liu, Quanquan Gu
cs.LGcs.AIarXiv:2606.04048v12026Self-Distilled Policy Gradient
Yifeng Liu, Shiyuan Zhang, Yifan Zhang +1
cs.LGarXiv:2606.04036v12026Neural Networks Provably Learn Spectral Representations for Group Composition
Jianliang He, Leda Wang, Fengzhuo Zhang +2
cs.LGmath.OCmath.RTarXiv:2606.02993v22026Unified Neural Scaling Laws
Ethan Caballero, Priyank Jaini, David Krueger +1
cs.LGcs.AIcs.NEarXiv:2605.26248v12026Uniform Diffusion Models Revisited: Leave-One-Out Denoiser and Absorbing State Reformulation
Samson Gourevitch, Yazid Janati, Dario Shariatian +4
cs.LGstat.MLarXiv:2605.22765v12026Efficient Agentic Reasoning Through Self-Regulated Simulative Planning
Mingkai Deng, Jinyu Hou, Lara Sá Neves +4
cs.AIcs.CLcs.LGarXiv:2605.22138v12026From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoning
Xitai Jiang, Zihan Tang, Wenze Lin +3
cs.LGcs.AIcs.CLarXiv:2605.22074v12026SAGA: A Sequence-Adaptive Generative Architecture for Multi-Horizon Probabilistic Forecasting with Adaptive Temporal Conformal Prediction
Gustav Olaf Yunus Laitinen-Fredriksson Lundström-Imanov, Hafize Gonca Cömert
cs.LGecon.EMstat.MLarXiv:2605.19014v12026DynMuon: A Dynamic Spectral Shaping View of Muon
Fangzhou Wu, Rikhav Shah, Sandeep Silwal +1
cs.LGcs.AIarXiv:2605.17109v32026FutureSim: Replaying World Events to Evaluate Adaptive Agents
Shashwat Goel, Nikhil Chandak, Arvindh Arun +5
cs.LGcs.AIcs.CLarXiv:2605.15188v12026MinT: Managed Infrastructure for Training and Serving Millions of LLMs
Mind Lab, :, Song Cao +60
cs.LGcs.AIcs.DCarXiv:2605.13779v22026The Extrapolation Cliff in On-Policy Distillation of Near-Deterministic Structured Outputs
Xin Li, Hao Jiang, Annan Wang +2
cs.LGcs.CLarXiv:2605.08737v120263DTV: A Feedforward Interpolation Network for Real-Time View Synthesis
Stefan Schulz, Fernando Edelstein, Hannah Dröge +2
cs.CVcs.LGcs.MMarXiv:2604.11211v12026Streaming Structured Inference with Flash-SemiCRF
Benjamin K. Johnson, Thomas Goralski, Ayush Semwal +2
cs.LGarXiv:2604.18780v12026Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer
Rafał Powalski, Łukasz Borchmann, Dawid Jurkiewicz +3
cs.CLcs.LGarXiv:2102.09550v32021From Truncation to Commitment: Persistent Context in Uniform Discrete Diffusion
Satoshi Hayakawa
cs.LGcs.AImath.PRarXiv:2609.01043v12026Not All Layers Are Created Equal: Adaptive LoRA Ranks for Personalized Image Generation
Donald Shenaj, Federico Errica, Antonio Carta
cs.CVcs.AIcs.LGarXiv:2603.21884v12026I Know What I Don't Know: Latent Posterior Factor Models for Multi-Evidence Probabilistic Reasoning
Aliyu Agboola Alege
cs.AIcs.LGarXiv:2603.15670v22026Beyond Single Tokens: Distilling Discrete Diffusion Models via Discrete MMD
Emiel Hoogeboom, David Ruhe, Jonathan Heek +2
cs.LGcs.CVstat.MLarXiv:2603.20155v12026Remasking Discrete Diffusion Models with Inference-Time Scaling
Guanghan Wang, Yair Schiff, Subham Sekhar Sahoo +1
cs.LGstat.MLarXiv:2503.00307v42025SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis
Shahriar Noroozizadeh, Xiaobin Shen, Jeremy C. Weiss +1
cs.LGcs.AIstat.MLarXiv:2603.05483v12026MedMentions: A Large Biomedical Corpus Annotated with UMLS Concepts
Sunil Mohan, Donghui Li
cs.CLcs.LGarXiv:1902.09476v12019Surprised by Attention: Predictable Query Dynamics for Time Series Anomaly Detection
Kadir-Kaan Özer, René Ebeling, Markus Enzweiler
cs.LGcs.AIarXiv:2603.12916v32026Denoising Diffusion Generative Models Secretly Calculate Attentions
Farzan Haddadi, Leila Monfared, Ebrahim Rezaii +3
cs.AIcs.CVcs.LGarXiv:2609.00885v12026Model-Based Reinforcement Learning with a Generative Model is Minimax Optimal
Alekh Agarwal, Sham Kakade, Lin F. Yang
cs.LGmath.PRstat.MLarXiv:1906.03804v32019Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
Arnas Uselis, Andrea Dittadi, Seong Joon Oh
cs.CVcs.LGarXiv:2602.24264v22026Transform-Invariant Generative Ray Path Sampling for Efficient Radio Propagation Modeling
Jérome Eertmans, Enrico M. Vitucci, Vittorio Degli-Esposti +3
cs.LGeess.SParXiv:2603.01655v22026On the Mechanism and Dynamics of Modular Addition: Fourier Features, Lottery Ticket, and Grokking
Jianliang He, Leda Wang, Siyu Chen +1
cs.LGmath.OCstat.MLarXiv:2602.16849v12026