Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

7,081 to 7,140 of 20,199

  1. GPG: A Simple and Strong Reinforcement Learning Baseline for Model Reasoning

    Xiangxiang Chu, Hailang Huang, Xiao Zhang +2

    cs.LGcs.AIarXiv:2504.02546v42025
  2. Beyond Local Power: Functional Connectivity Analysis for Subject-Independent Learning Style Recognition

    Wiga Maulana Baihaqi, Indriana Hidayah, Sri Kusrohmaniah +1

    q-bio.NCcs.LGeess.SParXiv:2608.12000v12026
  3. AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset

    Ivan Moshkov, Darragh Hanley, Ivan Sorokin +5

    cs.AIcs.CLcs.LGarXiv:2504.16891v12025
  4. ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

    Bill Yuchen Lin, Ronan Le Bras, Kyle Richardson +4

    cs.AIcs.CLcs.LGarXiv:2502.01100v22025
  5. Dion3: Full-Stack Orthogonal Updates

    Noah Amsel, Jack Zhang, Kwangjun Ahn +5

    cs.LGcs.AIarXiv:2608.11612v12026
  6. InSight-doc: Agentic Visual Perception for Long-Document Understanding

    Kaican Li, Weiyan Xie, Lewei Yao +4

    cs.CVcs.CLcs.LGarXiv:2608.10628v12026
  7. A deep learning model for estimating story points

    Morakot Choetkiertikul, Hoa Khanh Dam, Truyen Tran +3

    cs.SEcs.LGstat.MLarXiv:1609.00489v22016
  8. Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networks

    Ziwei Ji, Matus Telgarsky

    cs.LGmath.OCstat.MLarXiv:1909.12292v42019
  9. Network Randomization: A Simple Technique for Generalization in Deep Reinforcement Learning

    Kimin Lee, Kibok Lee, Jinwoo Shin +1

    cs.LGstat.MLarXiv:1910.05396v32019
  10. Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure

    Víctor Gallego

    cs.LGcs.AIarXiv:2608.08722v12026
    Summaries:한국어
  11. CEM-RL: Combining evolutionary and gradient-based methods for policy search

    Aloïs Pourchot, Olivier Sigaud

    cs.LGcs.NEstat.MLarXiv:1810.01222v32018
  12. A Hybrid Nested Harness for Decoupling Structure and Parameters in LLM-Driven Optimization

    Víctor Gallego

    cs.LGcs.NEarXiv:2608.08156v12026
  13. Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution

    Changzhi Liu, Yilun Liu, Sikuan Yan +2

    cs.AIcs.LGarXiv:2608.07645v12026
  14. Finite-Sample Metric Non-Collapse for Geometrically Supervised Latent World Models in Control

    Alain Bensoussan, Minh-Nhat Phung, Minh-Binh Tran

    math.OCcs.LGarXiv:2608.07265v22026
  15. End-to-End Lane Marker Detection via Row-wise Classification

    Seungwoo Yoo, Heeseok Lee, Heesoo Myeong +4

    cs.CVcs.LGarXiv:2005.08630v12020
  16. On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification

    Yongliang Wu, Yizhou Zhou, Zhou Ziheng +7

    cs.LGarXiv:2508.05629v32025
  17. Recursive Harness Self-Improvement

    Hyunin Lee, Jinglue Xu, Jeffrey Seely +3

    cs.LGcs.AIarXiv:2607.15524v12026
    Summaries:한국어
  18. ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

    Qingyu Zhang, Qianhao Yuan, Hongyu Lin +5

    cs.LGcs.AIcs.CLarXiv:2607.13124v22026
    Summaries:한국어
  19. Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning

    Chen Tang, Yizhou Wang, Jianyu Wu +26

    cs.CLcs.AIcs.CEarXiv:2607.07708v12026
    Summaries:한국어
  20. AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

    Andrey Podivilov, Vadim Lomshakov, Sergey Savin +4

    cs.AIcs.LGcs.SEarXiv:2607.06624v22026
  21. Multiplayer Interactive World Models with Representation Autoencoders

    Anthony Hu, Václav Volhejn, Adrien Ramanana Rahary +24

    cs.CVcs.AIcs.LGarXiv:2607.05352v22026
  22. TESSERA v2: Scaling Pixel-wise Earth Foundation Models

    Zhengpeng Feng, Sadiq Jaffer, Ira Shokar +12

    cs.CVcs.LGarXiv:2607.03949v22026
  23. OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers

    Donghyun Lee, Jitesh Chavan, Duy Nguyen +5

    cs.CVcs.AIcs.LGarXiv:2607.02461v12026
  24. Discrete Diffusion Language Models for Interactive Radiology Report Drafting

    Max Van Puyvelde, Halil Ibrahim Gulluk, Wim Van Criekinge +1

    cs.AIcs.LGarXiv:2607.01436v12026
  25. Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads

    Aryo Pradipta Gema, Beatrice Alex, Pasquale Minervini

    cs.CLcs.AIcs.LGarXiv:2607.01002v12026
  26. QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents

    Sergio Hernández-Gutiérrez, Matteo Merler, Ilze Amanda Auzina +3

    cs.LGcs.AIcs.CLarXiv:2606.32034v12026
  27. TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning

    Yuanda Xu, Zhengze Zhou, Hejian Sang +6

    cs.LGcs.AIarXiv:2606.32017v22026
  28. SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions

    Mohit Raghavendra, Anisha Gunjal, Aakash Sabharwal +1

    cs.LGarXiv:2606.30573v12026
  29. $μ_0$: A Scalable 3D Interaction-Trace World Model

    Seungjae Lee, Yoonkyo Jung, Jusuk Lee +6

    cs.ROcs.CVcs.LGarXiv:2606.13769v22026
  30. Entropy as a Structural Prior: How a Log-Barrier on DiT Belief Space Drives Musical Diversity and Development

    Zixi Li, Youzhen Li

    cs.SDcs.LGeess.ASarXiv:2606.07207v12026
  31. Regret Minimization with Adaptive Opponents in Repeated Games

    Mingyang Liu, Asuman Ozdaglar, Tiancheng Yu +1

    cs.LGcs.AIcs.GTarXiv:2606.06486v12026
  32. Reinforcement Learning from Rich Feedback with Distributional DAgger

    Rishabh Agrawal, Jacob Fein-Ashley, Paria Rashidinejad

    cs.LGcs.AIcs.CLarXiv:2606.05152v22026
  33. Unlocking Feature Learning in Gated Delta Networks at Scale

    Yifeng Liu, Quanquan Gu

    cs.LGcs.AIarXiv:2606.04048v12026
  34. Self-Distilled Policy Gradient

    Yifeng Liu, Shiyuan Zhang, Yifan Zhang +1

    cs.LGarXiv:2606.04036v12026
  35. Neural Networks Provably Learn Spectral Representations for Group Composition

    Jianliang He, Leda Wang, Fengzhuo Zhang +2

    cs.LGmath.OCmath.RTarXiv:2606.02993v22026
  36. Unified Neural Scaling Laws

    Ethan Caballero, Priyank Jaini, David Krueger +1

    cs.LGcs.AIcs.NEarXiv:2605.26248v12026
  37. Uniform Diffusion Models Revisited: Leave-One-Out Denoiser and Absorbing State Reformulation

    Samson Gourevitch, Yazid Janati, Dario Shariatian +4

    cs.LGstat.MLarXiv:2605.22765v12026
  38. Efficient Agentic Reasoning Through Self-Regulated Simulative Planning

    Mingkai Deng, Jinyu Hou, Lara Sá Neves +4

    cs.AIcs.CLcs.LGarXiv:2605.22138v12026
  39. From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoning

    Xitai Jiang, Zihan Tang, Wenze Lin +3

    cs.LGcs.AIcs.CLarXiv:2605.22074v12026
  40. SAGA: A Sequence-Adaptive Generative Architecture for Multi-Horizon Probabilistic Forecasting with Adaptive Temporal Conformal Prediction

    Gustav Olaf Yunus Laitinen-Fredriksson Lundström-Imanov, Hafize Gonca Cömert

    cs.LGecon.EMstat.MLarXiv:2605.19014v12026
  41. DynMuon: A Dynamic Spectral Shaping View of Muon

    Fangzhou Wu, Rikhav Shah, Sandeep Silwal +1

    cs.LGcs.AIarXiv:2605.17109v32026
  42. FutureSim: Replaying World Events to Evaluate Adaptive Agents

    Shashwat Goel, Nikhil Chandak, Arvindh Arun +5

    cs.LGcs.AIcs.CLarXiv:2605.15188v12026
  43. MinT: Managed Infrastructure for Training and Serving Millions of LLMs

    Mind Lab, :, Song Cao +60

    cs.LGcs.AIcs.DCarXiv:2605.13779v22026
  44. The Extrapolation Cliff in On-Policy Distillation of Near-Deterministic Structured Outputs

    Xin Li, Hao Jiang, Annan Wang +2

    cs.LGcs.CLarXiv:2605.08737v12026
  45. 3DTV: A Feedforward Interpolation Network for Real-Time View Synthesis

    Stefan Schulz, Fernando Edelstein, Hannah Dröge +2

    cs.CVcs.LGcs.MMarXiv:2604.11211v12026
  46. Streaming Structured Inference with Flash-SemiCRF

    Benjamin K. Johnson, Thomas Goralski, Ayush Semwal +2

    cs.LGarXiv:2604.18780v12026
  47. Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer

    Rafał Powalski, Łukasz Borchmann, Dawid Jurkiewicz +3

    cs.CLcs.LGarXiv:2102.09550v32021
  48. From Truncation to Commitment: Persistent Context in Uniform Discrete Diffusion

    Satoshi Hayakawa

    cs.LGcs.AImath.PRarXiv:2609.01043v12026
  49. Not All Layers Are Created Equal: Adaptive LoRA Ranks for Personalized Image Generation

    Donald Shenaj, Federico Errica, Antonio Carta

    cs.CVcs.AIcs.LGarXiv:2603.21884v12026
  50. I Know What I Don't Know: Latent Posterior Factor Models for Multi-Evidence Probabilistic Reasoning

    Aliyu Agboola Alege

    cs.AIcs.LGarXiv:2603.15670v22026
  51. Beyond Single Tokens: Distilling Discrete Diffusion Models via Discrete MMD

    Emiel Hoogeboom, David Ruhe, Jonathan Heek +2

    cs.LGcs.CVstat.MLarXiv:2603.20155v12026
  52. Remasking Discrete Diffusion Models with Inference-Time Scaling

    Guanghan Wang, Yair Schiff, Subham Sekhar Sahoo +1

    cs.LGstat.MLarXiv:2503.00307v42025
  53. SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis

    Shahriar Noroozizadeh, Xiaobin Shen, Jeremy C. Weiss +1

    cs.LGcs.AIstat.MLarXiv:2603.05483v12026
  54. MedMentions: A Large Biomedical Corpus Annotated with UMLS Concepts

    Sunil Mohan, Donghui Li

    cs.CLcs.LGarXiv:1902.09476v12019
  55. Surprised by Attention: Predictable Query Dynamics for Time Series Anomaly Detection

    Kadir-Kaan Özer, René Ebeling, Markus Enzweiler

    cs.LGcs.AIarXiv:2603.12916v32026
  56. Denoising Diffusion Generative Models Secretly Calculate Attentions

    Farzan Haddadi, Leila Monfared, Ebrahim Rezaii +3

    cs.AIcs.CVcs.LGarXiv:2609.00885v12026
  57. Model-Based Reinforcement Learning with a Generative Model is Minimax Optimal

    Alekh Agarwal, Sham Kakade, Lin F. Yang

    cs.LGmath.PRstat.MLarXiv:1906.03804v32019
  58. Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models

    Arnas Uselis, Andrea Dittadi, Seong Joon Oh

    cs.CVcs.LGarXiv:2602.24264v22026
  59. Transform-Invariant Generative Ray Path Sampling for Efficient Radio Propagation Modeling

    Jérome Eertmans, Enrico M. Vitucci, Vittorio Degli-Esposti +3

    cs.LGeess.SParXiv:2603.01655v22026
  60. On the Mechanism and Dynamics of Modular Addition: Fourier Features, Lottery Ticket, and Grokking

    Jianliang He, Leda Wang, Siyu Chen +1

    cs.LGmath.OCstat.MLarXiv:2602.16849v12026