Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1,201 to 1,260 of 20,192

  1. Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

    Xin Qiu, Yulu Gan, Conor F. Hayes +6

    cs.LGcs.AIcs.NEarXiv:2509.24372v32025
  2. VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models

    Guochao Jiang, Wenfeng Feng, Guofeng Quan +4

    cs.LGcs.CLarXiv:2509.19803v12025
  3. Search Self-play: Pushing the Frontier of Agent Capability without Supervision

    Hongliang Lu, Yuhang Wen, Pengyu Cheng +7

    cs.LGarXiv:2510.18821v32025
  4. Multi-view Vector-valued Manifold Regularization for Multi-label Image Classification

    Yong Luo, Dacheng Tao, Chang Xu +3

    stat.MLcs.CVcs.LGarXiv:1904.03921v12019
  5. Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization

    Nikita Kachaev, Mikhail Kolosov, Daniil Zelezetsky +2

    cs.LGcs.AIcs.ROarXiv:2510.25616v12025
  6. SimpleFold: Folding Proteins is Simpler than You Think

    Yuyang Wang, Jiarui Lu, Navdeep Jaitly +2

    cs.LGq-bio.QMarXiv:2509.18480v42025
  7. Deep Residual Learning in the JPEG Transform Domain

    Max Ehrlich, Larry Davis

    cs.LGcs.CVstat.MLarXiv:1812.11690v32018
  8. A Review of the Gumbel-max Trick and its Extensions for Discrete Stochasticity in Machine Learning

    Iris A. M. Huijben, Wouter Kool, Max B. Paulus +1

    cs.LGstat.MLarXiv:2110.01515v22021
  9. Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents

    Guoqing Wang, Sunhao Dai, Guangze Ye +5

    cs.CLcs.AIcs.LGarXiv:2510.14967v22025
  10. Stable Gaussian Process based Tracking Control of Euler-Lagrange Systems

    Thomas Beckers, Dana Kulić, Sandra Hirche

    cs.LGeess.SYstat.MLarXiv:1806.07190v22018
  11. DragFlow: Unleashing DiT Priors with Region Based Supervision for Drag Editing

    Zihan Zhou, Shilin Lu, Shuli Leng +4

    cs.CVcs.AIcs.LGarXiv:2510.02253v32025
  12. Algorithms for Dynamic Spectrum Access with Learning for Cognitive Radio

    Jayakrishnan Unnikrishnan, Venugopal Veeravalli

    cs.NIcs.LGarXiv:0807.2677v42008
  13. Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models

    Boxin Wang, Chankyu Lee, Nayeon Lee +9

    cs.CLcs.AIcs.LGarXiv:2512.13607v22025
  14. ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration

    Hongjin Su, Shizhe Diao, Ximing Lu +13

    cs.CLcs.AIcs.LGarXiv:2511.21689v12025
  15. RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

    Zhiyuan Zeng, Hamish Ivison, Yiping Wang +14

    cs.CLcs.LGarXiv:2511.07317v22025
  16. Eliciting Secret Knowledge from Language Models

    Bartosz Cywiński, Emil Ryd, Rowan Wang +4

    cs.LGarXiv:2510.01070v22025
  17. Patient2Vec: A Personalized Interpretable Deep Representation of the Longitudinal Electronic Health Record

    Jinghe Zhang, Kamran Kowsari, James H. Harrison +2

    q-bio.QMcs.AIcs.IRarXiv:1810.04793v32018
  18. Decision Trees for Decision-Making under the Predict-then-Optimize Framework

    Adam N. Elmachtoub, Jason Cheuk Nam Liang, Ryan McNellis

    cs.LGmath.OCstat.MLarXiv:2003.00360v22020
  19. LFM2 Technical Report

    Alexander Amini, Anna Banaszak, Harold Benoit +30

    cs.LGcs.AIarXiv:2511.23404v12025
  20. Kalman Filtering with Intermittent Observations: Weak Convergence to a Stationary Distribution

    Soummya Kar, Bruno Sinopoli, Jose M. F. Moura

    cs.ITcs.LGmath.STarXiv:0903.2890v22009
  21. FlowRL: Matching Reward Distributions for LLM Reasoning

    Xuekai Zhu, Daixuan Cheng, Dinghuai Zhang +20

    cs.LGcs.AIcs.CLarXiv:2509.15207v32025
  22. Monotonic Calibrated Interpolated Look-Up Tables

    Maya Gupta, Andrew Cotter, Jan Pfeifer +5

    cs.LGarXiv:1505.06378v32015
  23. Inference Compilation and Universal Probabilistic Programming

    Tuan Anh Le, Atilim Gunes Baydin, Frank Wood

    cs.AIcs.LGstat.MLarXiv:1610.09900v22016
  24. ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMs

    Wonjun Kang, Kevin Galim, Seunghyuk Oh +8

    cs.LGarXiv:2510.04767v22025
  25. AlphaFlow: Understanding and Improving MeanFlow Models

    Huijie Zhang, Aliaksandr Siarohin, Willi Menapace +4

    cs.CVcs.LGarXiv:2510.20771v12025
  26. Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training

    Junkai Zhang, Zihao Wang, Lin Gui +7

    cs.LGcs.AIarXiv:2509.21500v32025
  27. Model Compression with Adversarial Robustness: A Unified Optimization Framework

    Shupeng Gui, Haotao Wang, Chen Yu +3

    cs.LGstat.MLarXiv:1902.03538v32019
  28. Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

    Jiawei Wang, Jiacai Liu, Yuqian Fu +7

    cs.LGcs.CLarXiv:2509.09265v12025
  29. Residual Off-Policy RL for Finetuning Behavior Cloning Policies

    Lars Ankile, Zhenyu Jiang, Rocky Duan +3

    cs.ROcs.LGarXiv:2509.19301v22025
  30. Discovering Reinforcement Learning Algorithms

    Junhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki +4

    cs.LGcs.AIarXiv:2007.08794v32020
  31. TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times

    Jintao Zhang, Kaiwen Zheng, Kai Jiang +5

    cs.CVcs.AIcs.LGarXiv:2512.16093v12025
  32. Global Optimality in Tensor Factorization, Deep Learning, and Beyond

    Benjamin D. Haeffele, Rene Vidal

    math.NAcs.LGstat.MLarXiv:1506.07540v12015
  33. Ensemble Methods as a Defense to Adversarial Perturbations Against Deep Neural Networks

    Thilo Strauss, Markus Hanselmann, Andrej Junginger +1

    stat.MLcs.LGarXiv:1709.03423v22017
  34. Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization

    Vage Egiazarian, Roberto L. Castro, Denis Kuznedelev +8

    cs.LGarXiv:2509.23202v32025
  35. AnyUp: Universal Feature Upsampling

    Thomas Wimmer, Prune Truong, Marie-Julie Rakotosaona +4

    cs.CVcs.LGarXiv:2510.12764v22025
  36. Scaling laws for amplitude surrogates

    Henning Bahl, Victor Bresó-Pla, Anja Butter +1

    hep-phcs.LGarXiv:2601.13308v12026
  37. IQP Born Machines under Data-dependent and Agnostic Initialization Strategies

    Sacha Lerch, Joseph Bowles, Ricard Puig +3

    quant-phcs.LGstat.MLarXiv:2603.14576v12026
  38. OpenMAG: A Comprehensive Benchmark for Multimodal-Attributed Graph

    Chenxi Wan, Xunkai Li, Yilong Zuo +6

    cs.LGarXiv:2602.05576v12026
  39. Simple and Effective Zero-shot Cross-lingual Phoneme Recognition

    Qiantong Xu, Alexei Baevski, Michael Auli

    cs.CLcs.LGcs.SDarXiv:2109.11680v12021
  40. Sharp Bounds on the Approximation Rates, Metric Entropy, and $n$-widths of Shallow Neural Networks

    Jonathan W. Siegel, Jinchao Xu

    stat.MLcs.ITcs.LGarXiv:2101.12365v102021
  41. Particle Filter Networks with Application to Visual Localization

    Peter Karkus, David Hsu, Wee Sun Lee

    cs.ROcs.AIcs.CVarXiv:1805.08975v32018
  42. Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models

    Boyi Deng, Xu Wang, Yaoning Wang +15

    cs.CLcs.LGarXiv:2605.11887v12026
  43. Self-supervised Trajectory Representation Learning with Temporal Regularities and Travel Semantics

    Jiawei Jiang, Dayan Pan, Houxing Ren +3

    cs.LGarXiv:2211.09510v42022
  44. Generating Long Videos of Dynamic Scenes

    Brooks, Tim, Hellsten, Janne, Aittala, Miika +6

    cs.CVcs.AIcs.LGarXiv:2206.03429v22022
  45. Residual Energy-Based Models for Text Generation

    Deng, Yuntian, Bakhtin, Anton, Ott, Myle +2

    cs.CLcs.LGarXiv:2004.11714v12020
  46. A Definition of AGI

    Hendrycks, Dan, Song, Dawn, Szegedy, Christian +30

    cs.AIcs.LGarXiv:2510.18212v32025
  47. DeepHeart: Semi-Supervised Sequence Learning for Cardiovascular Risk Prediction

    Brandon Ballinger, Johnson Hsieh, Avesh Singh +8

    cs.LGcs.AIstat.MLarXiv:1802.02511v12018
  48. Recurrent Neural Network Transducer for Audio-Visual Speech Recognition

    Takaki Makino, Hank Liao, Yannis Assael +4

    eess.AScs.CLcs.CVarXiv:1911.04890v12019
  49. Dynamic Pricing in High-dimensions

    Adel Javanmard, Hamid Nazerzadeh

    stat.MLcs.LGarXiv:1609.07574v42016
  50. The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward

    Long Li, Zhijian Zhou, Jiaran Hao +9

    cs.LGcs.AIarXiv:2509.07430v42025
  51. ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases

    Ziqian Zhong, Aditi Raghunathan, Nicholas Carlini

    cs.LGcs.CLarXiv:2510.20270v12025
  52. Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward

    Eshwar Reddy M, Sourav Karmakar

    cs.AIcs.LGarXiv:2609.09776v12026
  53. Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation

    Yujun Zhou, Zhenwen Liang, Haolin Liu +7

    cs.LGcs.CLarXiv:2509.15194v32025
  54. Explainable AI for clinical and remote health applications: a survey on tabular and time series data

    Flavio Di Martino, Franca Delmastro

    cs.LGcs.AIarXiv:2209.06528v12022
  55. Evaluating Reinforcement Learning Algorithms in Observational Health Settings

    Omer Gottesman, Fredrik Johansson, Joshua Meier +17

    cs.LGstat.MLarXiv:1805.12298v12018
  56. Artificial Intelligence Algorithms for the Detection of Pathologies Related to Lung Cancer through Image Analysis using Convolutional Neural Networks and Data Augmentation: a systematic mapping of the literature

    Pablo Ramirez Amador

    cs.LGcs.CLarXiv:2609.10652v12026
  57. Stabilizing Reinforcement Learning with LLMs: Formulation and Practices

    Chujie Zheng, Kai Dang, Bowen Yu +10

    cs.LGcs.AIcs.CLarXiv:2512.01374v32025
  58. Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents

    Shuai Shao, Qihan Ren, Chen Qian +8

    cs.AIcs.CLcs.LGarXiv:2509.26354v22025
  59. Scaling LLM Multi-turn RL with End-to-end Summarization-based Context Management

    Miao Lu, Weiwei Sun, Weihua Du +4

    cs.CLcs.AIcs.LGarXiv:2510.06727v12025
  60. GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping

    Jing Wang, Jiajun Liang, Jie Liu +10

    cs.CVcs.LGarXiv:2510.22319v22025