Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1,441 to 1,500 of 20,454

  1. LaPred: Lane-Aware Prediction of Multi-Modal Future Trajectories of Dynamic Agents

    ByeoungDo Kim, Seong Hyeon Park, Seokhwan Lee +5

    cs.CVcs.LGarXiv:2104.00249v12021
  2. MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes

    Yu Ying Chiu, Michael S. Lee, Rachel Calcott +17

    cs.CLcs.AIcs.CYarXiv:2510.16380v22025
  3. Remote Labor Index: Measuring AI Automation of Remote Work

    Mantas Mazeika, Alice Gatti, Cristina Menghini +44

    cs.LGcs.AIcs.CLarXiv:2510.26787v12025
  4. VC Classes are Adversarially Robustly Learnable, but Only Improperly

    Omar Montasser, Steve Hanneke, Nathan Srebro

    cs.LGstat.MLarXiv:1902.04217v22019
  5. Learning Unmasking Policies for Diffusion Language Models

    Metod Jazbec, Theo X. Olausson, Louis Béthune +6

    cs.LGarXiv:2512.09106v42025
  6. pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation

    Hansheng Chen, Kai Zhang, Hao Tan +3

    cs.LGcs.AIcs.CVarXiv:2510.14974v32025
  7. OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation

    Henry Herzog, Favyen Bastani, Yawen Zhang +23

    cs.CVcs.LGarXiv:2511.13655v12025
  8. A PAC-Bayesian Tutorial with A Dropout Bound

    David McAllester

    cs.LGarXiv:1307.2118v12013
  9. Building a Foundational Guardrail for General Agentic Systems via Synthetic Data

    Yue Huang, Hang Hua, Yujun Zhou +11

    cs.LGcs.AIcs.CLarXiv:2510.09781v12025
  10. Skeleton Image Representation for 3D Action Recognition based on Tree Structure and Reference Joints

    Carlos Caetano, François Brémond, William Robson Schwartz

    cs.CVcs.LGarXiv:1909.05704v12019
  11. Scaling Behavior of Discrete Diffusion Language Models

    Dimitri von Rütte, Janis Fluri, Omead Pooladzandi +3

    cs.LGarXiv:2512.10858v32025
  12. TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models

    Zheng Ding, Weirui Ye

    cs.LGcs.AIcs.CVarXiv:2512.08153v12025
  13. Uniqueness of Low-Rank Matrix Completion by Rigidity Theory

    Amit Singer, Mihai Cucuringu

    cs.LGarXiv:0902.3846v12009
  14. Multi-Scale Dense Networks for Resource Efficient Image Classification

    Gao Huang, Danlu Chen, Tianhong Li +3

    cs.LGarXiv:1703.09844v52017
  15. Parsimonious Black-Box Adversarial Attacks via Efficient Combinatorial Optimization

    Seungyong Moon, Gaon An, Hyun Oh Song

    cs.LGcs.CRcs.CVarXiv:1905.06635v22019
  16. On the Relation Between the Sharpest Directions of DNN Loss and the SGD Step Length

    Stanisław Jastrzębski, Zachary Kenton, Nicolas Ballas +3

    stat.MLcs.LGarXiv:1807.05031v62018
  17. State-Relabeling Adversarial Active Learning

    Beichen Zhang, Liang Li, Shijie Yang +3

    cs.CVcs.LGarXiv:2004.04943v12020
  18. Attention Is All You Need for KV Cache in Diffusion LLMs

    Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen

    cs.CLcs.AIcs.LGarXiv:2510.14973v22025
  19. LongCat-Flash-Omni Technical Report

    Meituan LongCat Team, Bairui Wang, Bayan +130

    cs.MMcs.AIcs.CLarXiv:2511.00279v22025
  20. KeyPose: Multi-View 3D Labeling and Keypoint Estimation for Transparent Objects

    Xingyu Liu, Rico Jonschkowski, Anelia Angelova +1

    cs.CVcs.LGcs.ROarXiv:1912.02805v22019
  21. PAN: A World Model for General, Actionable, and Long-Horizon World Simulation

    PAN Team, Zihan Liu, Yi Gu +12

    cs.CVcs.AIcs.CLarXiv:2511.09057v42025
  22. See, Point, Fly: A Learning-Free VLM Framework for Universal Unmanned Aerial Navigation

    Chih Yao Hu, Yang-Sen Lin, Yuna Lee +7

    cs.ROcs.AIcs.CLarXiv:2509.22653v12025
  23. Semi-Implicit Variational Inference

    Mingzhang Yin, Mingyuan Zhou

    stat.MLcs.LGstat.COarXiv:1805.11183v12018
  24. Defending Against Physically Realizable Attacks on Image Classification

    Tong Wu, Liang Tong, Yevgeniy Vorobeychik

    cs.LGcs.AIcs.CVarXiv:1909.09552v22019
  25. Frequentist Regret Bounds for Randomized Least-Squares Value Iteration

    Andrea Zanette, David Brandfonbrener, Emma Brunskill +2

    cs.LGstat.MLarXiv:1911.00567v72019
  26. Learning Latent Space Energy-Based Prior Model

    Bo Pang, Tian Han, Erik Nijkamp +2

    stat.MLcs.LGarXiv:2006.08205v22020
  27. Tongyi DeepResearch Technical Report

    Tongyi DeepResearch Team, Baixuan Li, Bo Zhang +54

    cs.CLcs.AIcs.IRarXiv:2510.24701v32025
  28. An optimal algorithm for the Thresholding Bandit Problem

    Andrea Locatelli, Maurilio Gutzeit, Alexandra Carpentier

    stat.MLcs.LGarXiv:1605.08671v12016
  29. GIANT: Globally Improved Approximate Newton Method for Distributed Optimization

    Shusen Wang, Farbod Roosta-Khorasani, Peng Xu +1

    cs.LGcs.DCmath.OCarXiv:1709.03528v52017
  30. Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence

    Sean McLeish, Ang Li, John Kirchenbauer +7

    cs.CLcs.AIcs.LGarXiv:2511.07384v12025
  31. ICE-BeeM: Identifiable Conditional Energy-Based Deep Models Based on Nonlinear ICA

    Ilyes Khemakhem, Ricardo Pio Monti, Diederik P. Kingma +1

    stat.MLcs.LGarXiv:2002.11537v42020
  32. Online Convex Optimization in Adversarial Markov Decision Processes

    Aviv Rosenberg, Yishay Mansour

    cs.LGcs.AIstat.MLarXiv:1905.07773v12019
  33. Training AI Co-Scientists Using Rubric Rewards

    Shashwat Goel, Rishi Hazra, Dulhan Jayalath +8

    cs.LGcs.CLcs.HCarXiv:2512.23707v12025
  34. Agentic Entropy-Balanced Policy Optimization

    Guanting Dong, Licheng Bao, Zhongyuan Wang +11

    cs.LGcs.AIcs.CLarXiv:2510.14545v12025
  35. Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed

    Yonggan Fu, Lexington Whalen, Zhifan Ye +11

    cs.CLcs.AIcs.LGarXiv:2512.14067v22025
  36. Guided Self-Evolving LLMs with Minimal Human Supervision

    Wenhao Yu, Zhenwen Liang, Chengsong Huang +4

    cs.AIcs.CLcs.LGarXiv:2512.02472v12025
  37. Efficient Reinforcement Learning Using Recursive Least-Squares Methods

    H. He, D. Hu, X. Xu

    cs.LGcs.AIarXiv:1106.0707v12011
  38. Does FLUX Already Know How to Perform Physically Plausible Image Composition?

    Shilin Lu, Zhuming Lian, Zihan Zhou +3

    cs.CVcs.AIcs.LGarXiv:2509.21278v42025
  39. Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

    Xin Qiu, Yulu Gan, Conor F. Hayes +6

    cs.LGcs.AIcs.NEarXiv:2509.24372v32025
  40. VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models

    Guochao Jiang, Wenfeng Feng, Guofeng Quan +4

    cs.LGcs.CLarXiv:2509.19803v12025
  41. Search Self-play: Pushing the Frontier of Agent Capability without Supervision

    Hongliang Lu, Yuhang Wen, Pengyu Cheng +7

    cs.LGarXiv:2510.18821v32025
  42. Multi-view Vector-valued Manifold Regularization for Multi-label Image Classification

    Yong Luo, Dacheng Tao, Chang Xu +3

    stat.MLcs.CVcs.LGarXiv:1904.03921v12019
  43. Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization

    Nikita Kachaev, Mikhail Kolosov, Daniil Zelezetsky +2

    cs.LGcs.AIcs.ROarXiv:2510.25616v12025
  44. SimpleFold: Folding Proteins is Simpler than You Think

    Yuyang Wang, Jiarui Lu, Navdeep Jaitly +2

    cs.LGq-bio.QMarXiv:2509.18480v42025
  45. Deep Residual Learning in the JPEG Transform Domain

    Max Ehrlich, Larry Davis

    cs.LGcs.CVstat.MLarXiv:1812.11690v32018
  46. A Review of the Gumbel-max Trick and its Extensions for Discrete Stochasticity in Machine Learning

    Iris A. M. Huijben, Wouter Kool, Max B. Paulus +1

    cs.LGstat.MLarXiv:2110.01515v22021
  47. Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents

    Guoqing Wang, Sunhao Dai, Guangze Ye +5

    cs.CLcs.AIcs.LGarXiv:2510.14967v22025
  48. Stable Gaussian Process based Tracking Control of Euler-Lagrange Systems

    Thomas Beckers, Dana Kulić, Sandra Hirche

    cs.LGeess.SYstat.MLarXiv:1806.07190v22018
  49. DragFlow: Unleashing DiT Priors with Region Based Supervision for Drag Editing

    Zihan Zhou, Shilin Lu, Shuli Leng +4

    cs.CVcs.AIcs.LGarXiv:2510.02253v32025
  50. Algorithms for Dynamic Spectrum Access with Learning for Cognitive Radio

    Jayakrishnan Unnikrishnan, Venugopal Veeravalli

    cs.NIcs.LGarXiv:0807.2677v42008
  51. Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models

    Boxin Wang, Chankyu Lee, Nayeon Lee +9

    cs.CLcs.AIcs.LGarXiv:2512.13607v22025
  52. ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration

    Hongjin Su, Shizhe Diao, Ximing Lu +13

    cs.CLcs.AIcs.LGarXiv:2511.21689v12025
  53. RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

    Zhiyuan Zeng, Hamish Ivison, Yiping Wang +14

    cs.CLcs.LGarXiv:2511.07317v22025
  54. Eliciting Secret Knowledge from Language Models

    Bartosz Cywiński, Emil Ryd, Rowan Wang +4

    cs.LGarXiv:2510.01070v22025
  55. Patient2Vec: A Personalized Interpretable Deep Representation of the Longitudinal Electronic Health Record

    Jinghe Zhang, Kamran Kowsari, James H. Harrison +2

    q-bio.QMcs.AIcs.IRarXiv:1810.04793v32018
  56. Decision Trees for Decision-Making under the Predict-then-Optimize Framework

    Adam N. Elmachtoub, Jason Cheuk Nam Liang, Ryan McNellis

    cs.LGmath.OCstat.MLarXiv:2003.00360v22020
  57. LFM2 Technical Report

    Alexander Amini, Anna Banaszak, Harold Benoit +30

    cs.LGcs.AIarXiv:2511.23404v12025
  58. Kalman Filtering with Intermittent Observations: Weak Convergence to a Stationary Distribution

    Soummya Kar, Bruno Sinopoli, Jose M. F. Moura

    cs.ITcs.LGmath.STarXiv:0903.2890v22009
  59. FlowRL: Matching Reward Distributions for LLM Reasoning

    Xuekai Zhu, Daixuan Cheng, Dinghuai Zhang +20

    cs.LGcs.AIcs.CLarXiv:2509.15207v32025
  60. Monotonic Calibrated Interpolated Look-Up Tables

    Maya Gupta, Andrew Cotter, Jan Pfeifer +5

    cs.LGarXiv:1505.06378v32015