Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1,141 to 1,200 of 20,193

  1. NVIDIA Nemotron 3: Efficient and Open Intelligence

    NVIDIA, :, Aaron Blakeman +356

    cs.CLcs.AIcs.LGarXiv:2512.20856v12025
  2. Dual Supervised Learning

    Yingce Xia, Tao Qin, Wei Chen +3

    cs.LGstat.MLarXiv:1707.00415v12017
  3. StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?

    Yanxu Chen, Zijun Yao, Yantao Liu +5

    cs.LGcs.CLarXiv:2510.02209v22025
  4. Scaling Open-Ended Reasoning to Predict the Future

    Nikhil Chandak, Shashwat Goel, Ameya Prabhu +2

    cs.LGcs.CLarXiv:2512.25070v22025
  5. Neural Networks are Convex Regularizers: Exact Polynomial-time Convex Optimization Formulations for Two-layer Networks

    Mert Pilanci, Tolga Ergen

    cs.LGcs.CCstat.MLarXiv:2002.10553v22020
  6. Sublinear Optimization for Machine Learning

    Kenneth L. Clarkson, Elad Hazan, David P. Woodruff

    cs.LGarXiv:1010.4408v12010
  7. Thought Communication in Multiagent Collaboration

    Yujia Zheng, Zhuokai Zhao, Zijian Li +4

    cs.LGcs.AIcs.MAarXiv:2510.20733v12025
  8. Biologically Inspired Spiking Neurons : Piecewise Linear Models and Digital Implementation

    Hamid Soleimani, Arash Ahmadi, Mohammad Bavandpour

    cs.LGcs.NEq-bio.NCarXiv:1212.3765v12012
  9. The Information Geometry of Mirror Descent

    Garvesh Raskutti, Sayan Mukherjee

    stat.MLcs.LGarXiv:1310.7780v22013
  10. COCO-GAN: Generation by Parts via Conditional Coordinating

    Chieh Hubert Lin, Chia-Che Chang, Yu-Sheng Chen +3

    cs.LGcs.CVstat.MLarXiv:1904.00284v42019
  11. VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models

    Xinlei Yu, Chengming Xu, Guibin Zhang +7

    cs.CVcs.AIcs.LGarXiv:2511.11007v22025
  12. Practical Contextual Bandits with Regression Oracles

    Dylan J. Foster, Alekh Agarwal, Miroslav Dudík +2

    cs.LGstat.MLarXiv:1803.01088v12018
  13. Efficient Orthogonal Parametrisation of Recurrent Neural Networks Using Householder Reflections

    Zakaria Mhammedi, Andrew Hellicar, Ashfaqur Rahman +1

    cs.LGarXiv:1612.00188v52016
  14. DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning

    Shih-Yang Liu, Xin Dong, Ximing Lu +9

    cs.LGcs.AIcs.CLarXiv:2510.15110v12025
  15. Understanding the Effect of Out-of-distribution Examples and Interactive Explanations on Human-AI Decision Making

    Han Liu, Vivian Lai, Chenhao Tan

    cs.AIcs.CYcs.HCarXiv:2101.05303v42021
  16. Multilingual Routing in Mixture-of-Experts

    Lucas Bandarkar, Chenyuan Yang, Mohsen Fayyaz +2

    cs.CLcs.AIcs.LGarXiv:2510.04694v22025
  17. ToolUniverse: An open platform for democratizing AI scientists

    Shanghua Gao, Richard Zhu, Pengwei Sui +8

    cs.AIcs.LGarXiv:2509.23426v32025
  18. Distinguishing cause from effect using observational data: methods and benchmarks

    Joris M. Mooij, Jonas Peters, Dominik Janzing +2

    cs.LGcs.AIstat.MLarXiv:1412.3773v32014
  19. Muon Outperforms Adam in Tail-End Associative Memory Learning

    Shuche Wang, Fengzhuo Zhang, Jiaxiang Li +6

    cs.LGcs.AImath.OCarXiv:2509.26030v22025
  20. Learning to Optimize Multi-Objective Alignment Through Dynamic Reward Weighting

    Yining Lu, Zilong Wang, Shiyang Li +6

    cs.LGcs.CLarXiv:2509.11452v22025
  21. On the Pitfalls of Measuring Emergent Communication

    Ryan Lowe, Jakob Foerster, Y-Lan Boureau +2

    cs.LGcs.AIcs.CLarXiv:1903.05168v12019
  22. Meta-RL Induces Exploration in Language Agents

    Yulun Jiang, Liangze Jiang, Damien Teney +2

    cs.LGcs.AIarXiv:2512.16848v22025
  23. Correlational Neural Networks

    Sarath Chandar, Mitesh M. Khapra, Hugo Larochelle +1

    cs.CLcs.LGcs.NEarXiv:1504.07225v32015
  24. LaPred: Lane-Aware Prediction of Multi-Modal Future Trajectories of Dynamic Agents

    ByeoungDo Kim, Seong Hyeon Park, Seokhwan Lee +5

    cs.CVcs.LGarXiv:2104.00249v12021
  25. MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes

    Yu Ying Chiu, Michael S. Lee, Rachel Calcott +17

    cs.CLcs.AIcs.CYarXiv:2510.16380v22025
  26. Remote Labor Index: Measuring AI Automation of Remote Work

    Mantas Mazeika, Alice Gatti, Cristina Menghini +44

    cs.LGcs.AIcs.CLarXiv:2510.26787v12025
  27. VC Classes are Adversarially Robustly Learnable, but Only Improperly

    Omar Montasser, Steve Hanneke, Nathan Srebro

    cs.LGstat.MLarXiv:1902.04217v22019
  28. Learning Unmasking Policies for Diffusion Language Models

    Metod Jazbec, Theo X. Olausson, Louis Béthune +6

    cs.LGarXiv:2512.09106v42025
  29. pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation

    Hansheng Chen, Kai Zhang, Hao Tan +3

    cs.LGcs.AIcs.CVarXiv:2510.14974v32025
  30. OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation

    Henry Herzog, Favyen Bastani, Yawen Zhang +23

    cs.CVcs.LGarXiv:2511.13655v12025
  31. A PAC-Bayesian Tutorial with A Dropout Bound

    David McAllester

    cs.LGarXiv:1307.2118v12013
  32. Building a Foundational Guardrail for General Agentic Systems via Synthetic Data

    Yue Huang, Hang Hua, Yujun Zhou +11

    cs.LGcs.AIcs.CLarXiv:2510.09781v12025
  33. Skeleton Image Representation for 3D Action Recognition based on Tree Structure and Reference Joints

    Carlos Caetano, François Brémond, William Robson Schwartz

    cs.CVcs.LGarXiv:1909.05704v12019
  34. Scaling Behavior of Discrete Diffusion Language Models

    Dimitri von Rütte, Janis Fluri, Omead Pooladzandi +3

    cs.LGarXiv:2512.10858v32025
  35. TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models

    Zheng Ding, Weirui Ye

    cs.LGcs.AIcs.CVarXiv:2512.08153v12025
  36. Uniqueness of Low-Rank Matrix Completion by Rigidity Theory

    Amit Singer, Mihai Cucuringu

    cs.LGarXiv:0902.3846v12009
  37. Multi-Scale Dense Networks for Resource Efficient Image Classification

    Gao Huang, Danlu Chen, Tianhong Li +3

    cs.LGarXiv:1703.09844v52017
  38. Parsimonious Black-Box Adversarial Attacks via Efficient Combinatorial Optimization

    Seungyong Moon, Gaon An, Hyun Oh Song

    cs.LGcs.CRcs.CVarXiv:1905.06635v22019
  39. On the Relation Between the Sharpest Directions of DNN Loss and the SGD Step Length

    Stanisław Jastrzębski, Zachary Kenton, Nicolas Ballas +3

    stat.MLcs.LGarXiv:1807.05031v62018
  40. State-Relabeling Adversarial Active Learning

    Beichen Zhang, Liang Li, Shijie Yang +3

    cs.CVcs.LGarXiv:2004.04943v12020
  41. Attention Is All You Need for KV Cache in Diffusion LLMs

    Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen

    cs.CLcs.AIcs.LGarXiv:2510.14973v22025
  42. LongCat-Flash-Omni Technical Report

    Meituan LongCat Team, Bairui Wang, Bayan +130

    cs.MMcs.AIcs.CLarXiv:2511.00279v22025
  43. KeyPose: Multi-View 3D Labeling and Keypoint Estimation for Transparent Objects

    Xingyu Liu, Rico Jonschkowski, Anelia Angelova +1

    cs.CVcs.LGcs.ROarXiv:1912.02805v22019
  44. PAN: A World Model for General, Actionable, and Long-Horizon World Simulation

    PAN Team, Zihan Liu, Yi Gu +12

    cs.CVcs.AIcs.CLarXiv:2511.09057v42025
  45. See, Point, Fly: A Learning-Free VLM Framework for Universal Unmanned Aerial Navigation

    Chih Yao Hu, Yang-Sen Lin, Yuna Lee +7

    cs.ROcs.AIcs.CLarXiv:2509.22653v12025
  46. Semi-Implicit Variational Inference

    Mingzhang Yin, Mingyuan Zhou

    stat.MLcs.LGstat.COarXiv:1805.11183v12018
  47. Defending Against Physically Realizable Attacks on Image Classification

    Tong Wu, Liang Tong, Yevgeniy Vorobeychik

    cs.LGcs.AIcs.CVarXiv:1909.09552v22019
  48. Frequentist Regret Bounds for Randomized Least-Squares Value Iteration

    Andrea Zanette, David Brandfonbrener, Emma Brunskill +2

    cs.LGstat.MLarXiv:1911.00567v72019
  49. Learning Latent Space Energy-Based Prior Model

    Bo Pang, Tian Han, Erik Nijkamp +2

    stat.MLcs.LGarXiv:2006.08205v22020
  50. Tongyi DeepResearch Technical Report

    Tongyi DeepResearch Team, Baixuan Li, Bo Zhang +54

    cs.CLcs.AIcs.IRarXiv:2510.24701v32025
  51. An optimal algorithm for the Thresholding Bandit Problem

    Andrea Locatelli, Maurilio Gutzeit, Alexandra Carpentier

    stat.MLcs.LGarXiv:1605.08671v12016
  52. GIANT: Globally Improved Approximate Newton Method for Distributed Optimization

    Shusen Wang, Farbod Roosta-Khorasani, Peng Xu +1

    cs.LGcs.DCmath.OCarXiv:1709.03528v52017
  53. Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence

    Sean McLeish, Ang Li, John Kirchenbauer +7

    cs.CLcs.AIcs.LGarXiv:2511.07384v12025
  54. ICE-BeeM: Identifiable Conditional Energy-Based Deep Models Based on Nonlinear ICA

    Ilyes Khemakhem, Ricardo Pio Monti, Diederik P. Kingma +1

    stat.MLcs.LGarXiv:2002.11537v42020
  55. Online Convex Optimization in Adversarial Markov Decision Processes

    Aviv Rosenberg, Yishay Mansour

    cs.LGcs.AIstat.MLarXiv:1905.07773v12019
  56. Training AI Co-Scientists Using Rubric Rewards

    Shashwat Goel, Rishi Hazra, Dulhan Jayalath +8

    cs.LGcs.CLcs.HCarXiv:2512.23707v12025
  57. Agentic Entropy-Balanced Policy Optimization

    Guanting Dong, Licheng Bao, Zhongyuan Wang +11

    cs.LGcs.AIcs.CLarXiv:2510.14545v12025
  58. Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed

    Yonggan Fu, Lexington Whalen, Zhifan Ye +11

    cs.CLcs.AIcs.LGarXiv:2512.14067v22025
  59. Guided Self-Evolving LLMs with Minimal Human Supervision

    Wenhao Yu, Zhenwen Liang, Chengsong Huang +4

    cs.AIcs.CLcs.LGarXiv:2512.02472v12025
  60. Efficient Reinforcement Learning Using Recursive Least-Squares Methods

    H. He, D. Hu, X. Xu

    cs.LGcs.AIarXiv:1106.0707v12011