Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
1,141 to 1,200 of 20,193
NVIDIA Nemotron 3: Efficient and Open Intelligence
NVIDIA, :, Aaron Blakeman +356
cs.CLcs.AIcs.LGarXiv:2512.20856v12025Dual Supervised Learning
Yingce Xia, Tao Qin, Wei Chen +3
cs.LGstat.MLarXiv:1707.00415v12017StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
Yanxu Chen, Zijun Yao, Yantao Liu +5
cs.LGcs.CLarXiv:2510.02209v22025Scaling Open-Ended Reasoning to Predict the Future
Nikhil Chandak, Shashwat Goel, Ameya Prabhu +2
cs.LGcs.CLarXiv:2512.25070v22025Neural Networks are Convex Regularizers: Exact Polynomial-time Convex Optimization Formulations for Two-layer Networks
Mert Pilanci, Tolga Ergen
cs.LGcs.CCstat.MLarXiv:2002.10553v22020Sublinear Optimization for Machine Learning
Kenneth L. Clarkson, Elad Hazan, David P. Woodruff
cs.LGarXiv:1010.4408v12010Thought Communication in Multiagent Collaboration
Yujia Zheng, Zhuokai Zhao, Zijian Li +4
cs.LGcs.AIcs.MAarXiv:2510.20733v12025Biologically Inspired Spiking Neurons : Piecewise Linear Models and Digital Implementation
Hamid Soleimani, Arash Ahmadi, Mohammad Bavandpour
cs.LGcs.NEq-bio.NCarXiv:1212.3765v12012The Information Geometry of Mirror Descent
Garvesh Raskutti, Sayan Mukherjee
stat.MLcs.LGarXiv:1310.7780v22013COCO-GAN: Generation by Parts via Conditional Coordinating
Chieh Hubert Lin, Chia-Che Chang, Yu-Sheng Chen +3
cs.LGcs.CVstat.MLarXiv:1904.00284v42019VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
Xinlei Yu, Chengming Xu, Guibin Zhang +7
cs.CVcs.AIcs.LGarXiv:2511.11007v22025Practical Contextual Bandits with Regression Oracles
Dylan J. Foster, Alekh Agarwal, Miroslav Dudík +2
cs.LGstat.MLarXiv:1803.01088v12018Efficient Orthogonal Parametrisation of Recurrent Neural Networks Using Householder Reflections
Zakaria Mhammedi, Andrew Hellicar, Ashfaqur Rahman +1
cs.LGarXiv:1612.00188v52016DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning
Shih-Yang Liu, Xin Dong, Ximing Lu +9
cs.LGcs.AIcs.CLarXiv:2510.15110v12025Understanding the Effect of Out-of-distribution Examples and Interactive Explanations on Human-AI Decision Making
Han Liu, Vivian Lai, Chenhao Tan
cs.AIcs.CYcs.HCarXiv:2101.05303v42021Multilingual Routing in Mixture-of-Experts
Lucas Bandarkar, Chenyuan Yang, Mohsen Fayyaz +2
cs.CLcs.AIcs.LGarXiv:2510.04694v22025ToolUniverse: An open platform for democratizing AI scientists
Shanghua Gao, Richard Zhu, Pengwei Sui +8
cs.AIcs.LGarXiv:2509.23426v32025Distinguishing cause from effect using observational data: methods and benchmarks
Joris M. Mooij, Jonas Peters, Dominik Janzing +2
cs.LGcs.AIstat.MLarXiv:1412.3773v32014Muon Outperforms Adam in Tail-End Associative Memory Learning
Shuche Wang, Fengzhuo Zhang, Jiaxiang Li +6
cs.LGcs.AImath.OCarXiv:2509.26030v22025Learning to Optimize Multi-Objective Alignment Through Dynamic Reward Weighting
Yining Lu, Zilong Wang, Shiyang Li +6
cs.LGcs.CLarXiv:2509.11452v22025On the Pitfalls of Measuring Emergent Communication
Ryan Lowe, Jakob Foerster, Y-Lan Boureau +2
cs.LGcs.AIcs.CLarXiv:1903.05168v12019Meta-RL Induces Exploration in Language Agents
Yulun Jiang, Liangze Jiang, Damien Teney +2
cs.LGcs.AIarXiv:2512.16848v22025Correlational Neural Networks
Sarath Chandar, Mitesh M. Khapra, Hugo Larochelle +1
cs.CLcs.LGcs.NEarXiv:1504.07225v32015LaPred: Lane-Aware Prediction of Multi-Modal Future Trajectories of Dynamic Agents
ByeoungDo Kim, Seong Hyeon Park, Seokhwan Lee +5
cs.CVcs.LGarXiv:2104.00249v12021MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
Yu Ying Chiu, Michael S. Lee, Rachel Calcott +17
cs.CLcs.AIcs.CYarXiv:2510.16380v22025Remote Labor Index: Measuring AI Automation of Remote Work
Mantas Mazeika, Alice Gatti, Cristina Menghini +44
cs.LGcs.AIcs.CLarXiv:2510.26787v12025VC Classes are Adversarially Robustly Learnable, but Only Improperly
Omar Montasser, Steve Hanneke, Nathan Srebro
cs.LGstat.MLarXiv:1902.04217v22019Learning Unmasking Policies for Diffusion Language Models
Metod Jazbec, Theo X. Olausson, Louis Béthune +6
cs.LGarXiv:2512.09106v42025pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation
Hansheng Chen, Kai Zhang, Hao Tan +3
cs.LGcs.AIcs.CVarXiv:2510.14974v32025OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation
Henry Herzog, Favyen Bastani, Yawen Zhang +23
cs.CVcs.LGarXiv:2511.13655v12025A PAC-Bayesian Tutorial with A Dropout Bound
David McAllester
cs.LGarXiv:1307.2118v12013Building a Foundational Guardrail for General Agentic Systems via Synthetic Data
Yue Huang, Hang Hua, Yujun Zhou +11
cs.LGcs.AIcs.CLarXiv:2510.09781v12025Skeleton Image Representation for 3D Action Recognition based on Tree Structure and Reference Joints
Carlos Caetano, François Brémond, William Robson Schwartz
cs.CVcs.LGarXiv:1909.05704v12019Scaling Behavior of Discrete Diffusion Language Models
Dimitri von Rütte, Janis Fluri, Omead Pooladzandi +3
cs.LGarXiv:2512.10858v32025TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
Zheng Ding, Weirui Ye
cs.LGcs.AIcs.CVarXiv:2512.08153v12025Uniqueness of Low-Rank Matrix Completion by Rigidity Theory
Amit Singer, Mihai Cucuringu
cs.LGarXiv:0902.3846v12009Multi-Scale Dense Networks for Resource Efficient Image Classification
Gao Huang, Danlu Chen, Tianhong Li +3
cs.LGarXiv:1703.09844v52017Parsimonious Black-Box Adversarial Attacks via Efficient Combinatorial Optimization
Seungyong Moon, Gaon An, Hyun Oh Song
cs.LGcs.CRcs.CVarXiv:1905.06635v22019On the Relation Between the Sharpest Directions of DNN Loss and the SGD Step Length
Stanisław Jastrzębski, Zachary Kenton, Nicolas Ballas +3
stat.MLcs.LGarXiv:1807.05031v62018State-Relabeling Adversarial Active Learning
Beichen Zhang, Liang Li, Shijie Yang +3
cs.CVcs.LGarXiv:2004.04943v12020Attention Is All You Need for KV Cache in Diffusion LLMs
Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen
cs.CLcs.AIcs.LGarXiv:2510.14973v22025LongCat-Flash-Omni Technical Report
Meituan LongCat Team, Bairui Wang, Bayan +130
cs.MMcs.AIcs.CLarXiv:2511.00279v22025KeyPose: Multi-View 3D Labeling and Keypoint Estimation for Transparent Objects
Xingyu Liu, Rico Jonschkowski, Anelia Angelova +1
cs.CVcs.LGcs.ROarXiv:1912.02805v22019PAN: A World Model for General, Actionable, and Long-Horizon World Simulation
PAN Team, Zihan Liu, Yi Gu +12
cs.CVcs.AIcs.CLarXiv:2511.09057v42025See, Point, Fly: A Learning-Free VLM Framework for Universal Unmanned Aerial Navigation
Chih Yao Hu, Yang-Sen Lin, Yuna Lee +7
cs.ROcs.AIcs.CLarXiv:2509.22653v12025Semi-Implicit Variational Inference
Mingzhang Yin, Mingyuan Zhou
stat.MLcs.LGstat.COarXiv:1805.11183v12018Defending Against Physically Realizable Attacks on Image Classification
Tong Wu, Liang Tong, Yevgeniy Vorobeychik
cs.LGcs.AIcs.CVarXiv:1909.09552v22019Frequentist Regret Bounds for Randomized Least-Squares Value Iteration
Andrea Zanette, David Brandfonbrener, Emma Brunskill +2
cs.LGstat.MLarXiv:1911.00567v72019Learning Latent Space Energy-Based Prior Model
Bo Pang, Tian Han, Erik Nijkamp +2
stat.MLcs.LGarXiv:2006.08205v22020Tongyi DeepResearch Technical Report
Tongyi DeepResearch Team, Baixuan Li, Bo Zhang +54
cs.CLcs.AIcs.IRarXiv:2510.24701v32025An optimal algorithm for the Thresholding Bandit Problem
Andrea Locatelli, Maurilio Gutzeit, Alexandra Carpentier
stat.MLcs.LGarXiv:1605.08671v12016GIANT: Globally Improved Approximate Newton Method for Distributed Optimization
Shusen Wang, Farbod Roosta-Khorasani, Peng Xu +1
cs.LGcs.DCmath.OCarXiv:1709.03528v52017Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
Sean McLeish, Ang Li, John Kirchenbauer +7
cs.CLcs.AIcs.LGarXiv:2511.07384v12025ICE-BeeM: Identifiable Conditional Energy-Based Deep Models Based on Nonlinear ICA
Ilyes Khemakhem, Ricardo Pio Monti, Diederik P. Kingma +1
stat.MLcs.LGarXiv:2002.11537v42020Online Convex Optimization in Adversarial Markov Decision Processes
Aviv Rosenberg, Yishay Mansour
cs.LGcs.AIstat.MLarXiv:1905.07773v12019Training AI Co-Scientists Using Rubric Rewards
Shashwat Goel, Rishi Hazra, Dulhan Jayalath +8
cs.LGcs.CLcs.HCarXiv:2512.23707v12025Agentic Entropy-Balanced Policy Optimization
Guanting Dong, Licheng Bao, Zhongyuan Wang +11
cs.LGcs.AIcs.CLarXiv:2510.14545v12025Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed
Yonggan Fu, Lexington Whalen, Zhifan Ye +11
cs.CLcs.AIcs.LGarXiv:2512.14067v22025Guided Self-Evolving LLMs with Minimal Human Supervision
Wenhao Yu, Zhenwen Liang, Chengsong Huang +4
cs.AIcs.CLcs.LGarXiv:2512.02472v12025Efficient Reinforcement Learning Using Recursive Least-Squares Methods
H. He, D. Hu, X. Xu
cs.LGcs.AIarXiv:1106.0707v12011