Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
1,441 to 1,500 of 20,454
LaPred: Lane-Aware Prediction of Multi-Modal Future Trajectories of Dynamic Agents
ByeoungDo Kim, Seong Hyeon Park, Seokhwan Lee +5
cs.CVcs.LGarXiv:2104.00249v12021MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
Yu Ying Chiu, Michael S. Lee, Rachel Calcott +17
cs.CLcs.AIcs.CYarXiv:2510.16380v22025Remote Labor Index: Measuring AI Automation of Remote Work
Mantas Mazeika, Alice Gatti, Cristina Menghini +44
cs.LGcs.AIcs.CLarXiv:2510.26787v12025VC Classes are Adversarially Robustly Learnable, but Only Improperly
Omar Montasser, Steve Hanneke, Nathan Srebro
cs.LGstat.MLarXiv:1902.04217v22019Learning Unmasking Policies for Diffusion Language Models
Metod Jazbec, Theo X. Olausson, Louis Béthune +6
cs.LGarXiv:2512.09106v42025pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation
Hansheng Chen, Kai Zhang, Hao Tan +3
cs.LGcs.AIcs.CVarXiv:2510.14974v32025OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation
Henry Herzog, Favyen Bastani, Yawen Zhang +23
cs.CVcs.LGarXiv:2511.13655v12025A PAC-Bayesian Tutorial with A Dropout Bound
David McAllester
cs.LGarXiv:1307.2118v12013Building a Foundational Guardrail for General Agentic Systems via Synthetic Data
Yue Huang, Hang Hua, Yujun Zhou +11
cs.LGcs.AIcs.CLarXiv:2510.09781v12025Skeleton Image Representation for 3D Action Recognition based on Tree Structure and Reference Joints
Carlos Caetano, François Brémond, William Robson Schwartz
cs.CVcs.LGarXiv:1909.05704v12019Scaling Behavior of Discrete Diffusion Language Models
Dimitri von Rütte, Janis Fluri, Omead Pooladzandi +3
cs.LGarXiv:2512.10858v32025TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
Zheng Ding, Weirui Ye
cs.LGcs.AIcs.CVarXiv:2512.08153v12025Uniqueness of Low-Rank Matrix Completion by Rigidity Theory
Amit Singer, Mihai Cucuringu
cs.LGarXiv:0902.3846v12009Multi-Scale Dense Networks for Resource Efficient Image Classification
Gao Huang, Danlu Chen, Tianhong Li +3
cs.LGarXiv:1703.09844v52017Parsimonious Black-Box Adversarial Attacks via Efficient Combinatorial Optimization
Seungyong Moon, Gaon An, Hyun Oh Song
cs.LGcs.CRcs.CVarXiv:1905.06635v22019On the Relation Between the Sharpest Directions of DNN Loss and the SGD Step Length
Stanisław Jastrzębski, Zachary Kenton, Nicolas Ballas +3
stat.MLcs.LGarXiv:1807.05031v62018State-Relabeling Adversarial Active Learning
Beichen Zhang, Liang Li, Shijie Yang +3
cs.CVcs.LGarXiv:2004.04943v12020Attention Is All You Need for KV Cache in Diffusion LLMs
Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen
cs.CLcs.AIcs.LGarXiv:2510.14973v22025LongCat-Flash-Omni Technical Report
Meituan LongCat Team, Bairui Wang, Bayan +130
cs.MMcs.AIcs.CLarXiv:2511.00279v22025KeyPose: Multi-View 3D Labeling and Keypoint Estimation for Transparent Objects
Xingyu Liu, Rico Jonschkowski, Anelia Angelova +1
cs.CVcs.LGcs.ROarXiv:1912.02805v22019PAN: A World Model for General, Actionable, and Long-Horizon World Simulation
PAN Team, Zihan Liu, Yi Gu +12
cs.CVcs.AIcs.CLarXiv:2511.09057v42025See, Point, Fly: A Learning-Free VLM Framework for Universal Unmanned Aerial Navigation
Chih Yao Hu, Yang-Sen Lin, Yuna Lee +7
cs.ROcs.AIcs.CLarXiv:2509.22653v12025Semi-Implicit Variational Inference
Mingzhang Yin, Mingyuan Zhou
stat.MLcs.LGstat.COarXiv:1805.11183v12018Defending Against Physically Realizable Attacks on Image Classification
Tong Wu, Liang Tong, Yevgeniy Vorobeychik
cs.LGcs.AIcs.CVarXiv:1909.09552v22019Frequentist Regret Bounds for Randomized Least-Squares Value Iteration
Andrea Zanette, David Brandfonbrener, Emma Brunskill +2
cs.LGstat.MLarXiv:1911.00567v72019Learning Latent Space Energy-Based Prior Model
Bo Pang, Tian Han, Erik Nijkamp +2
stat.MLcs.LGarXiv:2006.08205v22020Tongyi DeepResearch Technical Report
Tongyi DeepResearch Team, Baixuan Li, Bo Zhang +54
cs.CLcs.AIcs.IRarXiv:2510.24701v32025An optimal algorithm for the Thresholding Bandit Problem
Andrea Locatelli, Maurilio Gutzeit, Alexandra Carpentier
stat.MLcs.LGarXiv:1605.08671v12016GIANT: Globally Improved Approximate Newton Method for Distributed Optimization
Shusen Wang, Farbod Roosta-Khorasani, Peng Xu +1
cs.LGcs.DCmath.OCarXiv:1709.03528v52017Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
Sean McLeish, Ang Li, John Kirchenbauer +7
cs.CLcs.AIcs.LGarXiv:2511.07384v12025ICE-BeeM: Identifiable Conditional Energy-Based Deep Models Based on Nonlinear ICA
Ilyes Khemakhem, Ricardo Pio Monti, Diederik P. Kingma +1
stat.MLcs.LGarXiv:2002.11537v42020Online Convex Optimization in Adversarial Markov Decision Processes
Aviv Rosenberg, Yishay Mansour
cs.LGcs.AIstat.MLarXiv:1905.07773v12019Training AI Co-Scientists Using Rubric Rewards
Shashwat Goel, Rishi Hazra, Dulhan Jayalath +8
cs.LGcs.CLcs.HCarXiv:2512.23707v12025Agentic Entropy-Balanced Policy Optimization
Guanting Dong, Licheng Bao, Zhongyuan Wang +11
cs.LGcs.AIcs.CLarXiv:2510.14545v12025Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed
Yonggan Fu, Lexington Whalen, Zhifan Ye +11
cs.CLcs.AIcs.LGarXiv:2512.14067v22025Guided Self-Evolving LLMs with Minimal Human Supervision
Wenhao Yu, Zhenwen Liang, Chengsong Huang +4
cs.AIcs.CLcs.LGarXiv:2512.02472v12025Efficient Reinforcement Learning Using Recursive Least-Squares Methods
H. He, D. Hu, X. Xu
cs.LGcs.AIarXiv:1106.0707v12011Does FLUX Already Know How to Perform Physically Plausible Image Composition?
Shilin Lu, Zhuming Lian, Zihan Zhou +3
cs.CVcs.AIcs.LGarXiv:2509.21278v42025Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
Xin Qiu, Yulu Gan, Conor F. Hayes +6
cs.LGcs.AIcs.NEarXiv:2509.24372v32025VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models
Guochao Jiang, Wenfeng Feng, Guofeng Quan +4
cs.LGcs.CLarXiv:2509.19803v12025Search Self-play: Pushing the Frontier of Agent Capability without Supervision
Hongliang Lu, Yuhang Wen, Pengyu Cheng +7
cs.LGarXiv:2510.18821v32025Multi-view Vector-valued Manifold Regularization for Multi-label Image Classification
Yong Luo, Dacheng Tao, Chang Xu +3
stat.MLcs.CVcs.LGarXiv:1904.03921v12019Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
Nikita Kachaev, Mikhail Kolosov, Daniil Zelezetsky +2
cs.LGcs.AIcs.ROarXiv:2510.25616v12025SimpleFold: Folding Proteins is Simpler than You Think
Yuyang Wang, Jiarui Lu, Navdeep Jaitly +2
cs.LGq-bio.QMarXiv:2509.18480v42025Deep Residual Learning in the JPEG Transform Domain
Max Ehrlich, Larry Davis
cs.LGcs.CVstat.MLarXiv:1812.11690v32018A Review of the Gumbel-max Trick and its Extensions for Discrete Stochasticity in Machine Learning
Iris A. M. Huijben, Wouter Kool, Max B. Paulus +1
cs.LGstat.MLarXiv:2110.01515v22021Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents
Guoqing Wang, Sunhao Dai, Guangze Ye +5
cs.CLcs.AIcs.LGarXiv:2510.14967v22025Stable Gaussian Process based Tracking Control of Euler-Lagrange Systems
Thomas Beckers, Dana Kulić, Sandra Hirche
cs.LGeess.SYstat.MLarXiv:1806.07190v22018DragFlow: Unleashing DiT Priors with Region Based Supervision for Drag Editing
Zihan Zhou, Shilin Lu, Shuli Leng +4
cs.CVcs.AIcs.LGarXiv:2510.02253v32025Algorithms for Dynamic Spectrum Access with Learning for Cognitive Radio
Jayakrishnan Unnikrishnan, Venugopal Veeravalli
cs.NIcs.LGarXiv:0807.2677v42008Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
Boxin Wang, Chankyu Lee, Nayeon Lee +9
cs.CLcs.AIcs.LGarXiv:2512.13607v22025ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
Hongjin Su, Shizhe Diao, Ximing Lu +13
cs.CLcs.AIcs.LGarXiv:2511.21689v12025RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
Zhiyuan Zeng, Hamish Ivison, Yiping Wang +14
cs.CLcs.LGarXiv:2511.07317v22025Eliciting Secret Knowledge from Language Models
Bartosz Cywiński, Emil Ryd, Rowan Wang +4
cs.LGarXiv:2510.01070v22025Patient2Vec: A Personalized Interpretable Deep Representation of the Longitudinal Electronic Health Record
Jinghe Zhang, Kamran Kowsari, James H. Harrison +2
q-bio.QMcs.AIcs.IRarXiv:1810.04793v32018Decision Trees for Decision-Making under the Predict-then-Optimize Framework
Adam N. Elmachtoub, Jason Cheuk Nam Liang, Ryan McNellis
cs.LGmath.OCstat.MLarXiv:2003.00360v22020LFM2 Technical Report
Alexander Amini, Anna Banaszak, Harold Benoit +30
cs.LGcs.AIarXiv:2511.23404v12025Kalman Filtering with Intermittent Observations: Weak Convergence to a Stationary Distribution
Soummya Kar, Bruno Sinopoli, Jose M. F. Moura
cs.ITcs.LGmath.STarXiv:0903.2890v22009FlowRL: Matching Reward Distributions for LLM Reasoning
Xuekai Zhu, Daixuan Cheng, Dinghuai Zhang +20
cs.LGcs.AIcs.CLarXiv:2509.15207v32025Monotonic Calibrated Interpolated Look-Up Tables
Maya Gupta, Andrew Cotter, Jan Pfeifer +5
cs.LGarXiv:1505.06378v32015