Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
19,021 to 19,080 of 20,193
Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles
Jinyang Wu, Guocheng Zhai, Ruihan Jin +7
cs.LGcs.CLarXiv:2605.22177v12026The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation
Yifan Lan, Yuanpu Cao, Hanyu Wang +2
cs.LGcs.AIarXiv:2605.21856v12026ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison
Tianle Li, Xuyang Shen, Yan Ma +7
cs.LGcs.AIcs.CVarXiv:2605.20278v22026PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents
Zhuohan Gu, Qizheng Zhang, Omar Khattab +1
cs.AIcs.CLcs.LGarXiv:2605.19932v12026Nexus : An Agentic Framework for Time Series Forecasting
Sarkar Snigdha Sarathi Das, Palash Goyal, Mihir Parmar +6
cs.AIcs.CLcs.LGarXiv:2605.14389v12026EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents
Jiaqi Liu, Xinyu Ye, Peng Xia +4
cs.LGcs.AIarXiv:2605.13941v12026Learning, Fast and Slow: Towards LLMs That Adapt Continually
Rishabh Tiwari, Kusha Sareen, Lakshya A Agrawal +6
cs.LGcs.AIarXiv:2605.12484v22026Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple Silicon
Víctor Gallego
cs.LGcs.AIcs.DCarXiv:2605.09708v22026Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction
Ngoc Bui, Hieu Trung Nguyen, Arman Cohan +1
cs.LGarXiv:2605.09649v12026RewardHarness: Self-Evolving Agentic Post-Training
Yuxuan Zhang, Penghui Du, Bo Li +11
cs.AIcs.CLcs.CVarXiv:2605.08703v12026Large Language Models over Networks: Collaborative Intelligence under Resource Constraints
Liangqi Yuan, Wenzhi Fang, Shiqiang Wang +2
eess.SPcs.DCcs.LGarXiv:2605.08626v12026STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation
Ying Shen, Tianrong Chen, Yuan Gao +6
cs.CVcs.LGarXiv:2605.08029v12026HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents
Guankai Li, Jiabin Chen, Yi Xu +2
cs.LGcs.AIarXiv:2605.07177v22026ModelLens: Finding the Best for Your Task from Myriads of Models
Rui Cai, Weijie Jacky Mo, Xiaofei Wen +5
cs.LGarXiv:2605.07075v12026RaguTeam at SemEval-2026 Task 8: Meno and Friends in a Judge-Orchestrated LLM Ensemble for Faithful Multi-Turn Response Generation
Ivan Bondarenko, Roman Derunets, Oleg Sedukhin +3
cs.CLcs.AIcs.LGarXiv:2605.04523v12026CGM-JEPA: Learning Consistent Continuous Glucose Monitor Representations via Predictive Self-Supervised Pretraining
Hada Melino Muhammad, Zechen Li, Flora Salim +1
cs.LGcs.AIarXiv:2605.00933v12026Scaling Continual Learning to 300+ Tasks with Bi-Level Routing Mixture-of-Experts
Meng Lou, Yunxiang Fu, Yizhou Yu
cs.LGcs.CVarXiv:2602.03473v22026Understanding the Behaviour of Contrastive Loss
Feng Wang, Huaping Liu
cs.LGarXiv:2012.09740v22020Summaries:한국어ProGen: Language Modeling for Protein Generation
Ali Madani, Bryan McCann, Nikhil Naik +5
q-bio.BMcs.LGstat.MLarXiv:2004.03497v12020nuScenes: A multimodal dataset for autonomous driving
Holger Caesar, Varun Bankiti, Alex H. Lang +7
cs.LGcs.CVcs.ROarXiv:1903.11027v52019Parameter-Efficient Transfer Learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski +5
cs.LGcs.CLstat.MLarXiv:1902.00751v22019On Calibration of Modern Neural Networks
Chuan Guo, Geoff Pleiss, Yu Sun +1
cs.LGarXiv:1706.04599v22017Asynchronous Methods for Deep Reinforcement Learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza +5
cs.LGarXiv:1602.01783v22016Striving for Simplicity: The All Convolutional Net
Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox +1
cs.LGcs.CVcs.NEarXiv:1412.6806v32014Diffusion-Pretrained Dense and Contextual Embeddings
Sedigheh Eslami, Maksim Gaiduk, Markus Krimmel +3
cs.LGcs.CLcs.IRarXiv:2602.11151v22026Summaries:한국어DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Zhihong Shao, Peiyi Wang, Qihao Zhu +8
cs.CLcs.AIcs.LGarXiv:2402.03300v32024RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Anthony Brohan, Noah Brown, Justice Carbajal +51
cs.ROcs.CLcs.CVarXiv:2307.15818v12023Robust Speech Recognition via Large-Scale Weak Supervision
Alec Radford, Jong Wook Kim, Tao Xu +3
eess.AScs.CLcs.LGarXiv:2212.04356v12022A Time Series is Worth 64 Words: Long-term Forecasting with Transformers
Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong +1
cs.LGcs.AIarXiv:2211.14730v22022Large Language Models are Zero-Shot Reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid +2
cs.CLcs.AIcs.LGarXiv:2205.11916v42022Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Yuntao Bai, Andy Jones, Kamal Ndousse +28
cs.CLcs.LGarXiv:2204.05862v12022TruthfulQA: Measuring How Models Mimic Human Falsehoods
Stephanie Lin, Jacob Hilton, Owain Evans
cs.CLcs.AIcs.CYarXiv:2109.07958v22021Barlow Twins: Self-Supervised Learning via Redundancy Reduction
Jure Zbontar, Li Jing, Ishan Misra +2
cs.CVcs.AIcs.LGarXiv:2103.03230v32021Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy +9
cs.CVcs.LGarXiv:2103.00020v12021Summaries:한국어Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation
David M. W. Powers
cs.LGstat.MEstat.MLarXiv:2010.16061v12020Supervised Contrastive Learning
Prannay Khosla, Piotr Teterwak, Chen Wang +6
cs.LGcs.CVstat.MLarXiv:2004.11362v52020Don't Stop Pretraining: Adapt Language Models to Domains and Tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta +4
cs.CLcs.LGarXiv:2004.10964v32020A Simple Framework for Contrastive Learning of Visual Representations
Ting Chen, Simon Kornblith, Mohammad Norouzi +1
cs.LGcs.CVstat.MLarXiv:2002.05709v32020Generative Modeling by Estimating Gradients of the Data Distribution
Yang Song, Stefano Ermon
cs.LGstat.MLarXiv:1907.05600v32019Deep learning in agriculture: A survey
Andreas Kamilaris, Francesc X. Prenafeta-Boldu
cs.LGcs.CVstat.MLarXiv:1807.11809v12018Glow: Generative Flow with Invertible 1x1 Convolutions
Diederik P. Kingma, Prafulla Dhariwal
stat.MLcs.AIcs.LGarXiv:1807.03039v22018An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling
Shaojie Bai, J. Zico Kolter, Vladlen Koltun
cs.LGcs.AIcs.CLarXiv:1803.01271v22018Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel +1
cs.LGcs.AIstat.MLarXiv:1801.01290v22018Convolutional 2D Knowledge Graph Embeddings
Tim Dettmers, Pasquale Minervini, Pontus Stenetorp +1
cs.LGarXiv:1707.01476v62017Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World
Josh Tobin, Rachel Fong, Alex Ray +3
cs.ROcs.LGarXiv:1703.06907v12017Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
Chelsea Finn, Pieter Abbeel, Sergey Levine
cs.LGcs.AIcs.CVarXiv:1703.03400v32017On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal +2
cs.LGmath.OCarXiv:1609.04836v22016XGBoost: A Scalable Tree Boosting System
Tianqi Chen, Carlos Guestrin
cs.LGarXiv:1603.02754v32016Summaries:한국어Prioritized Experience Replay
Tom Schaul, John Quan, Ioannis Antonoglou +1
cs.LGarXiv:1511.05952v42015Deep Unsupervised Learning using Nonequilibrium Thermodynamics
Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan +1
cs.LGcond-mat.dis-nnq-bio.NCarXiv:1503.03585v82015A Tutorial on Spectral Clustering
Ulrike von Luxburg
cs.DScs.LGarXiv:0711.0189v12007S-Bus: Automatic Read-Set Reconstruction for Multi-Agent LLM State Coordination
Sajjad Khan
cs.LGcs.AIcs.DCarXiv:2605.17076v22026Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning
Jiawei Wang, Ke Rui, Yushen Zuo +2
cs.LGcs.CVarXiv:2608.18746v12026An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models
Javier Aguilar Martín
cs.LGcs.AIeess.SYarXiv:2608.17956v12026Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments
Adam Karvonen, Euan Ong, Subhash Kantamneni +1
cs.LGcs.AIarXiv:2608.16747v12026Topology-Preserving Neural Operator Learning via Hodge Decomposition
Dongzhe Zheng, Tao Zhong, Christine Allen-Blanchette
cs.LGcs.AIcs.CGarXiv:2605.13834v22026Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia
Xiang Guan, Roger D. Newman-Norlund, Yong Yang +8
cs.LGcs.CLarXiv:2608.12717v12026Bitcoin Price Direction Prediction via Regime-Aware Multi-Modal Fusion of Social Sentiment and Technical Features
Muhammad Abdullah Haroon
cs.LGcs.CEecon.EMarXiv:2607.23370v12026RLDX-1 Technical Report
Dongyoung Kim, Huiwon Jang, Myungkyu Koo +65
cs.ROcs.AIcs.LGarXiv:2605.03269v22026Summaries:한국어Adaptive Heterogeneous Compression for Resource-Efficient Federated Knowledge Distillation
Chenwang Liu, Yijun Liu, Chang Liu +2
cs.DCcs.LGarXiv:2608.15660v12026