Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

19,021 to 19,080 of 20,193

  1. Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles

    Jinyang Wu, Guocheng Zhai, Ruihan Jin +7

    cs.LGcs.CLarXiv:2605.22177v12026
  2. The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation

    Yifan Lan, Yuanpu Cao, Hanyu Wang +2

    cs.LGcs.AIarXiv:2605.21856v12026
  3. ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison

    Tianle Li, Xuyang Shen, Yan Ma +7

    cs.LGcs.AIcs.CVarXiv:2605.20278v22026
  4. PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents

    Zhuohan Gu, Qizheng Zhang, Omar Khattab +1

    cs.AIcs.CLcs.LGarXiv:2605.19932v12026
  5. Nexus : An Agentic Framework for Time Series Forecasting

    Sarkar Snigdha Sarathi Das, Palash Goyal, Mihir Parmar +6

    cs.AIcs.CLcs.LGarXiv:2605.14389v12026
  6. EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents

    Jiaqi Liu, Xinyu Ye, Peng Xia +4

    cs.LGcs.AIarXiv:2605.13941v12026
  7. Learning, Fast and Slow: Towards LLMs That Adapt Continually

    Rishabh Tiwari, Kusha Sareen, Lakshya A Agrawal +6

    cs.LGcs.AIarXiv:2605.12484v22026
  8. Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple Silicon

    Víctor Gallego

    cs.LGcs.AIcs.DCarXiv:2605.09708v22026
  9. Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction

    Ngoc Bui, Hieu Trung Nguyen, Arman Cohan +1

    cs.LGarXiv:2605.09649v12026
  10. RewardHarness: Self-Evolving Agentic Post-Training

    Yuxuan Zhang, Penghui Du, Bo Li +11

    cs.AIcs.CLcs.CVarXiv:2605.08703v12026
  11. Large Language Models over Networks: Collaborative Intelligence under Resource Constraints

    Liangqi Yuan, Wenzhi Fang, Shiqiang Wang +2

    eess.SPcs.DCcs.LGarXiv:2605.08626v12026
  12. STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation

    Ying Shen, Tianrong Chen, Yuan Gao +6

    cs.CVcs.LGarXiv:2605.08029v12026
  13. HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents

    Guankai Li, Jiabin Chen, Yi Xu +2

    cs.LGcs.AIarXiv:2605.07177v22026
  14. ModelLens: Finding the Best for Your Task from Myriads of Models

    Rui Cai, Weijie Jacky Mo, Xiaofei Wen +5

    cs.LGarXiv:2605.07075v12026
  15. RaguTeam at SemEval-2026 Task 8: Meno and Friends in a Judge-Orchestrated LLM Ensemble for Faithful Multi-Turn Response Generation

    Ivan Bondarenko, Roman Derunets, Oleg Sedukhin +3

    cs.CLcs.AIcs.LGarXiv:2605.04523v12026
  16. CGM-JEPA: Learning Consistent Continuous Glucose Monitor Representations via Predictive Self-Supervised Pretraining

    Hada Melino Muhammad, Zechen Li, Flora Salim +1

    cs.LGcs.AIarXiv:2605.00933v12026
  17. Scaling Continual Learning to 300+ Tasks with Bi-Level Routing Mixture-of-Experts

    Meng Lou, Yunxiang Fu, Yizhou Yu

    cs.LGcs.CVarXiv:2602.03473v22026
  18. Understanding the Behaviour of Contrastive Loss

    Feng Wang, Huaping Liu

    cs.LGarXiv:2012.09740v22020
    Summaries:한국어
  19. ProGen: Language Modeling for Protein Generation

    Ali Madani, Bryan McCann, Nikhil Naik +5

    q-bio.BMcs.LGstat.MLarXiv:2004.03497v12020
  20. nuScenes: A multimodal dataset for autonomous driving

    Holger Caesar, Varun Bankiti, Alex H. Lang +7

    cs.LGcs.CVcs.ROarXiv:1903.11027v52019
  21. Parameter-Efficient Transfer Learning for NLP

    Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski +5

    cs.LGcs.CLstat.MLarXiv:1902.00751v22019
  22. On Calibration of Modern Neural Networks

    Chuan Guo, Geoff Pleiss, Yu Sun +1

    cs.LGarXiv:1706.04599v22017
  23. Asynchronous Methods for Deep Reinforcement Learning

    Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza +5

    cs.LGarXiv:1602.01783v22016
  24. Striving for Simplicity: The All Convolutional Net

    Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox +1

    cs.LGcs.CVcs.NEarXiv:1412.6806v32014
  25. Diffusion-Pretrained Dense and Contextual Embeddings

    Sedigheh Eslami, Maksim Gaiduk, Markus Krimmel +3

    cs.LGcs.CLcs.IRarXiv:2602.11151v22026
    Summaries:한국어
  26. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

    Zhihong Shao, Peiyi Wang, Qihao Zhu +8

    cs.CLcs.AIcs.LGarXiv:2402.03300v32024
  27. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

    Anthony Brohan, Noah Brown, Justice Carbajal +51

    cs.ROcs.CLcs.CVarXiv:2307.15818v12023
  28. Robust Speech Recognition via Large-Scale Weak Supervision

    Alec Radford, Jong Wook Kim, Tao Xu +3

    eess.AScs.CLcs.LGarXiv:2212.04356v12022
  29. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers

    Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong +1

    cs.LGcs.AIarXiv:2211.14730v22022
  30. Large Language Models are Zero-Shot Reasoners

    Takeshi Kojima, Shixiang Shane Gu, Machel Reid +2

    cs.CLcs.AIcs.LGarXiv:2205.11916v42022
  31. Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

    Yuntao Bai, Andy Jones, Kamal Ndousse +28

    cs.CLcs.LGarXiv:2204.05862v12022
  32. TruthfulQA: Measuring How Models Mimic Human Falsehoods

    Stephanie Lin, Jacob Hilton, Owain Evans

    cs.CLcs.AIcs.CYarXiv:2109.07958v22021
  33. Barlow Twins: Self-Supervised Learning via Redundancy Reduction

    Jure Zbontar, Li Jing, Ishan Misra +2

    cs.CVcs.AIcs.LGarXiv:2103.03230v32021
  34. Learning Transferable Visual Models From Natural Language Supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy +9

    cs.CVcs.LGarXiv:2103.00020v12021
    Summaries:한국어
  35. Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation

    David M. W. Powers

    cs.LGstat.MEstat.MLarXiv:2010.16061v12020
  36. Supervised Contrastive Learning

    Prannay Khosla, Piotr Teterwak, Chen Wang +6

    cs.LGcs.CVstat.MLarXiv:2004.11362v52020
  37. Don't Stop Pretraining: Adapt Language Models to Domains and Tasks

    Suchin Gururangan, Ana Marasović, Swabha Swayamdipta +4

    cs.CLcs.LGarXiv:2004.10964v32020
  38. A Simple Framework for Contrastive Learning of Visual Representations

    Ting Chen, Simon Kornblith, Mohammad Norouzi +1

    cs.LGcs.CVstat.MLarXiv:2002.05709v32020
  39. Generative Modeling by Estimating Gradients of the Data Distribution

    Yang Song, Stefano Ermon

    cs.LGstat.MLarXiv:1907.05600v32019
  40. Deep learning in agriculture: A survey

    Andreas Kamilaris, Francesc X. Prenafeta-Boldu

    cs.LGcs.CVstat.MLarXiv:1807.11809v12018
  41. Glow: Generative Flow with Invertible 1x1 Convolutions

    Diederik P. Kingma, Prafulla Dhariwal

    stat.MLcs.AIcs.LGarXiv:1807.03039v22018
  42. An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling

    Shaojie Bai, J. Zico Kolter, Vladlen Koltun

    cs.LGcs.AIcs.CLarXiv:1803.01271v22018
  43. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor

    Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel +1

    cs.LGcs.AIstat.MLarXiv:1801.01290v22018
  44. Convolutional 2D Knowledge Graph Embeddings

    Tim Dettmers, Pasquale Minervini, Pontus Stenetorp +1

    cs.LGarXiv:1707.01476v62017
  45. Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World

    Josh Tobin, Rachel Fong, Alex Ray +3

    cs.ROcs.LGarXiv:1703.06907v12017
  46. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks

    Chelsea Finn, Pieter Abbeel, Sergey Levine

    cs.LGcs.AIcs.CVarXiv:1703.03400v32017
  47. On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima

    Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal +2

    cs.LGmath.OCarXiv:1609.04836v22016
  48. XGBoost: A Scalable Tree Boosting System

    Tianqi Chen, Carlos Guestrin

    cs.LGarXiv:1603.02754v32016
    Summaries:한국어
  49. Prioritized Experience Replay

    Tom Schaul, John Quan, Ioannis Antonoglou +1

    cs.LGarXiv:1511.05952v42015
  50. Deep Unsupervised Learning using Nonequilibrium Thermodynamics

    Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan +1

    cs.LGcond-mat.dis-nnq-bio.NCarXiv:1503.03585v82015
  51. A Tutorial on Spectral Clustering

    Ulrike von Luxburg

    cs.DScs.LGarXiv:0711.0189v12007
  52. S-Bus: Automatic Read-Set Reconstruction for Multi-Agent LLM State Coordination

    Sajjad Khan

    cs.LGcs.AIcs.DCarXiv:2605.17076v22026
  53. Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning

    Jiawei Wang, Ke Rui, Yushen Zuo +2

    cs.LGcs.CVarXiv:2608.18746v12026
  54. An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models

    Javier Aguilar Martín

    cs.LGcs.AIeess.SYarXiv:2608.17956v12026
  55. Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments

    Adam Karvonen, Euan Ong, Subhash Kantamneni +1

    cs.LGcs.AIarXiv:2608.16747v12026
  56. Topology-Preserving Neural Operator Learning via Hodge Decomposition

    Dongzhe Zheng, Tao Zhong, Christine Allen-Blanchette

    cs.LGcs.AIcs.CGarXiv:2605.13834v22026
  57. Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia

    Xiang Guan, Roger D. Newman-Norlund, Yong Yang +8

    cs.LGcs.CLarXiv:2608.12717v12026
  58. Bitcoin Price Direction Prediction via Regime-Aware Multi-Modal Fusion of Social Sentiment and Technical Features

    Muhammad Abdullah Haroon

    cs.LGcs.CEecon.EMarXiv:2607.23370v12026
  59. RLDX-1 Technical Report

    Dongyoung Kim, Huiwon Jang, Myungkyu Koo +65

    cs.ROcs.AIcs.LGarXiv:2605.03269v22026
    Summaries:한국어
  60. Adaptive Heterogeneous Compression for Resource-Efficient Federated Knowledge Distillation

    Chenwang Liu, Yijun Liu, Chang Liu +2

    cs.DCcs.LGarXiv:2608.15660v12026