Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

17,641 to 17,700 of 20,193

  1. Rethinking the Trust Region in LLM Reinforcement Learning

    Penghui Qi, Xiangxin Zhou, Zichen Liu +4

    cs.LGcs.AIcs.CLarXiv:2602.04879v32026
  2. Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening

    Xiaotong Ji, Rasul Tutunov, Matthieu Zimmer +1

    cs.LGcs.AIarXiv:2601.21590v12026
  3. Continuous Deep Q-Learning with Model-based Acceleration

    Shixiang Gu, Timothy Lillicrap, Ilya Sutskever +1

    cs.LGcs.AIcs.ROarXiv:1603.00748v12016
  4. Meta-learning with differentiable closed-form solvers

    Luca Bertinetto, João F. Henriques, Philip H. S. Torr +1

    cs.CVcs.LGstat.MLarXiv:1805.08136v32018
  5. FinBERT: Financial Sentiment Analysis with Pre-trained Language Models

    Dogu Araci

    cs.CLcs.LGarXiv:1908.10063v12019
  6. Contrastive Learning with Hard Negative Samples

    Joshua Robinson, Ching-Yao Chuang, Suvrit Sra +1

    cs.LGstat.MLarXiv:2010.04592v22020
  7. Rethinking Few-Shot Image Classification: a Good Embedding Is All You Need?

    Yonglong Tian, Yue Wang, Dilip Krishnan +2

    cs.CVcs.LGarXiv:2003.11539v22020
  8. TranAD: Deep Transformer Networks for Anomaly Detection in Multivariate Time Series Data

    Shreshth Tuli, Giuliano Casale, Nicholas R. Jennings

    cs.LGarXiv:2201.07284v62022
  9. Using Self-Supervised Learning Can Improve Model Robustness and Uncertainty

    Dan Hendrycks, Mantas Mazeika, Saurav Kadavath +1

    cs.LGcs.CVstat.MLarXiv:1906.12340v22019
  10. FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization

    Chiyu Ma, Shuo Yang, Kexin Huang +7

    cs.LGarXiv:2603.19835v32026
  11. How to Construct Deep Recurrent Neural Networks

    Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho +1

    cs.NEcs.LGstat.MLarXiv:1312.6026v52013
  12. A survey of loss functions for semantic segmentation

    Shruti Jadon

    eess.IVcs.CVcs.LGarXiv:2006.14822v42020
  13. fastMRI: An Open Dataset and Benchmarks for Accelerated MRI

    Jure Zbontar, Florian Knoll, Anuroop Sriram +24

    cs.CVcs.LGeess.SParXiv:1811.08839v22018
  14. Large Language Model Reasoning Failures

    Peiyang Song, Pengrui Han, Noah Goodman

    cs.AIcs.CLcs.LGarXiv:2602.06176v12026
  15. A review on outlier/anomaly detection in time series data

    Ane Blázquez-García, Angel Conde, Usue Mori +1

    cs.LGstat.MLarXiv:2002.04236v12020
  16. ConViT: Improving Vision Transformers with Soft Convolutional Inductive Biases

    Stéphane d'Ascoli, Hugo Touvron, Matthew Leavitt +3

    cs.CVcs.LGstat.MLarXiv:2103.10697v22021
  17. GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields

    Michael Niemeyer, Andreas Geiger

    cs.CVcs.LGarXiv:2011.12100v22020
  18. QTRAN: Learning to Factorize with Transformation for Cooperative Multi-Agent Reinforcement Learning

    Kyunghwan Son, Daewoo Kim, Wan Ju Kang +2

    cs.LGcs.AIcs.MAarXiv:1905.05408v12019
  19. OmniGAIA: Towards Native Omni-Modal AI Agents

    Xiaoxi Li, Wenxiang Jiao, Jiarui Jin +10

    cs.AIcs.CLcs.CVarXiv:2602.22897v32026
  20. LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory

    Junyi Zhang, Charles Herrmann, Junhwa Hur +5

    cs.CVcs.LGarXiv:2603.03269v22026
  21. AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security

    Dongrui Liu, Qihan Ren, Chen Qian +40

    cs.AIcs.CCcs.CLarXiv:2601.18491v22026
  22. A Multimodal Anomaly Detector for Robot-Assisted Feeding Using an LSTM-based Variational Autoencoder

    Daehyung Park, Yuuna Hoshi, Charles C. Kemp

    cs.ROcs.LGarXiv:1711.00614v12017
  23. Distributed GraphLab: A Framework for Machine Learning in the Cloud

    Yucheng Low, Joseph Gonzalez, Aapo Kyrola +3

    cs.DBcs.LGarXiv:1204.6078v12012
  24. Sim-to-Real Transfer in Deep Reinforcement Learning for Robotics: a Survey

    Wenshuai Zhao, Jorge Peña Queralta, Tomi Westerlund

    cs.LGcs.ROarXiv:2009.13303v22020
  25. Revisiting the Platonic Representation Hypothesis: An Aristotelian View

    Fabian Gröger, Shuo Wen, Maria Brbić

    cs.LGcs.AIcs.CVarXiv:2602.14486v22026
  26. A Survey of Machine and Deep Learning Methods for Internet of Things (IoT) Security

    Mohammed Ali Al-Garadi, Amr Mohamed, Abdulla Al-Ali +2

    cs.CRcs.LGcs.NIarXiv:1807.11023v12018
  27. Trained Ternary Quantization

    Chenzhuo Zhu, Song Han, Huizi Mao +1

    cs.LGarXiv:1612.01064v32016
  28. MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification

    MiroMind Team, S. Bai, L. Bing +41

    cs.CLcs.AIcs.IRarXiv:2603.15726v12026
  29. UI-Venus-1.5 Technical Report

    Venus Team, Changlong Gao, Zhangxuan Gu +24

    cs.CVcs.AIcs.CLarXiv:2602.09082v22026
  30. TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics

    Shirui Chen, Cole Harrison, Ying-Chun Lee +6

    cs.ROcs.AIcs.LGarXiv:2602.19313v22026
  31. Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs

    Haoming Meng, Kexin Huang, Shaohang Wei +6

    cs.CLcs.AIcs.LGarXiv:2603.22446v12026
  32. Selective Classification for Deep Neural Networks

    Yonatan Geifman, Ran El-Yaniv

    cs.LGcs.AIarXiv:1705.08500v22017
  33. A General Theoretical Paradigm to Understand Learning from Human Preferences

    Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot +4

    cs.AIcs.LGstat.MLarXiv:2310.12036v22023
  34. FFJORD: Free-form Continuous Dynamics for Scalable Reversible Generative Models

    Will Grathwohl, Ricky T. Q. Chen, Jesse Bettencourt +2

    cs.LGcs.CVstat.MLarXiv:1810.01367v32018
  35. Deep Kernel Learning

    Andrew Gordon Wilson, Zhiting Hu, Ruslan Salakhutdinov +1

    cs.LGcs.AIstat.MEarXiv:1511.02222v12015
  36. Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond

    Jingfeng Yang, Hongye Jin, Ruixiang Tang +5

    cs.CLcs.AIcs.LGarXiv:2304.13712v22023
  37. Vector Quantized Diffusion Model for Text-to-Image Synthesis

    Shuyang Gu, Dong Chen, Jianmin Bao +5

    cs.CVcs.LGarXiv:2111.14822v32021
  38. Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

    Peiyi Wang, Lei Li, Zhihong Shao +6

    cs.AIcs.CLcs.LGarXiv:2312.08935v32023
  39. IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse

    Yushi Bai, Qian Dong, Ting Jiang +5

    cs.CLcs.LGarXiv:2603.12201v12026
  40. Show Your Work: Scratchpads for Intermediate Computation with Language Models

    Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari +9

    cs.LGcs.NEarXiv:2112.00114v12021
  41. eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

    Yogesh Balaji, Seungjun Nah, Xun Huang +10

    cs.CVcs.LGarXiv:2211.01324v52022
  42. Endless Terminals: Scaling RL Environments for Terminal Agents

    Kanishk Gandhi, Shivam Garg, Noah D. Goodman +1

    cs.LGcs.CLarXiv:2601.16443v32026
  43. Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution

    Rui Cai, Jun Guo, Xinze He +20

    cs.ROcs.LGarXiv:2602.12684v22026
  44. Video-to-Video Synthesis

    Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu +4

    cs.CVcs.GRcs.LGarXiv:1808.06601v22018
  45. Neural Network Dynamics for Model-Based Deep Reinforcement Learning with Model-Free Fine-Tuning

    Anusha Nagabandi, Gregory Kahn, Ronald S. Fearing +1

    cs.LGcs.AIcs.ROarXiv:1708.02596v22017
  46. Spatial-Temporal Fusion Graph Neural Networks for Traffic Flow Forecasting

    Mengzhang Li, Zhanxing Zhu

    cs.LGcs.AIarXiv:2012.09641v22020
  47. Learning to reinforcement learn

    Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala +6

    cs.LGcs.AIstat.MLarXiv:1611.05763v32016
  48. Multitask learning and benchmarking with clinical time series data

    Hrayr Harutyunyan, Hrant Khachatrian, David C. Kale +2

    stat.MLcs.LGarXiv:1703.07771v32017
  49. Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception

    Lai Wei, Liangbo He, Jun Lan +9

    cs.CVcs.AIcs.CLarXiv:2602.11858v22026
  50. Hindsight Credit Assignment for Long-Horizon LLM Agents

    Hui-Ze Tan, Xiao-Wen Yang, Hao Chen +7

    cs.LGcs.AIarXiv:2603.08754v12026
    Summaries:简体中文
  51. Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning

    Zhaoyang Wang, Canwen Xu, Boyi Liu +5

    cs.AIcs.CLcs.LGarXiv:2602.10090v32026
  52. How Much Knowledge Can You Pack Into the Parameters of a Language Model?

    Adam Roberts, Colin Raffel, Noam Shazeer

    cs.CLcs.LGstat.MLarXiv:2002.08910v42020
  53. Deep Packet: A Novel Approach For Encrypted Traffic Classification Using Deep Learning

    Mohammad Lotfollahi, Ramin Shirali Hossein Zade, Mahdi Jafari Siavoshani +1

    cs.LGcs.CRcs.NIarXiv:1709.02656v32017
  54. TS2Vec: Towards Universal Representation of Time Series

    Zhihan Yue, Yujing Wang, Juanyong Duan +4

    cs.LGcs.AIarXiv:2106.10466v42021
  55. Dynamic Filter Networks

    Bert De Brabandere, Xu Jia, Tinne Tuytelaars +1

    cs.LGcs.CVarXiv:1605.09673v22016
  56. Snapshot Ensembles: Train 1, get M for free

    Gao Huang, Yixuan Li, Geoff Pleiss +3

    cs.LGarXiv:1704.00109v12017
  57. Kimi K2.5: Visual Agentic Intelligence

    Kimi Team, Tongtong Bai, Yifan Bai +334

    cs.CLcs.AIcs.LGarXiv:2602.02276v22026
    Summaries:한국어
  58. Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

    Yuqian Fu, Haohuan Huang, Kaiwen Jiang +4

    cs.LGcs.AIcs.CLarXiv:2603.25562v22026
  59. Towards a Science of AI Agent Reliability

    Stephan Rabanser, Sayash Kapoor, Peter Kirgis +3

    cs.AIcs.CYcs.LGarXiv:2602.16666v32026
  60. Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

    Wenkai Yang, Weijie Liu, Ruobing Xie +3

    cs.LGcs.AIcs.CLarXiv:2602.12125v22026