Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
17,641 to 17,700 of 20,193
Rethinking the Trust Region in LLM Reinforcement Learning
Penghui Qi, Xiangxin Zhou, Zichen Liu +4
cs.LGcs.AIcs.CLarXiv:2602.04879v32026Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening
Xiaotong Ji, Rasul Tutunov, Matthieu Zimmer +1
cs.LGcs.AIarXiv:2601.21590v12026Continuous Deep Q-Learning with Model-based Acceleration
Shixiang Gu, Timothy Lillicrap, Ilya Sutskever +1
cs.LGcs.AIcs.ROarXiv:1603.00748v12016Meta-learning with differentiable closed-form solvers
Luca Bertinetto, João F. Henriques, Philip H. S. Torr +1
cs.CVcs.LGstat.MLarXiv:1805.08136v32018FinBERT: Financial Sentiment Analysis with Pre-trained Language Models
Dogu Araci
cs.CLcs.LGarXiv:1908.10063v12019Contrastive Learning with Hard Negative Samples
Joshua Robinson, Ching-Yao Chuang, Suvrit Sra +1
cs.LGstat.MLarXiv:2010.04592v22020Rethinking Few-Shot Image Classification: a Good Embedding Is All You Need?
Yonglong Tian, Yue Wang, Dilip Krishnan +2
cs.CVcs.LGarXiv:2003.11539v22020TranAD: Deep Transformer Networks for Anomaly Detection in Multivariate Time Series Data
Shreshth Tuli, Giuliano Casale, Nicholas R. Jennings
cs.LGarXiv:2201.07284v62022Using Self-Supervised Learning Can Improve Model Robustness and Uncertainty
Dan Hendrycks, Mantas Mazeika, Saurav Kadavath +1
cs.LGcs.CVstat.MLarXiv:1906.12340v22019FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization
Chiyu Ma, Shuo Yang, Kexin Huang +7
cs.LGarXiv:2603.19835v32026How to Construct Deep Recurrent Neural Networks
Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho +1
cs.NEcs.LGstat.MLarXiv:1312.6026v52013A survey of loss functions for semantic segmentation
Shruti Jadon
eess.IVcs.CVcs.LGarXiv:2006.14822v42020fastMRI: An Open Dataset and Benchmarks for Accelerated MRI
Jure Zbontar, Florian Knoll, Anuroop Sriram +24
cs.CVcs.LGeess.SParXiv:1811.08839v22018Large Language Model Reasoning Failures
Peiyang Song, Pengrui Han, Noah Goodman
cs.AIcs.CLcs.LGarXiv:2602.06176v12026A review on outlier/anomaly detection in time series data
Ane Blázquez-García, Angel Conde, Usue Mori +1
cs.LGstat.MLarXiv:2002.04236v12020ConViT: Improving Vision Transformers with Soft Convolutional Inductive Biases
Stéphane d'Ascoli, Hugo Touvron, Matthew Leavitt +3
cs.CVcs.LGstat.MLarXiv:2103.10697v22021GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields
Michael Niemeyer, Andreas Geiger
cs.CVcs.LGarXiv:2011.12100v22020QTRAN: Learning to Factorize with Transformation for Cooperative Multi-Agent Reinforcement Learning
Kyunghwan Son, Daewoo Kim, Wan Ju Kang +2
cs.LGcs.AIcs.MAarXiv:1905.05408v12019OmniGAIA: Towards Native Omni-Modal AI Agents
Xiaoxi Li, Wenxiang Jiao, Jiarui Jin +10
cs.AIcs.CLcs.CVarXiv:2602.22897v32026LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
Junyi Zhang, Charles Herrmann, Junhwa Hur +5
cs.CVcs.LGarXiv:2603.03269v22026AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
Dongrui Liu, Qihan Ren, Chen Qian +40
cs.AIcs.CCcs.CLarXiv:2601.18491v22026A Multimodal Anomaly Detector for Robot-Assisted Feeding Using an LSTM-based Variational Autoencoder
Daehyung Park, Yuuna Hoshi, Charles C. Kemp
cs.ROcs.LGarXiv:1711.00614v12017Distributed GraphLab: A Framework for Machine Learning in the Cloud
Yucheng Low, Joseph Gonzalez, Aapo Kyrola +3
cs.DBcs.LGarXiv:1204.6078v12012Sim-to-Real Transfer in Deep Reinforcement Learning for Robotics: a Survey
Wenshuai Zhao, Jorge Peña Queralta, Tomi Westerlund
cs.LGcs.ROarXiv:2009.13303v22020Revisiting the Platonic Representation Hypothesis: An Aristotelian View
Fabian Gröger, Shuo Wen, Maria Brbić
cs.LGcs.AIcs.CVarXiv:2602.14486v22026A Survey of Machine and Deep Learning Methods for Internet of Things (IoT) Security
Mohammed Ali Al-Garadi, Amr Mohamed, Abdulla Al-Ali +2
cs.CRcs.LGcs.NIarXiv:1807.11023v12018Trained Ternary Quantization
Chenzhuo Zhu, Song Han, Huizi Mao +1
cs.LGarXiv:1612.01064v32016MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification
MiroMind Team, S. Bai, L. Bing +41
cs.CLcs.AIcs.IRarXiv:2603.15726v12026UI-Venus-1.5 Technical Report
Venus Team, Changlong Gao, Zhangxuan Gu +24
cs.CVcs.AIcs.CLarXiv:2602.09082v22026TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics
Shirui Chen, Cole Harrison, Ying-Chun Lee +6
cs.ROcs.AIcs.LGarXiv:2602.19313v22026Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs
Haoming Meng, Kexin Huang, Shaohang Wei +6
cs.CLcs.AIcs.LGarXiv:2603.22446v12026Selective Classification for Deep Neural Networks
Yonatan Geifman, Ran El-Yaniv
cs.LGcs.AIarXiv:1705.08500v22017A General Theoretical Paradigm to Understand Learning from Human Preferences
Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot +4
cs.AIcs.LGstat.MLarXiv:2310.12036v22023FFJORD: Free-form Continuous Dynamics for Scalable Reversible Generative Models
Will Grathwohl, Ricky T. Q. Chen, Jesse Bettencourt +2
cs.LGcs.CVstat.MLarXiv:1810.01367v32018Deep Kernel Learning
Andrew Gordon Wilson, Zhiting Hu, Ruslan Salakhutdinov +1
cs.LGcs.AIstat.MEarXiv:1511.02222v12015Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond
Jingfeng Yang, Hongye Jin, Ruixiang Tang +5
cs.CLcs.AIcs.LGarXiv:2304.13712v22023Vector Quantized Diffusion Model for Text-to-Image Synthesis
Shuyang Gu, Dong Chen, Jianmin Bao +5
cs.CVcs.LGarXiv:2111.14822v32021Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Peiyi Wang, Lei Li, Zhihong Shao +6
cs.AIcs.CLcs.LGarXiv:2312.08935v32023IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
Yushi Bai, Qian Dong, Ting Jiang +5
cs.CLcs.LGarXiv:2603.12201v12026Show Your Work: Scratchpads for Intermediate Computation with Language Models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari +9
cs.LGcs.NEarXiv:2112.00114v12021eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers
Yogesh Balaji, Seungjun Nah, Xun Huang +10
cs.CVcs.LGarXiv:2211.01324v52022Endless Terminals: Scaling RL Environments for Terminal Agents
Kanishk Gandhi, Shivam Garg, Noah D. Goodman +1
cs.LGcs.CLarXiv:2601.16443v32026Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution
Rui Cai, Jun Guo, Xinze He +20
cs.ROcs.LGarXiv:2602.12684v22026Video-to-Video Synthesis
Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu +4
cs.CVcs.GRcs.LGarXiv:1808.06601v22018Neural Network Dynamics for Model-Based Deep Reinforcement Learning with Model-Free Fine-Tuning
Anusha Nagabandi, Gregory Kahn, Ronald S. Fearing +1
cs.LGcs.AIcs.ROarXiv:1708.02596v22017Spatial-Temporal Fusion Graph Neural Networks for Traffic Flow Forecasting
Mengzhang Li, Zhanxing Zhu
cs.LGcs.AIarXiv:2012.09641v22020Learning to reinforcement learn
Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala +6
cs.LGcs.AIstat.MLarXiv:1611.05763v32016Multitask learning and benchmarking with clinical time series data
Hrayr Harutyunyan, Hrant Khachatrian, David C. Kale +2
stat.MLcs.LGarXiv:1703.07771v32017Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
Lai Wei, Liangbo He, Jun Lan +9
cs.CVcs.AIcs.CLarXiv:2602.11858v22026Hindsight Credit Assignment for Long-Horizon LLM Agents
Hui-Ze Tan, Xiao-Wen Yang, Hao Chen +7
cs.LGcs.AIarXiv:2603.08754v12026Summaries:简体中文Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
Zhaoyang Wang, Canwen Xu, Boyi Liu +5
cs.AIcs.CLcs.LGarXiv:2602.10090v32026How Much Knowledge Can You Pack Into the Parameters of a Language Model?
Adam Roberts, Colin Raffel, Noam Shazeer
cs.CLcs.LGstat.MLarXiv:2002.08910v42020Deep Packet: A Novel Approach For Encrypted Traffic Classification Using Deep Learning
Mohammad Lotfollahi, Ramin Shirali Hossein Zade, Mahdi Jafari Siavoshani +1
cs.LGcs.CRcs.NIarXiv:1709.02656v32017TS2Vec: Towards Universal Representation of Time Series
Zhihan Yue, Yujing Wang, Juanyong Duan +4
cs.LGcs.AIarXiv:2106.10466v42021Dynamic Filter Networks
Bert De Brabandere, Xu Jia, Tinne Tuytelaars +1
cs.LGcs.CVarXiv:1605.09673v22016Snapshot Ensembles: Train 1, get M for free
Gao Huang, Yixuan Li, Geoff Pleiss +3
cs.LGarXiv:1704.00109v12017Kimi K2.5: Visual Agentic Intelligence
Kimi Team, Tongtong Bai, Yifan Bai +334
cs.CLcs.AIcs.LGarXiv:2602.02276v22026Summaries:한국어Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
Yuqian Fu, Haohuan Huang, Kaiwen Jiang +4
cs.LGcs.AIcs.CLarXiv:2603.25562v22026Towards a Science of AI Agent Reliability
Stephan Rabanser, Sayash Kapoor, Peter Kirgis +3
cs.AIcs.CYcs.LGarXiv:2602.16666v32026Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
Wenkai Yang, Weijie Liu, Ruobing Xie +3
cs.LGcs.AIcs.CLarXiv:2602.12125v22026