Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

5,221 to 5,280 of 20,454

  1. CRPO: A New Approach for Safe Reinforcement Learning with Convergence Guarantee

    Tengyu Xu, Yingbin Liang, Guanghui Lan

    cs.LGstat.MLarXiv:2011.05869v32020
  2. Are Diffusion Models Vulnerable to Membership Inference Attacks?

    Jinhao Duan, Fei Kong, Shiqi Wang +2

    cs.CVcs.AIcs.CRarXiv:2302.01316v22023
  3. Efficient Multi-view Clustering via Unified and Discrete Bipartite Graph Learning

    Si-Guo Fang, Dong Huang, Xiao-Sha Cai +3

    cs.LGcs.AIarXiv:2209.04187v22022
  4. Set You Straight: Auto-Steering Denoising Trajectories to Sidestep Unwanted Concepts

    Leyang Li, Shilin Lu, Yan Ren +1

    cs.CVcs.AIcs.CRarXiv:2504.12782v12025
  5. Machine Learning-Aided Operations and Communications of Unmanned Aerial Vehicles: A Contemporary Survey

    Harrison Kurunathan, Hailong Huang, Kai Li +2

    cs.ROcs.CVcs.LGarXiv:2211.04324v12022
  6. Physics Informed Neural Networks for Control Oriented Thermal Modeling of Buildings

    Gargya Gokhale, Bert Claessens, Chris Develder

    eess.SPcs.LGeess.SYarXiv:2111.12066v22021
  7. DALL-E-Bot: Introducing Web-Scale Diffusion Models to Robotics

    Ivan Kapelyukh, Vitalis Vosylius, Edward Johns

    cs.ROcs.CVcs.LGarXiv:2210.02438v32022
  8. Knowledge Transfer with Jacobian Matching

    Suraj Srinivas, Francois Fleuret

    cs.LGcs.CVarXiv:1803.00443v12018
  9. EDGE: Engine for Deterministic Graph Evaluation through Conversation Simulation from Graph Structured DSL Configuration

    Ram Kulathumani, Regunathan Radhakrishnan, Anupam Tripathi +5

    cs.AIcs.LGarXiv:2608.29971v12026
  10. Transformer-Based Flow Shop Scheduling Using MILP-Generated Training Data

    Roderich Wallrath

    math.OCcs.LGarXiv:2608.29690v12026
  11. All Roads Lead to Likelihood: The Value of Reinforcement Learning in Fine-Tuning

    Gokul Swamy, Sanjiban Choudhury, Wen Sun +2

    cs.LGarXiv:2503.01067v22025
  12. Towards a Systems Foundation for Agentic Skills: Architecture, Lifecycle, and Security

    Sanket Badhe, Deep Shah, Priyanka Tiwari +1

    cs.AIcs.LGcs.MAarXiv:2608.29596v12026
  13. Adaptive Doubly Robust Off-Policy Evaluation for Ranking Policies under Diverse User Behavior

    Kosuke Iguchi, Ren Kishimoto

    cs.LGcs.IRarXiv:2608.29600v12026
  14. Communication-Efficient and Distributed Learning Over Wireless Networks: Principles and Applications

    Jihong Park, Sumudu Samarakoon, Anis Elgabli +4

    cs.LGcs.ITcs.NIarXiv:2008.02608v12020
  15. End-to-End Deep Reinforcement Learning for Lane Keeping Assist

    Ahmad El Sallab, Mohammed Abdou, Etienne Perot +1

    stat.MLcs.LGcs.ROarXiv:1612.04340v12016
  16. Retrieval of aboveground crop nitrogen content with a hybrid machine learning method

    Katja Berger, Jochem Verrelst, Jean-Baptiste Féret +4

    q-bio.QMcs.LGarXiv:2012.05043v12020
  17. LiveMathematicianBench: A Live Benchmark for Mathematician-Level Reasoning with Proof Sketches

    Linyang He, Qiyao Yu, Hanze Dong +5

    cs.CLcs.AIcs.LGarXiv:2604.01754v12026
  18. Federated Learning for Healthcare Domain - Pipeline, Applications and Challenges

    Madhura Joshi, Ankit Pal, Malaikannan Sankarasubbu

    cs.LGcs.AIcs.CRarXiv:2211.07893v22022
  19. Controlled LLM Training on Spectral Sphere

    Tian Xie, Haoming Luo, Haoyu Tang +9

    cs.LGcs.AIarXiv:2601.08393v32026
  20. EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation

    Siyuan Huang, Liliang Chen, Pengfei Zhou +8

    cs.ROcs.CVcs.LGarXiv:2501.01895v32025
  21. Internet-augmented language models through few-shot prompting for open-domain question answering

    Angeliki Lazaridou, Elena Gribovskaya, Wojciech Stokowiec +1

    cs.CLcs.LGarXiv:2203.05115v22022
  22. A Survey on LoRA of Large Language Models

    Yuren Mao, Yuhang Ge, Yijiang Fan +4

    cs.LGcs.AIcs.CLarXiv:2407.11046v42024
  23. Multi-Agent Systems Execute Arbitrary Malicious Code

    Harold Triedman, Rishi Jha, Vitaly Shmatikov

    cs.CRcs.LGarXiv:2503.12188v22025
  24. Learning Combinatorial Optimization on Graphs: A Survey with Applications to Networking

    Natalia Vesselinova, Rebecca Steinert, Daniel F. Perez-Ramirez +1

    cs.LGcs.AIstat.MLarXiv:2005.11081v22020
  25. A Little Depth Goes a Long Way: The Expressive Power of Log-Depth Transformers

    William Merrill, Ashish Sabharwal

    cs.LGcs.CCarXiv:2503.03961v32025
  26. Distributed Deep Learning Using Synchronous Stochastic Gradient Descent

    Dipankar Das, Sasikanth Avancha, Dheevatsa Mudigere +5

    cs.DCcs.LGarXiv:1602.06709v12016
  27. A Hybrid Method for Traffic Flow Forecasting Using Multimodal Deep Learning

    Shengdong Du, Tianrui Li, Xun Gong +1

    cs.LGeess.SYarXiv:1803.02099v42018
  28. RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs

    Shreyas Chaudhari, Pranjal Aggarwal, Vishvak Murahari +5

    cs.LGcs.AIcs.CLarXiv:2404.08555v22024
  29. Deep Learning and Its Applications to Machine Health Monitoring: A Survey

    Rui Zhao, Ruqiang Yan, Zhenghua Chen +3

    cs.LGstat.MLarXiv:1612.07640v12016
  30. Recent Advances in Zero-shot Recognition

    Yanwei Fu, Tao Xiang, Yu-Gang Jiang +3

    cs.CVcs.AIcs.LGarXiv:1710.04837v12017
  31. A Comparative Survey of Recent Natural Language Interfaces for Databases

    Katrin Affolter, Kurt Stockinger, Abraham Bernstein

    cs.DBcs.CLcs.LGarXiv:1906.08990v12019
  32. Optimization for deep learning: theory and algorithms

    Ruoyu Sun

    cs.LGmath.OCstat.MLarXiv:1912.08957v12019
  33. Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research

    Bernard Koch, Emily Denton, Alex Hanna +1

    cs.LGcs.CLcs.CVarXiv:2112.01716v12021
  34. SkillGen: Verified Inference-Time Agent Skill Synthesis

    Yuchen Ma, Yue Huang, Han Bao +5

    cs.LGcs.AIcs.MAarXiv:2605.10999v12026
  35. Deception Abilities Emerged in Large Language Models

    Thilo Hagendorff

    cs.CLcs.AIcs.LGarXiv:2307.16513v22023
  36. DataSciBench: An LLM Agent Benchmark for Data Science

    Dan Zhang, Sining Zhoubian, Min Cai +7

    cs.CLcs.AIcs.LGarXiv:2502.13897v12025
  37. Deep Learning on FPGAs: Past, Present, and Future

    Griffin Lacey, Graham W. Taylor, Shawki Areibi

    cs.DCcs.LGstat.MLarXiv:1602.04283v12016
  38. Tiny Machine Learning: Progress and Futures

    Ji Lin, Ligeng Zhu, Wei-Ming Chen +2

    cs.LGcs.AIcs.CVarXiv:2403.19076v22024
  39. Rethinking Irregular Time Series Forecasting: A Simple yet Effective Baseline

    Xvyuan Liu, Xiangfei Qiu, Xingjian Wu +4

    cs.LGarXiv:2505.11250v42025
  40. Enhancing Multimodal Emotion Recognition via Multi-Feature Encoding and Attention-Based Fusion

    Xu Lin, Ke Wang, Hui Kang +1

    cs.CVcs.AIcs.LGarXiv:2609.04690v12026
  41. Label-Efficient Self-Supervised Federated Learning for Tackling Data Heterogeneity in Medical Imaging

    Rui Yan, Liangqiong Qu, Qingyue Wei +5

    cs.CVcs.LGarXiv:2205.08576v22022
  42. Fast and Effective On-policy Distillation from Reasoning Prefixes

    Dongxu Zhang, Zhichao Yang, Sepehr Janghorbani +6

    cs.LGcs.AIarXiv:2602.15260v12026
  43. PinnerSage: Multi-Modal User Embedding Framework for Recommendations at Pinterest

    Aditya Pal, Chantat Eksombatchai, Yitong Zhou +3

    cs.LGcs.IRcs.SIarXiv:2007.03634v12020
  44. HSVA: Hierarchical Semantic-Visual Adaptation for Zero-Shot Learning

    Shiming Chen, Guo-Sen Xie, Yang Liu +5

    cs.CVcs.LGarXiv:2109.15163v22021
  45. When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning

    Xiaogeng Liu, Xinyan Wang, Yingzi Ma +2

    cs.LGcs.AIarXiv:2605.21606v12026
  46. Quantum-Assisted Memory-Efficient Training for Parameter-Intensive Wi-Fi-Based Human Activity Recognition

    To Truong An, Jie Zhang, Guolin Yin +4

    cs.LGarXiv:2609.04271v12026
  47. Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems

    Xupeng Miao, Gabriele Oliaro, Zhihao Zhang +4

    cs.LGcs.AIcs.DCarXiv:2312.15234v22023
  48. Sim-and-Real Co-Training: A Simple Recipe for Vision-Based Robotic Manipulation

    Abhiram Maddukuri, Zhenyu Jiang, Lawrence Yunliang Chen +12

    cs.ROcs.AIcs.LGarXiv:2503.24361v22025
  49. Continual Learning via Neural Pruning

    Siavash Golkar, Michael Kagan, Kyunghyun Cho

    cs.LGcs.NEq-bio.NCarXiv:1903.04476v12019
  50. Policy Learning with Observational Data

    Susan Athey, Stefan Wager

    math.STcs.LGecon.EMarXiv:1702.02896v62017
  51. Fast Transformers with Clustered Attention

    Apoorv Vyas, Angelos Katharopoulos, François Fleuret

    cs.LGstat.MLarXiv:2007.04825v22020
  52. Gym-Anything: Turn any Software into an Agent Environment

    Pranjal Aggarwal, Graham Neubig, Sean Welleck

    cs.LGcs.AIarXiv:2604.06126v12026
  53. One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention

    Arvind Mahankali, Tatsunori B. Hashimoto, Tengyu Ma

    cs.LGarXiv:2307.03576v12023
  54. Dynamic Heterogeneous Graph Representation Learning: A Survey

    Huan Liu, Pengfei Jiao, Jie Yin +2

    cs.LGcs.AIcs.SIarXiv:2609.04779v12026
  55. Learned-Norm Pooling for Deep Feedforward and Recurrent Neural Networks

    Caglar Gulcehre, Kyunghyun Cho, Razvan Pascanu +1

    cs.NEcs.LGstat.MLarXiv:1311.1780v72013
  56. Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry

    Sai Sumedh R. Hindupur, Ekdeep Singh Lubana, Thomas Fel +1

    cs.LGcs.AIarXiv:2503.01822v22025
  57. When Does LeJEPA Learn a World Model?

    David Klindt, Yann LeCun, Randall Balestriero

    stat.MLcs.LGarXiv:2605.26379v12026
  58. Embedding Text in Hyperbolic Spaces

    Bhuwan Dhingra, Christopher J. Shallue, Mohammad Norouzi +2

    cs.CLcs.LGarXiv:1806.04313v12018
  59. E-LSTM-D: A Deep Learning Framework for Dynamic Network Link Prediction

    Jinyin Chen, Jian Zhang, Xuanheng Xu +4

    cs.SIcs.LGarXiv:1902.08329v12019
  60. On the Position Bias of On-Policy Distillation

    Yan Xie, Sijie Zhu, Tiansheng Wen +2

    cs.LGcs.AIarXiv:2606.22600v32026