Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

7,621 to 7,680 of 20,153

  1. Tighter Theory for Local SGD on Identical and Heterogeneous Data

    Ahmed Khaled, Konstantin Mishchenko, Peter Richtárik

    cs.LGcs.DCmath.NAarXiv:1909.04746v42019
  2. On the Variance of the Adaptive Learning Rate and Beyond

    Liyuan Liu, Haoming Jiang, Pengcheng He +4

    cs.LGcs.CLstat.MLarXiv:1908.03265v42019
  3. GraphSAINT: Graph Sampling Based Inductive Learning Method

    Hanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava +2

    cs.LGstat.MLarXiv:1907.04931v42019
  4. Large Scale Adversarial Representation Learning

    Jeff Donahue, Karen Simonyan

    cs.CVcs.LGstat.MLarXiv:1907.02544v22019
  5. On the Convergence of FedAvg on Non-IID Data

    Xiang Li, Kaixuan Huang, Wenhao Yang +2

    stat.MLcs.LGmath.OCarXiv:1907.02189v42019
  6. Does Learning Require Memorization? A Short Tale about a Long Tail

    Vitaly Feldman

    cs.LGstat.MLarXiv:1906.05271v42019
  7. Graph Neural Tangent Kernel: Fusing Graph Neural Networks with Graph Kernels

    Simon S. Du, Kangcheng Hou, Barnabás Póczos +3

    cs.LGcs.AIcs.CVarXiv:1905.13192v22019
  8. Text Classification Algorithms: A Survey

    Kamran Kowsari, Kiana Jafari Meimandi, Mojtaba Heidarysafa +3

    cs.LGcs.AIcs.CLarXiv:1904.08067v52019
  9. Provably Powerful Graph Networks

    Haggai Maron, Heli Ben-Hamu, Hadar Serviansky +1

    cs.LGstat.MLarXiv:1905.11136v42019
  10. Deep Reinforcement Learning for Sepsis Treatment

    Aniruddh Raghu, Matthieu Komorowski, Imran Ahmed +3

    cs.AIcs.LGarXiv:1711.09602v12017
  11. Mercury: Ultra-Fast Language Models Based on Diffusion

    Inception Labs, Samar Khanna, Siddhant Kharbanda +10

    cs.CLcs.AIcs.LGarXiv:2506.17298v12025
  12. On Exact Computation with an Infinitely Wide Neural Net

    Sanjeev Arora, Simon S. Du, Wei Hu +3

    cs.LGcs.CVcs.NEarXiv:1904.11955v22019
  13. Embarrassingly Shallow Autoencoders for Sparse Data

    Harald Steck

    cs.IRcs.LGstat.MLarXiv:1905.03375v12019
  14. CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMs

    Maryam Alshehyari, Dushyant Singh Chauhan, Samuele Poppi +3

    cs.LGarXiv:2609.01161v12026
  15. On the Convergence of Adam and Beyond

    Sashank J. Reddi, Satyen Kale, Sanjiv Kumar

    cs.LGmath.OCstat.MLarXiv:1904.09237v12019
  16. A Survey on Traffic Signal Control Methods

    Hua Wei, Guanjie Zheng, Vikash Gayah +1

    cs.LGcs.AIstat.MLarXiv:1904.08117v32019
  17. Surprises in High-Dimensional Ridgeless Least Squares Interpolation

    Trevor Hastie, Andrea Montanari, Saharon Rosset +1

    math.STcs.LGstat.MLarXiv:1903.08560v52019
  18. Three scenarios for continual learning

    Gido M. van de Ven, Andreas S. Tolias

    cs.LGcs.AIcs.CVarXiv:1904.07734v12019
  19. ICLabel: An automated electroencephalographic independent component classifier, dataset, and website

    Luca Pion-Tonachini, Ken Kreutz-Delgado, Scott Makeig

    eess.SPcs.LGstat.MLarXiv:1901.07915v22019
  20. Large Batch Optimization for Deep Learning: Training BERT in 76 minutes

    Yang You, Jing Li, Sashank Reddi +7

    cs.LGcs.AIcs.CLarXiv:1904.00962v52019
  21. Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent

    Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz +4

    stat.MLcs.LGarXiv:1902.06720v42019
  22. Graph Neural Networks for Social Recommendation

    Wenqi Fan, Yao Ma, Qing Li +4

    cs.IRcs.LGcs.SIarXiv:1902.07243v22019
  23. Adaptive Gradient Methods with Dynamic Bound of Learning Rate

    Liangchen Luo, Yuanhao Xiong, Yan Liu +1

    cs.LGstat.MLarXiv:1902.09843v12019
  24. Theoretically Principled Trade-off between Robustness and Accuracy

    Hongyang Zhang, Yaodong Yu, Jiantao Jiao +3

    cs.LGstat.MLarXiv:1901.08573v32019
  25. Physics-Constrained Deep Learning for High-dimensional Surrogate Modeling and Uncertainty Quantification without Labeled Data

    Yinhao Zhu, Nicholas Zabaras, Phaedon-Stelios Koutsourelakis +1

    physics.comp-phcs.CVcs.LGarXiv:1901.06314v12019
  26. Hybrid Recommender Systems: A Systematic Literature Review

    Erion Çano, Maurizio Morisio

    cs.IRcs.CYcs.LGarXiv:1901.03888v12019
  27. A Survey of Unsupervised Deep Domain Adaptation

    Garrett Wilson, Diane J. Cook

    cs.LGstat.MLarXiv:1812.02849v32018
  28. Soft Actor-Critic Algorithms and Applications

    Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen +8

    cs.LGcs.AIcs.ROarXiv:1812.05905v22018
  29. Optimistic mirror descent in saddle-point problems: Going the extra (gradient) mile

    Panayotis Mertikopoulos, Bruno Lecouat, Houssam Zenati +3

    cs.LGcs.GTmath.OCarXiv:1807.02629v22018
  30. Deep Neural Networks for Estimation and Inference

    Max H. Farrell, Tengyuan Liang, Sanjog Misra

    econ.EMcs.LGmath.STarXiv:1809.09953v32018
  31. GPyTorch: Blackbox Matrix-Matrix Gaussian Process Inference with GPU Acceleration

    Jacob R. Gardner, Geoff Pleiss, David Bindel +2

    cs.LGstat.MLarXiv:1809.11165v62018
  32. Gradient Descent Provably Optimizes Over-parameterized Neural Networks

    Simon S. Du, Xiyu Zhai, Barnabas Poczos +1

    cs.LGmath.OCstat.MLarXiv:1810.02054v22018
  33. Learning deep representations by mutual information estimation and maximization

    R Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon +4

    stat.MLcs.LGarXiv:1808.06670v52018
  34. Parallel Restarted SGD with Faster Convergence and Less Communication: Demystifying Why Model Averaging Works for Deep Learning

    Hao Yu, Sen Yang, Shenghuo Zhu

    math.OCcs.DCcs.LGarXiv:1807.06629v32018
  35. The relativistic discriminator: a key element missing from standard GAN

    Alexia Jolicoeur-Martineau

    cs.LGcs.AIcs.CRarXiv:1807.00734v32018
  36. Understanding Batch Normalization

    Johan Bjorck, Carla Gomes, Bart Selman +1

    cs.LGcs.AIstat.MLarXiv:1806.02375v42018
  37. DARTS: Differentiable Architecture Search

    Hanxiao Liu, Karen Simonyan, Yiming Yang

    cs.LGcs.CLcs.CVarXiv:1806.09055v22018
  38. A General Framework for Inference-time Scaling and Steering of Diffusion Models

    Raghav Singhal, Zachary Horvitz, Ryan Teehan +4

    cs.LGcs.CLcs.CVarXiv:2501.06848v52025
  39. Robustness May Be at Odds with Accuracy

    Dimitris Tsipras, Shibani Santurkar, Logan Engstrom +2

    stat.MLcs.CVcs.LGarXiv:1805.12152v52018
  40. TADAM: Task dependent adaptive metric for improved few-shot learning

    Boris N. Oreshkin, Pau Rodriguez, Alexandre Lacoste

    cs.LGcs.AIcs.CVarXiv:1805.10123v42018
  41. Hyperbolic Neural Networks

    Octavian-Eugen Ganea, Gary Bécigneul, Thomas Hofmann

    cs.LGstat.MLarXiv:1805.09112v22018
  42. Data-Efficient Hierarchical Reinforcement Learning

    Ofir Nachum, Shixiang Gu, Honglak Lee +1

    cs.LGcs.AIstat.MLarXiv:1805.08296v42018
  43. Byzantine-Robust Distributed Learning: Towards Optimal Statistical Rates

    Dong Yin, Yudong Chen, Kannan Ramchandran +1

    cs.LGcs.CRcs.DCarXiv:1803.01498v22018
  44. Scalable Private Learning with PATE

    Nicolas Papernot, Shuang Song, Ilya Mironov +3

    stat.MLcs.CRcs.LGarXiv:1802.08908v12018
  45. Mean Field Multi-Agent Reinforcement Learning

    Yaodong Yang, Rui Luo, Minne Li +3

    cs.MAcs.AIcs.LGarXiv:1802.05438v52018
  46. Bayesian Deep Convolutional Encoder-Decoder Networks for Surrogate Modeling and Uncertainty Quantification

    Yinhao Zhu, Nicholas Zabaras

    physics.comp-phcs.CVcs.LGarXiv:1801.06879v12018
  47. Reasoning with Latent Thoughts: On the Power of Looped Transformers

    Nikunj Saunshi, Nishanth Dikkala, Zhiyuan Li +2

    cs.CLcs.AIcs.LGarXiv:2502.17416v12025
  48. Demystifying MMD GANs

    Mikołaj Bińkowski, Danica J. Sutherland, Michael Arbel +1

    stat.MLcs.LGarXiv:1801.01401v52018
  49. Size-Independent Sample Complexity of Neural Networks

    Noah Golowich, Alexander Rakhlin, Ohad Shamir

    cs.LGcs.NEstat.MLarXiv:1712.06541v52017
  50. REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers

    Xingjian Leng, Jaskirat Singh, Yunzhong Hou +3

    cs.CVcs.LGarXiv:2504.10483v32025
  51. Towards Accurate Binary Convolutional Neural Network

    Xiaofan Lin, Cong Zhao, Wei Pan

    cs.LGstat.MLarXiv:1711.11294v12017
  52. Deep Learning for Physical Processes: Incorporating Prior Scientific Knowledge

    Emmanuel de Bezenac, Arthur Pajot, Patrick Gallinari

    cs.AIcs.LGstat.MLarXiv:1711.07970v22017
  53. Weakly-Supervised Neural Text Classification

    Yu Meng, Jiaming Shen, Chao Zhang +1

    cs.IRcs.CLcs.LGarXiv:1809.01478v22018
  54. Exploring Speech Enhancement with Generative Adversarial Networks for Robust Speech Recognition

    Chris Donahue, Bo Li, Rohit Prabhavalkar

    cs.SDcs.LGcs.NEarXiv:1711.05747v22017
  55. The Implicit Bias of Gradient Descent on Separable Data

    Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson +2

    stat.MLcs.LGarXiv:1710.10345v72017
  56. Ensembles of Multiple Models and Architectures for Robust Brain Tumour Segmentation

    Konstantinos Kamnitsas, Wenjia Bai, Enzo Ferrante +8

    cs.CVcs.AIcs.LGarXiv:1711.01468v12017
  57. Self-Normalizing Neural Networks

    Günter Klambauer, Thomas Unterthiner, Andreas Mayr +1

    cs.LGstat.MLarXiv:1706.02515v52017
  58. A systematic study of the class imbalance problem in convolutional neural networks

    Mateusz Buda, Atsuto Maki, Maciej A. Mazurowski

    cs.CVcs.AIcs.LGarXiv:1710.05381v22017
  59. A Tutorial on Thompson Sampling

    Daniel Russo, Benjamin Van Roy, Abbas Kazerouni +2

    cs.LGarXiv:1707.02038v32017
  60. Deep Potential Molecular Dynamics: a scalable model with the accuracy of quantum mechanics

    Linfeng Zhang, Jiequn Han, Han Wang +2

    physics.comp-phcs.LGphysics.chem-pharXiv:1707.09571v22017