Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
7,621 to 7,680 of 20,153
Tighter Theory for Local SGD on Identical and Heterogeneous Data
Ahmed Khaled, Konstantin Mishchenko, Peter Richtárik
cs.LGcs.DCmath.NAarXiv:1909.04746v42019On the Variance of the Adaptive Learning Rate and Beyond
Liyuan Liu, Haoming Jiang, Pengcheng He +4
cs.LGcs.CLstat.MLarXiv:1908.03265v42019GraphSAINT: Graph Sampling Based Inductive Learning Method
Hanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava +2
cs.LGstat.MLarXiv:1907.04931v42019Large Scale Adversarial Representation Learning
Jeff Donahue, Karen Simonyan
cs.CVcs.LGstat.MLarXiv:1907.02544v22019On the Convergence of FedAvg on Non-IID Data
Xiang Li, Kaixuan Huang, Wenhao Yang +2
stat.MLcs.LGmath.OCarXiv:1907.02189v42019Does Learning Require Memorization? A Short Tale about a Long Tail
Vitaly Feldman
cs.LGstat.MLarXiv:1906.05271v42019Graph Neural Tangent Kernel: Fusing Graph Neural Networks with Graph Kernels
Simon S. Du, Kangcheng Hou, Barnabás Póczos +3
cs.LGcs.AIcs.CVarXiv:1905.13192v22019Text Classification Algorithms: A Survey
Kamran Kowsari, Kiana Jafari Meimandi, Mojtaba Heidarysafa +3
cs.LGcs.AIcs.CLarXiv:1904.08067v52019Provably Powerful Graph Networks
Haggai Maron, Heli Ben-Hamu, Hadar Serviansky +1
cs.LGstat.MLarXiv:1905.11136v42019Deep Reinforcement Learning for Sepsis Treatment
Aniruddh Raghu, Matthieu Komorowski, Imran Ahmed +3
cs.AIcs.LGarXiv:1711.09602v12017Mercury: Ultra-Fast Language Models Based on Diffusion
Inception Labs, Samar Khanna, Siddhant Kharbanda +10
cs.CLcs.AIcs.LGarXiv:2506.17298v12025On Exact Computation with an Infinitely Wide Neural Net
Sanjeev Arora, Simon S. Du, Wei Hu +3
cs.LGcs.CVcs.NEarXiv:1904.11955v22019Embarrassingly Shallow Autoencoders for Sparse Data
Harald Steck
cs.IRcs.LGstat.MLarXiv:1905.03375v12019CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMs
Maryam Alshehyari, Dushyant Singh Chauhan, Samuele Poppi +3
cs.LGarXiv:2609.01161v12026On the Convergence of Adam and Beyond
Sashank J. Reddi, Satyen Kale, Sanjiv Kumar
cs.LGmath.OCstat.MLarXiv:1904.09237v12019A Survey on Traffic Signal Control Methods
Hua Wei, Guanjie Zheng, Vikash Gayah +1
cs.LGcs.AIstat.MLarXiv:1904.08117v32019Surprises in High-Dimensional Ridgeless Least Squares Interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset +1
math.STcs.LGstat.MLarXiv:1903.08560v52019Three scenarios for continual learning
Gido M. van de Ven, Andreas S. Tolias
cs.LGcs.AIcs.CVarXiv:1904.07734v12019ICLabel: An automated electroencephalographic independent component classifier, dataset, and website
Luca Pion-Tonachini, Ken Kreutz-Delgado, Scott Makeig
eess.SPcs.LGstat.MLarXiv:1901.07915v22019Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
Yang You, Jing Li, Sashank Reddi +7
cs.LGcs.AIcs.CLarXiv:1904.00962v52019Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent
Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz +4
stat.MLcs.LGarXiv:1902.06720v42019Graph Neural Networks for Social Recommendation
Wenqi Fan, Yao Ma, Qing Li +4
cs.IRcs.LGcs.SIarXiv:1902.07243v22019Adaptive Gradient Methods with Dynamic Bound of Learning Rate
Liangchen Luo, Yuanhao Xiong, Yan Liu +1
cs.LGstat.MLarXiv:1902.09843v12019Theoretically Principled Trade-off between Robustness and Accuracy
Hongyang Zhang, Yaodong Yu, Jiantao Jiao +3
cs.LGstat.MLarXiv:1901.08573v32019Physics-Constrained Deep Learning for High-dimensional Surrogate Modeling and Uncertainty Quantification without Labeled Data
Yinhao Zhu, Nicholas Zabaras, Phaedon-Stelios Koutsourelakis +1
physics.comp-phcs.CVcs.LGarXiv:1901.06314v12019Hybrid Recommender Systems: A Systematic Literature Review
Erion Çano, Maurizio Morisio
cs.IRcs.CYcs.LGarXiv:1901.03888v12019A Survey of Unsupervised Deep Domain Adaptation
Garrett Wilson, Diane J. Cook
cs.LGstat.MLarXiv:1812.02849v32018Soft Actor-Critic Algorithms and Applications
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen +8
cs.LGcs.AIcs.ROarXiv:1812.05905v22018Optimistic mirror descent in saddle-point problems: Going the extra (gradient) mile
Panayotis Mertikopoulos, Bruno Lecouat, Houssam Zenati +3
cs.LGcs.GTmath.OCarXiv:1807.02629v22018Deep Neural Networks for Estimation and Inference
Max H. Farrell, Tengyuan Liang, Sanjog Misra
econ.EMcs.LGmath.STarXiv:1809.09953v32018GPyTorch: Blackbox Matrix-Matrix Gaussian Process Inference with GPU Acceleration
Jacob R. Gardner, Geoff Pleiss, David Bindel +2
cs.LGstat.MLarXiv:1809.11165v62018Gradient Descent Provably Optimizes Over-parameterized Neural Networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos +1
cs.LGmath.OCstat.MLarXiv:1810.02054v22018Learning deep representations by mutual information estimation and maximization
R Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon +4
stat.MLcs.LGarXiv:1808.06670v52018Parallel Restarted SGD with Faster Convergence and Less Communication: Demystifying Why Model Averaging Works for Deep Learning
Hao Yu, Sen Yang, Shenghuo Zhu
math.OCcs.DCcs.LGarXiv:1807.06629v32018The relativistic discriminator: a key element missing from standard GAN
Alexia Jolicoeur-Martineau
cs.LGcs.AIcs.CRarXiv:1807.00734v32018Understanding Batch Normalization
Johan Bjorck, Carla Gomes, Bart Selman +1
cs.LGcs.AIstat.MLarXiv:1806.02375v42018DARTS: Differentiable Architecture Search
Hanxiao Liu, Karen Simonyan, Yiming Yang
cs.LGcs.CLcs.CVarXiv:1806.09055v22018A General Framework for Inference-time Scaling and Steering of Diffusion Models
Raghav Singhal, Zachary Horvitz, Ryan Teehan +4
cs.LGcs.CLcs.CVarXiv:2501.06848v52025Robustness May Be at Odds with Accuracy
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom +2
stat.MLcs.CVcs.LGarXiv:1805.12152v52018TADAM: Task dependent adaptive metric for improved few-shot learning
Boris N. Oreshkin, Pau Rodriguez, Alexandre Lacoste
cs.LGcs.AIcs.CVarXiv:1805.10123v42018Hyperbolic Neural Networks
Octavian-Eugen Ganea, Gary Bécigneul, Thomas Hofmann
cs.LGstat.MLarXiv:1805.09112v22018Data-Efficient Hierarchical Reinforcement Learning
Ofir Nachum, Shixiang Gu, Honglak Lee +1
cs.LGcs.AIstat.MLarXiv:1805.08296v42018Byzantine-Robust Distributed Learning: Towards Optimal Statistical Rates
Dong Yin, Yudong Chen, Kannan Ramchandran +1
cs.LGcs.CRcs.DCarXiv:1803.01498v22018Scalable Private Learning with PATE
Nicolas Papernot, Shuang Song, Ilya Mironov +3
stat.MLcs.CRcs.LGarXiv:1802.08908v12018Mean Field Multi-Agent Reinforcement Learning
Yaodong Yang, Rui Luo, Minne Li +3
cs.MAcs.AIcs.LGarXiv:1802.05438v52018Bayesian Deep Convolutional Encoder-Decoder Networks for Surrogate Modeling and Uncertainty Quantification
Yinhao Zhu, Nicholas Zabaras
physics.comp-phcs.CVcs.LGarXiv:1801.06879v12018Reasoning with Latent Thoughts: On the Power of Looped Transformers
Nikunj Saunshi, Nishanth Dikkala, Zhiyuan Li +2
cs.CLcs.AIcs.LGarXiv:2502.17416v12025Demystifying MMD GANs
Mikołaj Bińkowski, Danica J. Sutherland, Michael Arbel +1
stat.MLcs.LGarXiv:1801.01401v52018Size-Independent Sample Complexity of Neural Networks
Noah Golowich, Alexander Rakhlin, Ohad Shamir
cs.LGcs.NEstat.MLarXiv:1712.06541v52017REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
Xingjian Leng, Jaskirat Singh, Yunzhong Hou +3
cs.CVcs.LGarXiv:2504.10483v32025Towards Accurate Binary Convolutional Neural Network
Xiaofan Lin, Cong Zhao, Wei Pan
cs.LGstat.MLarXiv:1711.11294v12017Deep Learning for Physical Processes: Incorporating Prior Scientific Knowledge
Emmanuel de Bezenac, Arthur Pajot, Patrick Gallinari
cs.AIcs.LGstat.MLarXiv:1711.07970v22017Weakly-Supervised Neural Text Classification
Yu Meng, Jiaming Shen, Chao Zhang +1
cs.IRcs.CLcs.LGarXiv:1809.01478v22018Exploring Speech Enhancement with Generative Adversarial Networks for Robust Speech Recognition
Chris Donahue, Bo Li, Rohit Prabhavalkar
cs.SDcs.LGcs.NEarXiv:1711.05747v22017The Implicit Bias of Gradient Descent on Separable Data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson +2
stat.MLcs.LGarXiv:1710.10345v72017Ensembles of Multiple Models and Architectures for Robust Brain Tumour Segmentation
Konstantinos Kamnitsas, Wenjia Bai, Enzo Ferrante +8
cs.CVcs.AIcs.LGarXiv:1711.01468v12017Self-Normalizing Neural Networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr +1
cs.LGstat.MLarXiv:1706.02515v52017A systematic study of the class imbalance problem in convolutional neural networks
Mateusz Buda, Atsuto Maki, Maciej A. Mazurowski
cs.CVcs.AIcs.LGarXiv:1710.05381v22017A Tutorial on Thompson Sampling
Daniel Russo, Benjamin Van Roy, Abbas Kazerouni +2
cs.LGarXiv:1707.02038v32017Deep Potential Molecular Dynamics: a scalable model with the accuracy of quantum mechanics
Linfeng Zhang, Jiequn Han, Han Wang +2
physics.comp-phcs.LGphysics.chem-pharXiv:1707.09571v22017