Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

7,741 to 7,800 of 20,193

  1. An Actor-Critic Algorithm for Sequence Prediction

    Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu +5

    cs.LGarXiv:1607.07086v32016
  2. Generative Adversarial Imitation Learning

    Jonathan Ho, Stefano Ermon

    cs.LGcs.AIarXiv:1606.03476v12016
  3. Unifying Count-Based Exploration and Intrinsic Motivation

    Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski +3

    cs.AIcs.LGstat.MLarXiv:1606.01868v22016
  4. Adversarial Feature Learning

    Jeff Donahue, Philipp Krähenbühl, Trevor Darrell

    cs.LGcs.AIcs.CVarXiv:1605.09782v72016
  5. Fairness in Learning: Classic and Contextual Bandits

    Matthew Joseph, Michael Kearns, Jamie Morgenstern +1

    cs.LGstat.MLarXiv:1605.07139v22016
  6. Deep Learning without Poor Local Minima

    Kenji Kawaguchi

    stat.MLcs.LGmath.OCarXiv:1605.07110v32016
  7. Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization

    Lisha Li, Kevin Jamieson, Giulia DeSalvo +2

    cs.LGstat.MLarXiv:1603.06560v42016
  8. Bayesian Consensus Clustering

    Eric F. Lock, David B. Dunson

    stat.MLcs.LGarXiv:1302.7280v12013
  9. Rényi Divergence Variational Inference

    Yingzhen Li, Richard E. Turner

    stat.MLcs.LGarXiv:1602.02311v32016
  10. Feature Selection: A Data Perspective

    Jundong Li, Kewei Cheng, Suhang Wang +4

    cs.LGarXiv:1601.07996v52016
  11. From Softmax to Sparsemax: A Sparse Model of Attention and Multi-Label Classification

    André F. T. Martins, Ramón Fernandez Astudillo

    cs.CLcs.LGstat.MLarXiv:1602.02068v22016
  12. DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

    Yuxiang Zheng, Dayuan Fu, Xiangkun Hu +4

    cs.AIcs.CLcs.LGarXiv:2504.03160v42025
  13. Gated Graph Sequence Neural Networks

    Yujia Li, Daniel Tarlow, Marc Brockschmidt +1

    cs.LGcs.AIcs.NEarXiv:1511.05493v42015
  14. The Ubuntu Dialogue Corpus: A Large Dataset for Research in Unstructured Multi-Turn Dialogue Systems

    Ryan Lowe, Nissan Pow, Iulian Serban +1

    cs.CLcs.AIcs.LGarXiv:1506.08909v32015
  15. Hinge-Loss Markov Random Fields and Probabilistic Soft Logic

    Stephen H. Bach, Matthias Broecheler, Bert Huang +1

    cs.LGcs.AIstat.MLarXiv:1505.04406v32015
  16. Escaping From Saddle Points --- Online Stochastic Gradient for Tensor Decomposition

    Rong Ge, Furong Huang, Chi Jin +1

    cs.LGmath.OCstat.MLarXiv:1503.02101v12015
  17. Automatic differentiation in machine learning: a survey

    Atilim Gunes Baydin, Barak A. Pearlmutter, Alexey Andreyevich Radul +1

    cs.SCcs.LGstat.MLarXiv:1502.05767v42015
  18. Non-stochastic Best Arm Identification and Hyperparameter Optimization

    Kevin Jamieson, Ameet Talwalkar

    cs.LGstat.MLarXiv:1502.07943v12015
  19. Deep Learning with Limited Numerical Precision

    Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan +1

    cs.LGcs.NEstat.MLarXiv:1502.02551v12015
  20. New insights and perspectives on the natural gradient method

    James Martens

    cs.LGstat.MLarXiv:1412.1193v112014
  21. The Loss Surfaces of Multilayer Networks

    Anna Choromanska, Mikael Henaff, Michael Mathieu +2

    cs.LGarXiv:1412.0233v32014
  22. Breaking the Curse of Dimensionality with Convex Neural Networks

    Francis Bach

    cs.LGmath.OCmath.STarXiv:1412.8690v22014
  23. Making Gradient Descent Optimal for Strongly Convex Stochastic Optimization

    Alexander Rakhlin, Ohad Shamir, Karthik Sridharan

    cs.LGmath.OCarXiv:1109.5647v72011
  24. Generalized Low Rank Models

    Madeleine Udell, Corinne Horn, Reza Zadeh +1

    stat.MLcs.LGmath.OCarXiv:1410.0342v42014
  25. On the Complexity of Best Arm Identification in Multi-Armed Bandit Models

    Emilie Kaufmann, Olivier Cappé, Aurélien Garivier

    stat.MLcs.LGarXiv:1407.4443v22014
  26. Minimizing Finite Sums with the Stochastic Average Gradient

    Mark Schmidt, Nicolas Le Roux, Francis Bach

    math.OCcs.LGstat.COarXiv:1309.2388v22013
  27. Stochastic Gradient Descent, Weighted Sampling, and the Randomized Kaczmarz algorithm

    Deanna Needell, Nathan Srebro, Rachel Ward

    math.NAcs.CVcs.LGarXiv:1310.5715v52013
  28. Phase Retrieval using Alternating Minimization

    Praneeth Netrapalli, Prateek Jain, Sujay Sanghavi

    stat.MLcs.ITcs.LGarXiv:1306.0160v22013
  29. Bandits with Knapsacks

    Ashwinkumar Badanidiyuru, Robert Kleinberg, Aleksandrs Slivkins

    cs.DScs.LGarXiv:1305.2545v82013
  30. Tensor decompositions for learning latent variable models

    Anima Anandkumar, Rong Ge, Daniel Hsu +2

    cs.LGmath.NAstat.MLarXiv:1210.7559v42012
  31. A Widely Applicable Bayesian Information Criterion

    Sumio Watanabe

    cs.LGstat.MLarXiv:1208.6338v12012
  32. Sparse Approximation via Penalty Decomposition Methods

    Zhaosong Lu, Yong Zhang

    cs.LGmath.OCstat.COarXiv:1205.2334v22012
  33. Learning high-dimensional directed acyclic graphs with latent and selection variables

    Diego Colombo, Marloes H. Maathuis, Markus Kalisch +1

    stat.MEcs.LGmath.STarXiv:1104.5617v32011
  34. Noisy matrix decomposition via convex relaxation: Optimal rates in high dimensions

    Alekh Agarwal, Sahand N. Negahban, Martin J. Wainwright

    stat.MLcs.ITcs.LGarXiv:1102.4807v32011
  35. Optimization with Sparsity-Inducing Penalties

    Francis Bach, Rodolphe Jenatton, Julien Mairal +1

    cs.LGmath.OCstat.MLarXiv:1108.0775v22011
  36. HOGWILD!: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent

    Feng Niu, Benjamin Recht, Christopher Re +1

    math.OCcs.LGarXiv:1106.5730v22011
  37. Adaptive Submodularity: Theory and Applications in Active Learning and Stochastic Optimization

    Daniel Golovin, Andreas Krause

    cs.LGcs.AIcs.DSarXiv:1003.3967v52010
  38. A Dynamic Near-Optimal Algorithm for Online Linear Programming

    Shipra Agrawal, Zizhuo Wang, Yinyu Ye

    cs.DScs.LGarXiv:0911.2974v32009
  39. Contextual Bandits with Similarity Information

    Aleksandrs Slivkins

    cs.DScs.LGarXiv:0907.3986v52009
  40. Matrix Completion from Noisy Entries

    Raghunandan H. Keshavan, Andrea Montanari, Sewoong Oh

    cs.LGstat.MLarXiv:0906.2027v22009
  41. HVI: A New Color Space for Low-light Image Enhancement

    Qingsen Yan, Yixu Feng, Cheng Zhang +6

    cs.CVcs.AIcs.LGarXiv:2502.20272v22025
  42. Masked Face Recognition for Secure Authentication

    Aqeel Anwar, Arijit Raychowdhury

    cs.CVcs.LGeess.IVarXiv:2008.11104v12020
  43. Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions

    Jaeyeon Kim, Kulin Shah, Vasilis Kontonis +2

    cs.LGarXiv:2502.06768v32025
  44. Automated Detection and Forecasting of COVID-19 using Deep Learning Techniques: A Review

    Afshin Shoeibi, Marjane Khodatars, Mahboobeh Jafari +12

    cs.LGeess.IVarXiv:2007.10785v72020
  45. Scalable Best-of-N Selection for Large Language Models via Self-Certainty

    Zhewei Kang, Xuandong Zhao, Dawn Song

    cs.CLcs.AIcs.LGarXiv:2502.18581v32025
  46. Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

    Lucy Xiaoyang Shi, Brian Ichter, Michael Equi +12

    cs.ROcs.AIcs.LGarXiv:2502.19417v22025
  47. Learning Debiased Representation via Disentangled Feature Augmentation

    Jungsoo Lee, Eungyeup Kim, Juyoung Lee +2

    cs.LGarXiv:2107.01372v22021
  48. Path Integral Sampler: a stochastic control approach for sampling

    Qinsheng Zhang, Yongxin Chen

    cs.LGarXiv:2111.15141v22021
  49. d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning

    Siyan Zhao, Devaansh Gupta, Qinqing Zheng +1

    cs.CLcs.LGarXiv:2504.12216v22025
  50. ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills

    Tairan He, Jiawei Gao, Wenli Xiao +15

    cs.ROcs.AIcs.LGarXiv:2502.01143v32025
  51. Goedel-Prover-V2: Scaling Formal Theorem Proving with Scaffolded Data Synthesis and Self-Correction

    Yong Lin, Shange Tang, Bohan Lyu +17

    cs.LGcs.AIarXiv:2508.03613v12025
  52. Quantum-Grassmann-Plucker Token Mixing for Deep Learning-Based Post-Disaster Damage Assessment

    Kooroush Farahkhah, Umut Lagap, Taha Rezaei +1

    cs.CVcs.LGarXiv:2608.30633v12026
  53. Neural ODE enhanced linear mixed effect models for estimating complex association patterns of time-varying covariates with the marker trajectory

    Zhe Aurore Li, Quentin Clairon, Cécilia Samieri +3

    stat.MLcs.LGarXiv:2608.29714v12026
  54. Muon with Finite Newton-Schulz: The Smoothing Benefit in Nonsmooth Nonconvex Optimization

    Mingyi Li, Taira Tsuchiya

    cs.LGmath.OCstat.MLarXiv:2608.26288v12026
  55. Sharp Minimax Regret for Infinite-Memory Logistic Prediction

    Vaneet Aggarwal

    cs.ITcs.LGarXiv:2608.26515v12026
  56. Modeling spatio-temporal locality in multi-step forecasting of geo-referenced time series

    Annunziata D'Aversa, Gianvito Pio, Michelangelo Ceci

    cs.LGarXiv:2608.25698v12026
  57. Tropospheric temperature and humidity profile retrieval from Meteosat Flexible Combined Imager based on deep learning

    Alejandro Salgueiro, Johannes Rausch, Julie Thérèse Villinger +1

    cs.LGphysics.ao-pharXiv:2608.25700v12026
  58. LION: A Clifford Neural Paradigm for Multimodal-Attributed Graph Learning

    Xunkai Li, Zekai Chen, Zhengyu Wu +6

    cs.LGarXiv:2608.24795v12026
  59. Decoupling candidate dual AGN from chance superpositions in the GOTHIC survey via a deep-learning framework

    Bhavesh Mukheja, Snehanshu Saha, Anwesh Bhattacharya +3

    astro-ph.GAcs.CVcs.LGarXiv:2608.24164v12026
  60. Optimal Alternating Regret for Online Learning and Games

    Yixin Tao, Weiqiang Zheng

    cs.LGcs.GTstat.MLarXiv:2608.24731v12026