Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

241 to 300 of 19,938

  1. q-Learning in Continuous Time

    Yanwei Jia, Xun Yu Zhou

    cs.LGcs.AIq-fin.CParXiv:2207.00713v42022
  2. AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs

    Florian Grötschla, Luis Müller, Jan Tönshoff +2

    cs.MAcs.LGarXiv:2507.08616v12025
  3. ResearchTown: Simulator of Human Research Community

    Haofei Yu, Zhaochen Hong, Zirui Cheng +5

    cs.CLcs.LGarXiv:2412.17767v22024
  4. STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems

    Akash Bonagiri, Gerard Janno Anderias, Saee Patil +6

    cs.LGcs.AIarXiv:2605.02122v22026
  5. Order in the Court: Explainable AI Methods Prone to Disagreement

    Michael Neely, Stefan F. Schouten, Maurits J. R. Bleeker +1

    cs.LGcs.CLarXiv:2105.03287v32021
  6. RepMLPNet: Hierarchical Vision MLP with Re-parameterized Locality

    Xiaohan Ding, Honghao Chen, Xiangyu Zhang +2

    cs.CVcs.AIcs.LGarXiv:2112.11081v22021
  7. When to Show a Suggestion? Integrating Human Feedback in AI-Assisted Programming

    Hussein Mozannar, Gagan Bansal, Adam Fourney +1

    cs.HCcs.LGcs.SEarXiv:2306.04930v32023
    Summaries:한국어
  8. State of the Art Control of Atari Games Using Shallow Reinforcement Learning

    Yitao Liang, Marlos C. Machado, Erik Talvitie +1

    cs.LGarXiv:1512.01563v22015
  9. SoundReactor: Frame-level Online Video-to-Audio Generation

    Koichi Saito, Julian Tanke, Christian Simon +7

    cs.SDcs.LGeess.ASarXiv:2510.02110v12025
  10. Fair and Diverse DPP-based Data Summarization

    L. Elisa Celis, Vijay Keswani, Damian Straszak +3

    cs.LGcs.CYcs.IRarXiv:1802.04023v12018
  11. Progressive Distillation for Fast Sampling of Diffusion Models

    Tim Salimans, Jonathan Ho

    cs.LGcs.AIstat.MLarXiv:2202.00512v22022
  12. Entropy-SGD: Biasing Gradient Descent Into Wide Valleys

    Pratik Chaudhari, Anna Choromanska, Stefano Soatto +6

    cs.LGarXiv:1611.01838v42016
  13. Learning to Reason for Hallucination Span Detection

    Hsuan Su, Ting-Yao Hu, Hema Swetha Koppula +7

    cs.CLcs.AIcs.LGarXiv:2510.02173v22025
  14. R-Transformer: Recurrent Neural Network Enhanced Transformer

    Zhiwei Wang, Yao Ma, Zitao Liu +1

    cs.LGcs.CLcs.CVarXiv:1907.05572v12019
  15. A Near-Linear Time Algorithm for the Chamfer Distance

    Ainesh Bakshi, Piotr Indyk, Rajesh Jayaram +2

    cs.DScs.CGcs.GRarXiv:2307.03043v12023
  16. Lipschitz Bandits with Stochastic Delayed Feedback

    Zhongxuan Liu, Yue Kang, Thomas C. M. Lee

    cs.LGstat.MLarXiv:2510.00309v22025
  17. Measuring training variability from stochastic optimization using robust nonparametric testing

    Sinjini Banerjee, Tim Marrinan, Reilly Cannon +2

    stat.MLcs.LGarXiv:2406.08307v22024
  18. Normalization Propagation: A Parametric Technique for Removing Internal Covariate Shift in Deep Networks

    Devansh Arpit, Yingbo Zhou, Bhargava U. Kota +1

    stat.MLcs.LGarXiv:1603.01431v62016
  19. DiagrammerGPT: Generating Open-Domain, Open-Platform Diagrams via LLM Planning

    Abhay Zala, Han Lin, Jaemin Cho +1

    cs.CVcs.AIcs.CLarXiv:2310.12128v22023
  20. Weight decay induces low-rank attention layers

    Seijin Kobayashi, Yassir Akram, Johannes Von Oswald

    cs.LGarXiv:2410.23819v12024
  21. Think in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using Agents

    Utsav Garg, Sungjin Hong, Jason Jung +6

    cs.AIcs.LGarXiv:2606.31648v12026
  22. Online convex optimization in the bandit setting: gradient descent without a gradient

    Abraham D. Flaxman, Adam Tauman Kalai, H. Brendan McMahan

    cs.LGcs.CCarXiv:cs/0408007v12004
  23. Tight Differential Privacy for Discrete-Valued Mechanisms and for the Subsampled Gaussian Mechanism Using FFT

    Antti Koskela, Joonas Jälkö, Lukas Prediger +1

    stat.MLcs.CRcs.LGarXiv:2006.07134v32020
  24. Addressing Some Limitations of Transformers with Feedback Memory

    Angela Fan, Thibaut Lavril, Edouard Grave +2

    cs.LGcs.CLstat.MLarXiv:2002.09402v32020
  25. Fair Adversarial Gradient Tree Boosting

    Vincent Grari, Boris Ruf, Sylvain Lamprier +1

    cs.LGcs.AIcs.CYarXiv:1911.05369v22019
  26. Estimating individual treatment effect: generalization bounds and algorithms

    Uri Shalit, Fredrik D. Johansson, David Sontag

    stat.MLcs.AIcs.LGarXiv:1606.03976v52016
  27. Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

    Xin-Qiang Cai, Wei Wang, Feng Liu +3

    cs.LGcs.AIarXiv:2510.00915v42025
  28. Geometric Operator Learning with Optimal Transport

    Xinyi Li, Zongyi Li, Nikola Kovachki +1

    cs.LGarXiv:2507.20065v12025
  29. OceanLight: Efficient Global Ocean Forecasting via Geometry-Adaptive Unstructured Mesh Representation

    Wei Wu, Xiang Wang, Hongze Leng +3

    cs.LGcs.AIarXiv:2608.16070v12026
  30. Paper2Agent: Reimagining Research Papers As Interactive and Reliable AI Agents

    Jiacheng Miao, Joe R. Davis, Yaohui Zhang +2

    cs.AIcs.CLcs.LGarXiv:2509.06917v22025
  31. Multimodal Whole Slide Foundation Model for Pathology

    Tong Ding, Sophia J. Wagner, Andrew H. Song +20

    eess.IVcs.AIcs.CVarXiv:2411.19666v12024
  32. DeBERTa: Decoding-enhanced BERT with Disentangled Attention

    Pengcheng He, Xiaodong Liu, Jianfeng Gao +1

    cs.CLcs.LGarXiv:2006.03654v62020
  33. On the Value of Out-of-Distribution Testing: An Example of Goodhart's Law

    Damien Teney, Kushal Kafle, Robik Shrestha +3

    cs.CVcs.LGarXiv:2005.09241v12020
  34. Smoothed Dilated Convolutions for Improved Dense Prediction

    Zhengyang Wang, Shuiwang Ji

    cs.CVcs.LGarXiv:1808.08931v22018
  35. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

    Evan Hubinger, Carson Denison, Jesse Mu +36

    cs.CRcs.AIcs.CLarXiv:2401.05566v32024
  36. Addressing the Item Cold-start Problem by Attribute-driven Active Learning

    Yu Zhu, Jinhao Lin, Shibi He +4

    cs.IRcs.LGstat.MLarXiv:1805.09023v12018
  37. MotifNet: a motif-based Graph Convolutional Network for directed graphs

    Federico Monti, Karl Otness, Michael M. Bronstein

    cs.LGarXiv:1802.01572v12018
  38. Wav2Letter: an End-to-End ConvNet-based Speech Recognition System

    Ronan Collobert, Christian Puhrsch, Gabriel Synnaeve

    cs.LGcs.AIcs.CLarXiv:1609.03193v22016
  39. Sparse MoEs meet Efficient Ensembles

    James Urquhart Allingham, Florian Wenzel, Zelda E Mariet +10

    cs.LGcs.CVstat.MLarXiv:2110.03360v22021
  40. Graph of Thoughts: Solving Elaborate Problems with Large Language Models

    Maciej Besta, Nils Blach, Ales Kubicek +8

    cs.CLcs.AIcs.LGarXiv:2308.09687v42023
  41. The Impact of Positional Encoding on Length Generalization in Transformers

    Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy +2

    cs.CLcs.AIcs.LGarXiv:2305.19466v22023
  42. ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation

    Zhengyi Wang, Cheng Lu, Yikai Wang +4

    cs.LGcs.CVarXiv:2305.16213v22023
  43. Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

    Tony Z. Zhao, Vikash Kumar, Sergey Levine +1

    cs.ROcs.LGarXiv:2304.13705v12023
  44. StoRM: A Diffusion-based Stochastic Regeneration Model for Speech Enhancement and Dereverberation

    Jean-Marie Lemercier, Julius Richter, Simon Welker +1

    eess.AScs.LGcs.SDarXiv:2212.11851v22022
  45. Beyond neural scaling laws: beating power law scaling via data pruning

    Ben Sorscher, Robert Geirhos, Shashank Shekhar +2

    cs.LGcs.AIcs.CVarXiv:2206.14486v62022
  46. FedBABU: Towards Enhanced Representation for Federated Image Classification

    Jaehoon Oh, Sangmook Kim, Se-Young Yun

    cs.LGarXiv:2106.06042v32021
  47. Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

    A. Sophia Koepke, Daniil Zverev, Shiry Ginosar +1

    cs.CVcs.AIcs.LGarXiv:2604.18572v22026
  48. Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers

    Harsh Kohli, Srinivasan Parthasarathy, Huan Sun +1

    cs.CLcs.AIcs.LGarXiv:2604.07822v22026
  49. PLDR-LLMs Reason At Self-Organized Criticality

    Burc Gokden

    cs.AIcs.CLcs.LGarXiv:2603.23539v12026
  50. A Neurosymbolic Approach for Constructing Planning Domain Models from Clinical Narratives

    Ranveer Singh, Saurabh Mathur, Michael Skinner +3

    cs.LGarXiv:2608.21186v12026
  51. NerVE: Nonlinear Eigenspectrum Dynamics in LLM Feed-Forward Networks

    Nandan Kumar Jha, Brandon Reagen

    cs.LGarXiv:2603.06922v22026
  52. Reinforcement Learning for Code Optimization

    Pierre Chambon, Kunhao Zheng, Juliette Decugis +2

    cs.LGcs.AIarXiv:2607.25970v12026
  53. A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects

    Zewen Li, Wenjie Yang, Shouheng Peng +1

    cs.CVcs.LGeess.IVarXiv:2004.02806v12020
  54. SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?

    Azmine Toushik Wasi, Wahid Faisal, Abdur Rahman +12

    cs.CVcs.CEcs.CLarXiv:2602.03916v32026
  55. SINDy-PI: A Robust Algorithm for Parallel Implicit Sparse Identification of Nonlinear Dynamics

    Kadierdan Kaheman, J. Nathan Kutz, Steven L. Brunton

    cs.LGphysics.comp-phstat.MLarXiv:2004.02322v22020
  56. Meta Label Correction for Noisy Label Learning

    Guoqing Zheng, Ahmed Hassan Awadallah, Susan Dumais

    cs.LGstat.MLarXiv:1911.03809v22019
  57. The Born Supremacy: Quantum Advantage and Training of an Ising Born Machine

    Brian Coyle, Daniel Mills, Vincent Danos +1

    quant-phcs.LGarXiv:1904.02214v42019
  58. Federated Optimization in Heterogeneous Networks

    Tian Li, Anit Kumar Sahu, Manzil Zaheer +3

    cs.LGstat.MLarXiv:1812.06127v52018
  59. Automating the Design of Embodied Agent Architectures

    Jian Zhou, Sihao Lin, Jin Li +3

    cs.ROcs.AIcs.LGarXiv:2606.30111v22026
  60. TROPT: An Open Framework for Unifying and Advancing Discrete Text Optimization

    Matan Ben-Tov, Mahmood Sharif

    cs.LGcs.CRarXiv:2606.23496v12026