Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
241 to 300 of 19,938
q-Learning in Continuous Time
Yanwei Jia, Xun Yu Zhou
cs.LGcs.AIq-fin.CParXiv:2207.00713v42022AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs
Florian Grötschla, Luis Müller, Jan Tönshoff +2
cs.MAcs.LGarXiv:2507.08616v12025ResearchTown: Simulator of Human Research Community
Haofei Yu, Zhaochen Hong, Zirui Cheng +5
cs.CLcs.LGarXiv:2412.17767v22024STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems
Akash Bonagiri, Gerard Janno Anderias, Saee Patil +6
cs.LGcs.AIarXiv:2605.02122v22026Order in the Court: Explainable AI Methods Prone to Disagreement
Michael Neely, Stefan F. Schouten, Maurits J. R. Bleeker +1
cs.LGcs.CLarXiv:2105.03287v32021RepMLPNet: Hierarchical Vision MLP with Re-parameterized Locality
Xiaohan Ding, Honghao Chen, Xiangyu Zhang +2
cs.CVcs.AIcs.LGarXiv:2112.11081v22021When to Show a Suggestion? Integrating Human Feedback in AI-Assisted Programming
Hussein Mozannar, Gagan Bansal, Adam Fourney +1
cs.HCcs.LGcs.SEarXiv:2306.04930v32023Summaries:한국어State of the Art Control of Atari Games Using Shallow Reinforcement Learning
Yitao Liang, Marlos C. Machado, Erik Talvitie +1
cs.LGarXiv:1512.01563v22015SoundReactor: Frame-level Online Video-to-Audio Generation
Koichi Saito, Julian Tanke, Christian Simon +7
cs.SDcs.LGeess.ASarXiv:2510.02110v12025Fair and Diverse DPP-based Data Summarization
L. Elisa Celis, Vijay Keswani, Damian Straszak +3
cs.LGcs.CYcs.IRarXiv:1802.04023v12018Progressive Distillation for Fast Sampling of Diffusion Models
Tim Salimans, Jonathan Ho
cs.LGcs.AIstat.MLarXiv:2202.00512v22022Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto +6
cs.LGarXiv:1611.01838v42016Learning to Reason for Hallucination Span Detection
Hsuan Su, Ting-Yao Hu, Hema Swetha Koppula +7
cs.CLcs.AIcs.LGarXiv:2510.02173v22025R-Transformer: Recurrent Neural Network Enhanced Transformer
Zhiwei Wang, Yao Ma, Zitao Liu +1
cs.LGcs.CLcs.CVarXiv:1907.05572v12019A Near-Linear Time Algorithm for the Chamfer Distance
Ainesh Bakshi, Piotr Indyk, Rajesh Jayaram +2
cs.DScs.CGcs.GRarXiv:2307.03043v12023Lipschitz Bandits with Stochastic Delayed Feedback
Zhongxuan Liu, Yue Kang, Thomas C. M. Lee
cs.LGstat.MLarXiv:2510.00309v22025Measuring training variability from stochastic optimization using robust nonparametric testing
Sinjini Banerjee, Tim Marrinan, Reilly Cannon +2
stat.MLcs.LGarXiv:2406.08307v22024Normalization Propagation: A Parametric Technique for Removing Internal Covariate Shift in Deep Networks
Devansh Arpit, Yingbo Zhou, Bhargava U. Kota +1
stat.MLcs.LGarXiv:1603.01431v62016DiagrammerGPT: Generating Open-Domain, Open-Platform Diagrams via LLM Planning
Abhay Zala, Han Lin, Jaemin Cho +1
cs.CVcs.AIcs.CLarXiv:2310.12128v22023Weight decay induces low-rank attention layers
Seijin Kobayashi, Yassir Akram, Johannes Von Oswald
cs.LGarXiv:2410.23819v12024Think in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using Agents
Utsav Garg, Sungjin Hong, Jason Jung +6
cs.AIcs.LGarXiv:2606.31648v12026Online convex optimization in the bandit setting: gradient descent without a gradient
Abraham D. Flaxman, Adam Tauman Kalai, H. Brendan McMahan
cs.LGcs.CCarXiv:cs/0408007v12004Tight Differential Privacy for Discrete-Valued Mechanisms and for the Subsampled Gaussian Mechanism Using FFT
Antti Koskela, Joonas Jälkö, Lukas Prediger +1
stat.MLcs.CRcs.LGarXiv:2006.07134v32020Addressing Some Limitations of Transformers with Feedback Memory
Angela Fan, Thibaut Lavril, Edouard Grave +2
cs.LGcs.CLstat.MLarXiv:2002.09402v32020Fair Adversarial Gradient Tree Boosting
Vincent Grari, Boris Ruf, Sylvain Lamprier +1
cs.LGcs.AIcs.CYarXiv:1911.05369v22019Estimating individual treatment effect: generalization bounds and algorithms
Uri Shalit, Fredrik D. Johansson, David Sontag
stat.MLcs.AIcs.LGarXiv:1606.03976v52016Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
Xin-Qiang Cai, Wei Wang, Feng Liu +3
cs.LGcs.AIarXiv:2510.00915v42025Geometric Operator Learning with Optimal Transport
Xinyi Li, Zongyi Li, Nikola Kovachki +1
cs.LGarXiv:2507.20065v12025OceanLight: Efficient Global Ocean Forecasting via Geometry-Adaptive Unstructured Mesh Representation
Wei Wu, Xiang Wang, Hongze Leng +3
cs.LGcs.AIarXiv:2608.16070v12026Paper2Agent: Reimagining Research Papers As Interactive and Reliable AI Agents
Jiacheng Miao, Joe R. Davis, Yaohui Zhang +2
cs.AIcs.CLcs.LGarXiv:2509.06917v22025Multimodal Whole Slide Foundation Model for Pathology
Tong Ding, Sophia J. Wagner, Andrew H. Song +20
eess.IVcs.AIcs.CVarXiv:2411.19666v12024DeBERTa: Decoding-enhanced BERT with Disentangled Attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao +1
cs.CLcs.LGarXiv:2006.03654v62020On the Value of Out-of-Distribution Testing: An Example of Goodhart's Law
Damien Teney, Kushal Kafle, Robik Shrestha +3
cs.CVcs.LGarXiv:2005.09241v12020Smoothed Dilated Convolutions for Improved Dense Prediction
Zhengyang Wang, Shuiwang Ji
cs.CVcs.LGarXiv:1808.08931v22018Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Evan Hubinger, Carson Denison, Jesse Mu +36
cs.CRcs.AIcs.CLarXiv:2401.05566v32024Addressing the Item Cold-start Problem by Attribute-driven Active Learning
Yu Zhu, Jinhao Lin, Shibi He +4
cs.IRcs.LGstat.MLarXiv:1805.09023v12018MotifNet: a motif-based Graph Convolutional Network for directed graphs
Federico Monti, Karl Otness, Michael M. Bronstein
cs.LGarXiv:1802.01572v12018Wav2Letter: an End-to-End ConvNet-based Speech Recognition System
Ronan Collobert, Christian Puhrsch, Gabriel Synnaeve
cs.LGcs.AIcs.CLarXiv:1609.03193v22016Sparse MoEs meet Efficient Ensembles
James Urquhart Allingham, Florian Wenzel, Zelda E Mariet +10
cs.LGcs.CVstat.MLarXiv:2110.03360v22021Graph of Thoughts: Solving Elaborate Problems with Large Language Models
Maciej Besta, Nils Blach, Ales Kubicek +8
cs.CLcs.AIcs.LGarXiv:2308.09687v42023The Impact of Positional Encoding on Length Generalization in Transformers
Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy +2
cs.CLcs.AIcs.LGarXiv:2305.19466v22023ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation
Zhengyi Wang, Cheng Lu, Yikai Wang +4
cs.LGcs.CVarXiv:2305.16213v22023Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
Tony Z. Zhao, Vikash Kumar, Sergey Levine +1
cs.ROcs.LGarXiv:2304.13705v12023StoRM: A Diffusion-based Stochastic Regeneration Model for Speech Enhancement and Dereverberation
Jean-Marie Lemercier, Julius Richter, Simon Welker +1
eess.AScs.LGcs.SDarXiv:2212.11851v22022Beyond neural scaling laws: beating power law scaling via data pruning
Ben Sorscher, Robert Geirhos, Shashank Shekhar +2
cs.LGcs.AIcs.CVarXiv:2206.14486v62022FedBABU: Towards Enhanced Representation for Federated Image Classification
Jaehoon Oh, Sangmook Kim, Se-Young Yun
cs.LGarXiv:2106.06042v32021Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
A. Sophia Koepke, Daniil Zverev, Shiry Ginosar +1
cs.CVcs.AIcs.LGarXiv:2604.18572v22026Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
Harsh Kohli, Srinivasan Parthasarathy, Huan Sun +1
cs.CLcs.AIcs.LGarXiv:2604.07822v22026PLDR-LLMs Reason At Self-Organized Criticality
Burc Gokden
cs.AIcs.CLcs.LGarXiv:2603.23539v12026A Neurosymbolic Approach for Constructing Planning Domain Models from Clinical Narratives
Ranveer Singh, Saurabh Mathur, Michael Skinner +3
cs.LGarXiv:2608.21186v12026NerVE: Nonlinear Eigenspectrum Dynamics in LLM Feed-Forward Networks
Nandan Kumar Jha, Brandon Reagen
cs.LGarXiv:2603.06922v22026Reinforcement Learning for Code Optimization
Pierre Chambon, Kunhao Zheng, Juliette Decugis +2
cs.LGcs.AIarXiv:2607.25970v12026A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects
Zewen Li, Wenjie Yang, Shouheng Peng +1
cs.CVcs.LGeess.IVarXiv:2004.02806v12020SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?
Azmine Toushik Wasi, Wahid Faisal, Abdur Rahman +12
cs.CVcs.CEcs.CLarXiv:2602.03916v32026SINDy-PI: A Robust Algorithm for Parallel Implicit Sparse Identification of Nonlinear Dynamics
Kadierdan Kaheman, J. Nathan Kutz, Steven L. Brunton
cs.LGphysics.comp-phstat.MLarXiv:2004.02322v22020Meta Label Correction for Noisy Label Learning
Guoqing Zheng, Ahmed Hassan Awadallah, Susan Dumais
cs.LGstat.MLarXiv:1911.03809v22019The Born Supremacy: Quantum Advantage and Training of an Ising Born Machine
Brian Coyle, Daniel Mills, Vincent Danos +1
quant-phcs.LGarXiv:1904.02214v42019Federated Optimization in Heterogeneous Networks
Tian Li, Anit Kumar Sahu, Manzil Zaheer +3
cs.LGstat.MLarXiv:1812.06127v52018Automating the Design of Embodied Agent Architectures
Jian Zhou, Sihao Lin, Jin Li +3
cs.ROcs.AIcs.LGarXiv:2606.30111v22026TROPT: An Open Framework for Unifying and Advancing Discrete Text Optimization
Matan Ben-Tov, Mahmood Sharif
cs.LGcs.CRarXiv:2606.23496v12026