Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

18,241 to 18,300 of 20,199

  1. One weird trick for parallelizing convolutional neural networks

    Alex Krizhevsky

    cs.NEcs.DCcs.LGarXiv:1404.5997v22014
  2. Convolutional Radio Modulation Recognition Networks

    Timothy J O'Shea, Johnathan Corgan, T. Charles Clancy

    cs.LGcs.CVarXiv:1602.04105v32016
  3. Horovod: fast and easy distributed deep learning in TensorFlow

    Alexander Sergeev, Mike Del Balso

    cs.LGstat.MLarXiv:1802.05799v32018
  4. N-BaIoT: Network-based Detection of IoT Botnet Attacks Using Deep Autoencoders

    Yair Meidan, Michael Bohadana, Yael Mathov +4

    cs.CRcs.LGarXiv:1805.03409v12018
  5. Federated Learning with Matched Averaging

    Hongyi Wang, Mikhail Yurochkin, Yuekai Sun +2

    cs.LGstat.MLarXiv:2002.06440v12020
  6. Dark Experience for General Continual Learning: a Strong, Simple Baseline

    Pietro Buzzega, Matteo Boschini, Angelo Porrello +2

    stat.MLcs.LGarXiv:2004.07211v22020
  7. Deep Learning in Bioinformatics

    Seonwoo Min, Byunghan Lee, Sungroh Yoon

    cs.LGq-bio.GNarXiv:1603.06430v52016
  8. Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression

    Yao Du, Shanshan Song, Xiaomeng Li

    cs.CLcs.CVcs.LGarXiv:2605.01402v22026
  9. DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices

    Chenyang Song, Weilin Zhao, Xu Han +3

    cs.LGcs.CLarXiv:2605.10933v32026
  10. Unsupervised Process Reward Models

    Artyom Gadetsky, Maxim Kodryan, Siba Smarak Panigrahi +2

    cs.LGarXiv:2605.10158v12026
  11. Neural Factorization Machines for Sparse Predictive Analytics

    Xiangnan He, Tat-Seng Chua

    cs.LGarXiv:1708.05027v12017
  12. Urban-ImageNet: A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception

    Yiwei Ou, Chung Ching Cheung, Jun Yang Ang +5

    cs.CVcs.IRcs.LGarXiv:2605.09936v12026
  13. Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning

    Haokun Liu, Derek Tam, Mohammed Muqeeth +4

    cs.LGcs.AIcs.CLarXiv:2205.05638v22022
  14. On the Number of Linear Regions of Deep Neural Networks

    Guido Montúfar, Razvan Pascanu, Kyunghyun Cho +1

    stat.MLcs.LGcs.NEarXiv:1402.1869v22014
  15. Semi-Supervised Learning with Ladder Networks

    Antti Rasmus, Harri Valpola, Mikko Honkala +2

    cs.NEcs.LGstat.MLarXiv:1507.02672v22015
  16. Theano: new features and speed improvements

    Frédéric Bastien, Pascal Lamblin, Razvan Pascanu +6

    cs.SCcs.LGarXiv:1211.5590v12012
  17. Deep Learning for Anomaly Detection: A Review

    Guansong Pang, Chunhua Shen, Longbing Cao +1

    cs.LGcs.CVstat.MLarXiv:2007.02500v32020
  18. SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding

    Anton Plaksin, Sergei Krutikov, Sergei Skvortsov +1

    cs.LGcs.CLarXiv:2605.10453v12026
  19. Flood Prediction Using Machine Learning Models: Literature Review

    Amir Mosavi, Pinar Ozturk, Kwok-wing Chau

    cs.LGstat.MLarXiv:1908.02781v12019
  20. Explaining Machine Learning Classifiers through Diverse Counterfactual Explanations

    Ramaravind Kommiya Mothilal, Amit Sharma, Chenhao Tan

    cs.LGcs.CYstat.MLarXiv:1905.07697v22019
  21. Measuring and Relieving the Over-smoothing Problem for Graph Neural Networks from the Topological View

    Deli Chen, Yankai Lin, Wei Li +3

    cs.LGcs.SIstat.MLarXiv:1909.03211v22019
  22. Modeling polypharmacy side effects with graph convolutional networks

    Marinka Zitnik, Monica Agrawal, Jure Leskovec

    cs.LGq-bio.MNstat.MLarXiv:1802.00543v22018
  23. Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory

    Daniel Goldstein, Navneel Singhal, Eugene Cheah

    cs.LGcs.AIcs.CLarXiv:2605.09877v52026
  24. Active Tabular Augmentation via Policy-Guided Diffusion Inpainting

    Zheyu Zhang, Shuo Yang, Bardh Prenkaj +1

    cs.LGcs.AIarXiv:2605.10315v12026
  25. GLiNER-Relex: A Unified Framework for Joint Named Entity Recognition and Relation Extraction

    Ihor Stepanov, Oleksandr Lukashov, Mykhailo Shtopko +1

    cs.CLcs.LGarXiv:2605.10108v12026
  26. GAIN: Missing Data Imputation using Generative Adversarial Nets

    Jinsung Yoon, James Jordon, Mihaela van der Schaar

    cs.LGstat.MLarXiv:1806.02920v12018
  27. AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward

    Runhui Huang, Jie Wu, Rui Yang +2

    cs.CVcs.AIcs.LGarXiv:2605.12495v12026
  28. Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics

    Jishnu Sethumadhavan Nair, Patrice Bechard, Rishabh Maheshwary +14

    cs.AIcs.CLcs.LGarXiv:2605.12178v12026
  29. CoQA: A Conversational Question Answering Challenge

    Siva Reddy, Danqi Chen, Christopher D. Manning

    cs.CLcs.AIcs.LGarXiv:1808.07042v22018
  30. Learning to Communicate Locally for Large-Scale Multi-Agent Pathfinding

    Valeriy Vyaltsev, Alsu Sagirova, Anton Andreychuk +5

    cs.AIcs.LGcs.MAarXiv:2605.07637v22026
  31. Predicting Decisions of AI Agents from Limited Interaction through Text-Tabular Modeling

    Eilam Shapira, Moshe Tennenholtz, Roi Reichart

    cs.LGcs.AIcs.CLarXiv:2605.12411v12026
  32. End-to-end representation learning for Correlation Filter based tracking

    Jack Valmadre, Luca Bertinetto, João F. Henriques +2

    cs.CVcs.LGarXiv:1704.06036v12017
  33. Sequence-Level Knowledge Distillation

    Yoon Kim, Alexander M. Rush

    cs.CLcs.LGcs.NEarXiv:1606.07947v42016
  34. Over the Air Deep Learning Based Radio Signal Classification

    Timothy J. O'Shea, Tamoghna Roy, T. Charles Clancy

    cs.LGeess.SParXiv:1712.04578v12017
  35. Deep clustering: Discriminative embeddings for segmentation and separation

    John R. Hershey, Zhuo Chen, Jonathan Le Roux +1

    cs.NEcs.LGstat.MLarXiv:1508.04306v12015
  36. Debiased Model-based Representations for Sample-efficient Continuous Control

    Jiafei Lyu, Zichuan Lin, Scott Fujimoto +5

    cs.LGcs.AIarXiv:2605.11711v12026
  37. Exploring Generalization in Deep Learning

    Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester +1

    cs.LGarXiv:1706.08947v22017
  38. Meta-Learning with Differentiable Convex Optimization

    Kwonjoon Lee, Subhransu Maji, Avinash Ravichandran +1

    cs.CVcs.LGarXiv:1904.03758v22019
  39. Meta-Learning for Semi-Supervised Few-Shot Classification

    Mengye Ren, Eleni Triantafillou, Sachin Ravi +5

    cs.LGcs.CVstat.MLarXiv:1803.00676v12018
  40. Improved Knowledge Distillation via Teacher Assistant

    Seyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li +3

    cs.LGcs.AIstat.MLarXiv:1902.03393v22019
  41. WriteSAE: Sparse Autoencoders for Recurrent State

    Jack Young

    cs.LGcs.AIcs.CLarXiv:2605.12770v42026
  42. Large-Scale Long-Tailed Recognition in an Open World

    Ziwei Liu, Zhongqi Miao, Xiaohang Zhan +3

    cs.CVcs.LGarXiv:1904.05160v22019
  43. Concept Bottleneck Models

    Pang Wei Koh, Thao Nguyen, Yew Siang Tang +4

    cs.LGstat.MLarXiv:2007.04612v32020
  44. Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion

    Chien Van Nguyen, Chaitra Hegde, Van Cuong Pham +3

    cs.LGcs.AIarXiv:2605.12825v22026
  45. Ditto: Fair and Robust Federated Learning Through Personalization

    Tian Li, Shengyuan Hu, Ahmad Beirami +1

    cs.LGstat.MLarXiv:2012.04221v32020
  46. Scaled-YOLOv4: Scaling Cross Stage Partial Network

    Chien-Yao Wang, Alexey Bochkovskiy, Hong-Yuan Mark Liao

    cs.CVcs.LGarXiv:2011.08036v22020
  47. BloombergGPT: A Large Language Model for Finance

    Shijie Wu, Ozan Irsoy, Steven Lu +6

    cs.LGcs.AIcs.CLarXiv:2303.17564v32023
  48. DocAtlas: Multilingual Document Understanding Across 80+ Languages

    Ahmed Heakl, Youssef Mohamed, Abdullah Sohail +6

    cs.CLcs.CVcs.LGarXiv:2605.12623v22026
  49. MedMNIST v2 -- A large-scale lightweight benchmark for 2D and 3D biomedical image classification

    Jiancheng Yang, Rui Shi, Donglai Wei +5

    cs.CVcs.AIcs.LGarXiv:2110.14795v22021
  50. MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities

    Mingqiao Ye, Zhaochong An, Zhitong Gao +11

    cs.CVcs.AIcs.LGarXiv:2607.25948v12026
  51. Frequency Bias and OOD Generalization in Neural Operators under a Variable-Coefficient Wave Equation

    Runlong Xie, An Luo

    cs.LGarXiv:2605.12997v12026
  52. Code-Guided Reasoning for Small Language Models: Evaluating Executable MCQA Scaffolds

    Prateek Biswas, Dhaval Patel, Vedant Khandelwal +2

    cs.IRcs.LGcs.PLarXiv:2605.18827v12026
  53. SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot

    Elias Frantar, Dan Alistarh

    cs.LGarXiv:2301.00774v32023
  54. Financial Time Series Forecasting with Deep Learning : A Systematic Literature Review: 2005-2019

    Omer Berat Sezer, Mehmet Ugur Gudelek, Ahmet Murat Ozbayoglu

    cs.LGq-fin.CPstat.MLarXiv:1911.13288v12019
  55. Large-scale Point Cloud Semantic Segmentation with Superpoint Graphs

    Loic Landrieu, Martin Simonovsky

    cs.CVcs.LGcs.NEarXiv:1711.09869v22017
  56. Diffusion-LM Improves Controllable Text Generation

    Xiang Lisa Li, John Thickstun, Ishaan Gulrajani +2

    cs.CLcs.AIcs.LGarXiv:2205.14217v12022
  57. Do Vision Transformers See Like Convolutional Neural Networks?

    Maithra Raghu, Thomas Unterthiner, Simon Kornblith +2

    cs.CVcs.AIcs.LGarXiv:2108.08810v22021
  58. StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery

    Or Patashnik, Zongze Wu, Eli Shechtman +2

    cs.CVcs.CLcs.GRarXiv:2103.17249v12021
  59. LoREnc: Low-Rank Encryption for Securing Foundation Models and LoRA Adapters

    Beomjin Ahn, Jungmin Kwon, Chanyong Jung +1

    cs.CRcs.CVcs.LGarXiv:2605.13163v12026
  60. Training Large Language Models to Predict Clinical Events

    Benjamin Turtel, Paul Wilczewski, Kris Skotheim

    cs.LGcs.AIcs.CLarXiv:2605.12817v12026