Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
18,241 to 18,300 of 20,199
One weird trick for parallelizing convolutional neural networks
Alex Krizhevsky
cs.NEcs.DCcs.LGarXiv:1404.5997v22014Convolutional Radio Modulation Recognition Networks
Timothy J O'Shea, Johnathan Corgan, T. Charles Clancy
cs.LGcs.CVarXiv:1602.04105v32016Horovod: fast and easy distributed deep learning in TensorFlow
Alexander Sergeev, Mike Del Balso
cs.LGstat.MLarXiv:1802.05799v32018N-BaIoT: Network-based Detection of IoT Botnet Attacks Using Deep Autoencoders
Yair Meidan, Michael Bohadana, Yael Mathov +4
cs.CRcs.LGarXiv:1805.03409v12018Federated Learning with Matched Averaging
Hongyi Wang, Mikhail Yurochkin, Yuekai Sun +2
cs.LGstat.MLarXiv:2002.06440v12020Dark Experience for General Continual Learning: a Strong, Simple Baseline
Pietro Buzzega, Matteo Boschini, Angelo Porrello +2
stat.MLcs.LGarXiv:2004.07211v22020Deep Learning in Bioinformatics
Seonwoo Min, Byunghan Lee, Sungroh Yoon
cs.LGq-bio.GNarXiv:1603.06430v52016Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression
Yao Du, Shanshan Song, Xiaomeng Li
cs.CLcs.CVcs.LGarXiv:2605.01402v22026DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
Chenyang Song, Weilin Zhao, Xu Han +3
cs.LGcs.CLarXiv:2605.10933v32026Unsupervised Process Reward Models
Artyom Gadetsky, Maxim Kodryan, Siba Smarak Panigrahi +2
cs.LGarXiv:2605.10158v12026Neural Factorization Machines for Sparse Predictive Analytics
Xiangnan He, Tat-Seng Chua
cs.LGarXiv:1708.05027v12017Urban-ImageNet: A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception
Yiwei Ou, Chung Ching Cheung, Jun Yang Ang +5
cs.CVcs.IRcs.LGarXiv:2605.09936v12026Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning
Haokun Liu, Derek Tam, Mohammed Muqeeth +4
cs.LGcs.AIcs.CLarXiv:2205.05638v22022On the Number of Linear Regions of Deep Neural Networks
Guido Montúfar, Razvan Pascanu, Kyunghyun Cho +1
stat.MLcs.LGcs.NEarXiv:1402.1869v22014Semi-Supervised Learning with Ladder Networks
Antti Rasmus, Harri Valpola, Mikko Honkala +2
cs.NEcs.LGstat.MLarXiv:1507.02672v22015Theano: new features and speed improvements
Frédéric Bastien, Pascal Lamblin, Razvan Pascanu +6
cs.SCcs.LGarXiv:1211.5590v12012Deep Learning for Anomaly Detection: A Review
Guansong Pang, Chunhua Shen, Longbing Cao +1
cs.LGcs.CVstat.MLarXiv:2007.02500v32020SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding
Anton Plaksin, Sergei Krutikov, Sergei Skvortsov +1
cs.LGcs.CLarXiv:2605.10453v12026Flood Prediction Using Machine Learning Models: Literature Review
Amir Mosavi, Pinar Ozturk, Kwok-wing Chau
cs.LGstat.MLarXiv:1908.02781v12019Explaining Machine Learning Classifiers through Diverse Counterfactual Explanations
Ramaravind Kommiya Mothilal, Amit Sharma, Chenhao Tan
cs.LGcs.CYstat.MLarXiv:1905.07697v22019Measuring and Relieving the Over-smoothing Problem for Graph Neural Networks from the Topological View
Deli Chen, Yankai Lin, Wei Li +3
cs.LGcs.SIstat.MLarXiv:1909.03211v22019Modeling polypharmacy side effects with graph convolutional networks
Marinka Zitnik, Monica Agrawal, Jure Leskovec
cs.LGq-bio.MNstat.MLarXiv:1802.00543v22018Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory
Daniel Goldstein, Navneel Singhal, Eugene Cheah
cs.LGcs.AIcs.CLarXiv:2605.09877v52026Active Tabular Augmentation via Policy-Guided Diffusion Inpainting
Zheyu Zhang, Shuo Yang, Bardh Prenkaj +1
cs.LGcs.AIarXiv:2605.10315v12026GLiNER-Relex: A Unified Framework for Joint Named Entity Recognition and Relation Extraction
Ihor Stepanov, Oleksandr Lukashov, Mykhailo Shtopko +1
cs.CLcs.LGarXiv:2605.10108v12026GAIN: Missing Data Imputation using Generative Adversarial Nets
Jinsung Yoon, James Jordon, Mihaela van der Schaar
cs.LGstat.MLarXiv:1806.02920v12018AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward
Runhui Huang, Jie Wu, Rui Yang +2
cs.CVcs.AIcs.LGarXiv:2605.12495v12026Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics
Jishnu Sethumadhavan Nair, Patrice Bechard, Rishabh Maheshwary +14
cs.AIcs.CLcs.LGarXiv:2605.12178v12026CoQA: A Conversational Question Answering Challenge
Siva Reddy, Danqi Chen, Christopher D. Manning
cs.CLcs.AIcs.LGarXiv:1808.07042v22018Learning to Communicate Locally for Large-Scale Multi-Agent Pathfinding
Valeriy Vyaltsev, Alsu Sagirova, Anton Andreychuk +5
cs.AIcs.LGcs.MAarXiv:2605.07637v22026Predicting Decisions of AI Agents from Limited Interaction through Text-Tabular Modeling
Eilam Shapira, Moshe Tennenholtz, Roi Reichart
cs.LGcs.AIcs.CLarXiv:2605.12411v12026End-to-end representation learning for Correlation Filter based tracking
Jack Valmadre, Luca Bertinetto, João F. Henriques +2
cs.CVcs.LGarXiv:1704.06036v12017Sequence-Level Knowledge Distillation
Yoon Kim, Alexander M. Rush
cs.CLcs.LGcs.NEarXiv:1606.07947v42016Over the Air Deep Learning Based Radio Signal Classification
Timothy J. O'Shea, Tamoghna Roy, T. Charles Clancy
cs.LGeess.SParXiv:1712.04578v12017Deep clustering: Discriminative embeddings for segmentation and separation
John R. Hershey, Zhuo Chen, Jonathan Le Roux +1
cs.NEcs.LGstat.MLarXiv:1508.04306v12015Debiased Model-based Representations for Sample-efficient Continuous Control
Jiafei Lyu, Zichuan Lin, Scott Fujimoto +5
cs.LGcs.AIarXiv:2605.11711v12026Exploring Generalization in Deep Learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester +1
cs.LGarXiv:1706.08947v22017Meta-Learning with Differentiable Convex Optimization
Kwonjoon Lee, Subhransu Maji, Avinash Ravichandran +1
cs.CVcs.LGarXiv:1904.03758v22019Meta-Learning for Semi-Supervised Few-Shot Classification
Mengye Ren, Eleni Triantafillou, Sachin Ravi +5
cs.LGcs.CVstat.MLarXiv:1803.00676v12018Improved Knowledge Distillation via Teacher Assistant
Seyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li +3
cs.LGcs.AIstat.MLarXiv:1902.03393v22019WriteSAE: Sparse Autoencoders for Recurrent State
Jack Young
cs.LGcs.AIcs.CLarXiv:2605.12770v42026Large-Scale Long-Tailed Recognition in an Open World
Ziwei Liu, Zhongqi Miao, Xiaohang Zhan +3
cs.CVcs.LGarXiv:1904.05160v22019Concept Bottleneck Models
Pang Wei Koh, Thao Nguyen, Yew Siang Tang +4
cs.LGstat.MLarXiv:2007.04612v32020Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion
Chien Van Nguyen, Chaitra Hegde, Van Cuong Pham +3
cs.LGcs.AIarXiv:2605.12825v22026Ditto: Fair and Robust Federated Learning Through Personalization
Tian Li, Shengyuan Hu, Ahmad Beirami +1
cs.LGstat.MLarXiv:2012.04221v32020Scaled-YOLOv4: Scaling Cross Stage Partial Network
Chien-Yao Wang, Alexey Bochkovskiy, Hong-Yuan Mark Liao
cs.CVcs.LGarXiv:2011.08036v22020BloombergGPT: A Large Language Model for Finance
Shijie Wu, Ozan Irsoy, Steven Lu +6
cs.LGcs.AIcs.CLarXiv:2303.17564v32023DocAtlas: Multilingual Document Understanding Across 80+ Languages
Ahmed Heakl, Youssef Mohamed, Abdullah Sohail +6
cs.CLcs.CVcs.LGarXiv:2605.12623v22026MedMNIST v2 -- A large-scale lightweight benchmark for 2D and 3D biomedical image classification
Jiancheng Yang, Rui Shi, Donglai Wei +5
cs.CVcs.AIcs.LGarXiv:2110.14795v22021MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
Mingqiao Ye, Zhaochong An, Zhitong Gao +11
cs.CVcs.AIcs.LGarXiv:2607.25948v12026Frequency Bias and OOD Generalization in Neural Operators under a Variable-Coefficient Wave Equation
Runlong Xie, An Luo
cs.LGarXiv:2605.12997v12026Code-Guided Reasoning for Small Language Models: Evaluating Executable MCQA Scaffolds
Prateek Biswas, Dhaval Patel, Vedant Khandelwal +2
cs.IRcs.LGcs.PLarXiv:2605.18827v12026SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
Elias Frantar, Dan Alistarh
cs.LGarXiv:2301.00774v32023Financial Time Series Forecasting with Deep Learning : A Systematic Literature Review: 2005-2019
Omer Berat Sezer, Mehmet Ugur Gudelek, Ahmet Murat Ozbayoglu
cs.LGq-fin.CPstat.MLarXiv:1911.13288v12019Large-scale Point Cloud Semantic Segmentation with Superpoint Graphs
Loic Landrieu, Martin Simonovsky
cs.CVcs.LGcs.NEarXiv:1711.09869v22017Diffusion-LM Improves Controllable Text Generation
Xiang Lisa Li, John Thickstun, Ishaan Gulrajani +2
cs.CLcs.AIcs.LGarXiv:2205.14217v12022Do Vision Transformers See Like Convolutional Neural Networks?
Maithra Raghu, Thomas Unterthiner, Simon Kornblith +2
cs.CVcs.AIcs.LGarXiv:2108.08810v22021StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery
Or Patashnik, Zongze Wu, Eli Shechtman +2
cs.CVcs.CLcs.GRarXiv:2103.17249v12021LoREnc: Low-Rank Encryption for Securing Foundation Models and LoRA Adapters
Beomjin Ahn, Jungmin Kwon, Chanyong Jung +1
cs.CRcs.CVcs.LGarXiv:2605.13163v12026Training Large Language Models to Predict Clinical Events
Benjamin Turtel, Paul Wilczewski, Kris Skotheim
cs.LGcs.AIcs.CLarXiv:2605.12817v12026