Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
19,321 to 19,380 of 20,199
Accurate Image Super-Resolution Using Very Deep Convolutional Networks
Jiwon Kim, Jung Kwon Lee, Kyoung Mu Lee
cs.CVcs.LGarXiv:1511.04587v22015An overview of gradient descent optimization algorithms
Sebastian Ruder
cs.LGarXiv:1609.04747v22016Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson
Tanay Chowdhury, Saeideh Shahrokh Esfahani
cs.AIcs.LGarXiv:2608.18531v12026Analyzing and Improving the Image Quality of StyleGAN
Tero Karras, Samuli Laine, Miika Aittala +3
cs.CVcs.LGcs.NEarXiv:1912.04958v22019Gaussian Error Linear Units (GELUs)
Dan Hendrycks, Kevin Gimpel
cs.LGarXiv:1606.08415v52016Character-level Convolutional Networks for Text Classification
Xiang Zhang, Junbo Zhao, Yann LeCun
cs.LGcs.CLarXiv:1509.01626v32015Addressing Function Approximation Error in Actor-Critic Methods
Scott Fujimoto, Herke van Hoof, David Meger
cs.AIcs.LGstat.MLarXiv:1802.09477v32018EfficientDet: Scalable and Efficient Object Detection
Mingxing Tan, Ruoming Pang, Quoc V. Le
cs.CVcs.LGeess.IVarXiv:1911.09070v72019Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Yonghui Wu, Mike Schuster, Zhifeng Chen +28
cs.CLcs.AIcs.LGarXiv:1609.08144v22016SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition
Dae Lee, Mihai Delgeanu, Adel Youssef
cs.AIcs.LGarXiv:2608.18303v12026Neural Ordinary Differential Equations
Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt +1
cs.LGcs.AIstat.MLarXiv:1806.07366v52018Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application
Elias Schubert, Felix Bießmann
cs.AIcs.CRcs.IRarXiv:2608.18289v12026Cacheable by Design? Training Mixture-of-Experts Routers for Locality Against the Edge Memory-Bandwidth Wall: A Pre-Registered Negative Result with a Systems Measurement Study
Shriniwas Ramesh Suram
cs.AIcs.LGarXiv:2608.18261v12026Neural Discrete Representation Learning
Aaron van den Oord, Oriol Vinyals, Koray Kavukcuoglu
cs.LGarXiv:1711.00937v22017Learning both Weights and Connections for Efficient Neural Networks
Song Han, Jeff Pool, John Tran +1
cs.NEcs.CVcs.LGarXiv:1506.02626v32015WaveNet: A Generative Model for Raw Audio
Aaron van den Oord, Sander Dieleman, Heiga Zen +6
cs.SDcs.LGarXiv:1609.03499v22016Axiomatic Attribution for Deep Networks
Mukund Sundararajan, Ankur Taly, Qiqi Yan
cs.LGarXiv:1703.01365v22017Advances and Open Problems in Federated Learning
Peter Kairouz, H. Brendan McMahan, Brendan Avent +56
cs.LGcs.CRstat.MLarXiv:1912.04977v32019SW-ProxyCE: Zero-Query Adversarial Transfer from Public EEG Encoders to Private Downstream Models
Linhua Cong, Dingkun Liu, Dongrui Wu
cs.LGarXiv:2608.16931v12026Mr.Dec: Daily-Scale Longitudinal Multimodal Modeling for 30-Day Readmission Prediction
Minjun Kim, Jong Hak Moon
cs.LGarXiv:2608.16929v12026FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud
Dheeraj Mohandas Pai, Lu Xian
cs.AIcs.LGarXiv:2608.18136v12026Trust Region Policy Optimization
John Schulman, Sergey Levine, Philipp Moritz +2
cs.LGarXiv:1502.05477v52015Matching Networks for One Shot Learning
Oriol Vinyals, Charles Blundell, Timothy Lillicrap +2
cs.LGstat.MLarXiv:1606.04080v22016Progressive Growing of GANs for Improved Quality, Stability, and Variation
Tero Karras, Timo Aila, Samuli Laine +1
cs.NEcs.LGstat.MLarXiv:1710.10196v32017Convolutional Neural Networks on Graphs with Fast Localized Spectral Filtering
Michaël Defferrard, Xavier Bresson, Pierre Vandergheynst
cs.LGstat.MLarXiv:1606.09375v32016Position: Current Model Cards Are Insufficient for Downstream Governance of Open-Weight Foundation Models
Sungwon Chae, Keonwoo Kim, Hoki Kim +2
cs.AIcs.LGarXiv:2608.18086v12026Wide Residual Networks
Sergey Zagoruyko, Nikos Komodakis
cs.CVcs.LGcs.NEarXiv:1605.07146v42016Benchmarking Classical and Transformer-Based Models for Document Sensitivity Classification
Aleesha Zainab, Muhammad Ahmed Khalid, Faheem Ullah Khan +1
cs.LGarXiv:2608.16928v12026SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers
Enze Xie, Wenhai Wang, Zhiding Yu +3
cs.CVcs.LGarXiv:2105.15203v32021Measuring Massive Multitask Language Understanding
Dan Hendrycks, Collin Burns, Steven Basart +4
cs.CYcs.AIcs.CLarXiv:2009.03300v32020Explainable Artificial Intelligence (XAI): Concepts, Taxonomies, Opportunities and Challenges toward Responsible AI
Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser +9
cs.AIcs.LGcs.NEarXiv:1910.10045v22019Fine-tuning Multi-modal LLMs with ART: Art-based Reinforcement Training
Michal Chudoba, Sergey Alyaev, Petra Galuscakova +1
cs.LGcs.AIcs.CLarXiv:2606.11854v12026Dynamic Regime-Aware Conformal Calibration for Reliable Economic Forecast Intervals under Multiple Distribution Shifts
Bogdan Oancea
cs.LGarXiv:2608.17079v12026Network Denoising Revisited: A Ricci-Flow-Inspired Graph Diffusion Method
Ye Fang, Chuan-Xian Ren
cs.SIcs.LGarXiv:2608.16923v12026XLNet: Generalized Autoregressive Pretraining for Language Understanding
Zhilin Yang, Zihang Dai, Yiming Yang +3
cs.CLcs.LGarXiv:1906.08237v22019Causal Discovery in Equal Variance Linear Gaussian DAGs via SURE-Tuned Ridge Regression
Sambit Mishra, Urbashi Mitra
cs.LGeess.SPstat.MLarXiv:2608.17132v12026Diagonal Multi-omics Integration of Heterogenous Datasets
Maksim V. Kukushkin, Mikhail S. Arbatskiy, Dmitriy E. Balandin +1
stat.MLcs.LGmath.FAarXiv:2608.16968v12026DOW-KE: Anchor-Free Multi-Layer Knowledge Editing via Direct End-to-End Weight Optimization
Ran Chen, Junbo Zhang, Qianli Zhou +2
cs.LGarXiv:2608.16932v12026When is Your LLM Steerable?
Chenrui Fan, Yize Cheng, Ming Li +2
cs.CLcs.LGarXiv:2606.11599v12026APEX: A Network-Native Time-Series Foundation Model for Forecasting and Anomaly Detection for Wireless Edge Operations
Swadhin Pradhan, Niloo Bahadori, Peiman Amini
cs.LGarXiv:2606.11553v12026Data-DPO: Direct Preference Optimization for Target Model Data Selection in LLM Post-Training
Peng Sun, Yi Yang, Antong Zhang +7
cs.LGarXiv:2608.16926v12026Neural Message Passing for Quantum Chemistry
Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley +2
cs.LGarXiv:1704.01212v22017PianoKontext: Expressive Performance Rendering from Deadpan Context
Dmitrii Gavrilev
cs.SDcs.LGarXiv:2606.12282v12026MITRE-SAGE: A Multi-Agent Cybersecurity Question-Answering Model
Ali Habibzadeh, Farid Feyzi, Reza Ebrahimi Atani
cs.IRcs.CRcs.LGarXiv:2608.16921v22026Deep Reinforcement Learning with Double Q-learning
Hado van Hasselt, Arthur Guez, David Silver
cs.LGarXiv:1509.06461v32015Hierarchical Data Selection via Manifold Coverage and Sparse Feature Coverage in LLM Post-training
Peng Sun, Yi Yang, Antong Zhang +7
cs.LGcs.CVarXiv:2608.16927v12026How transferable are features in deep neural networks?
Jason Yosinski, Jeff Clune, Yoshua Bengio +1
cs.LGcs.NEarXiv:1411.1792v12014SPSA Hyperparameter Tuning for Variational Quantum Natural Language Inference
Nayan D'Souza, Christopher J. Agostino
quant-phcs.LGarXiv:2608.16939v12026UNet++: A Nested U-Net Architecture for Medical Image Segmentation
Zongwei Zhou, Md Mahfuzur Rahman Siddiquee, Nima Tajbakhsh +1
cs.CVcs.LGeess.IVarXiv:1807.10165v12018Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders
Gleb Gerasimov, Timofei Rusalev, Nikita Balagansky +3
cs.LGcs.AIcs.CLarXiv:2606.12138v12026Information Spreading in Diffusion Models from Effective Field Theory
Navonil Neogi, Nabil Iqbal
hep-thcond-mat.stat-mechcs.LGarXiv:2608.14308v12026MagViT: Interpretable Multi-Magnification Transformers with Patient-Level Model Selection for Breast Histopathology
Nabil Ashab, Soumit Kumar Kundu, Saif Mahmud Parvez +3
eess.IVcs.CVcs.LGarXiv:2608.16959v12026Pessimistic Meta-Induction and Its Limits: Lessons from Frequentist Statistics and Machine Learning Theory
Hanti Lin
cs.LGstat.MEarXiv:2608.17213v12026MultiSigBERT: Beyond Survival Analysis through Multimodal and Sequential Modeling in Oncology
Paul Minchella, Stéphane Chrétien, Guillaume Metzler +2
cs.LGarXiv:2608.16972v12026On Subquadratic Architectures: From Applications to Principles
Anamaria-Roberta Hartl, Levente Zólyomi, David Stap +6
cs.LGarXiv:2606.12364v12026VLCP: Vision Language Control Policy Closed-Loop Code Replanning for Robot Manipulation
Dhia Naouali, Minghan Wu, Claudia Wong +2
cs.ROcs.LGarXiv:2608.16978v12026Agents unlock new capabilities through Switching LoRA Adapters as a Tool (SLAaaT)
Kenneth Ge
cs.LGarXiv:2608.17034v12026Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
Joao Carreira, Andrew Zisserman
cs.CVcs.LGarXiv:1705.07750v32017Distributed Representations of Sentences and Documents
Quoc V. Le, Tomas Mikolov
cs.CLcs.AIcs.LGarXiv:1405.4053v22014Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents
Yujun Zhou, Kehan Guo, Haomin Zhuang +8
cs.LGcs.CLarXiv:2606.13174v12026