Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
17,941 to 18,000 of 20,188
SimPO: Simple Preference Optimization with a Reference-Free Reward
Yu Meng, Mengzhou Xia, Danqi Chen
cs.CLcs.LGarXiv:2405.14734v32024Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
Yue Zhang, Yafu Li, Leyang Cui +13
cs.CLcs.AIcs.CYarXiv:2309.01219v32023It's Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners
Timo Schick, Hinrich Schütze
cs.CLcs.AIcs.LGarXiv:2009.07118v22020Low-rank Matrix Completion using Alternating Minimization
Prateek Jain, Praneeth Netrapalli, Sujay Sanghavi
stat.MLcs.LGmath.OCarXiv:1212.0467v12012Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting
Benjamin Wilson, William Qi, Tanmay Agarwal +10
cs.CVcs.AIcs.LGarXiv:2301.00493v12023Contrastive Learning of Medical Visual Representations from Paired Images and Text
Yuhao Zhang, Hang Jiang, Yasuhide Miura +2
cs.CVcs.CLcs.LGarXiv:2010.00747v22020Joint Deep Modeling of Users and Items Using Reviews for Recommendation
Lei Zheng, Vahid Noroozi, Philip S. Yu
cs.LGcs.IRarXiv:1701.04783v12017QANet: Combining Local Convolution with Global Self-Attention for Reading Comprehension
Adams Wei Yu, David Dohan, Minh-Thang Luong +4
cs.CLcs.AIcs.LGarXiv:1804.09541v12018BERTweet: A pre-trained language model for English Tweets
Dat Quoc Nguyen, Thanh Vu, Anh Tuan Nguyen
cs.CLcs.LGarXiv:2005.10200v22020StateSMix: Online Lossless Compression via Mamba State Space Models and Sparse N-gram Context Mixing
Roberto Tacconelli
cs.LGcs.ITarXiv:2605.02904v12026A Generalist Agent
Scott Reed, Konrad Zolna, Emilio Parisotto +17
cs.AIcs.CLcs.LGarXiv:2205.06175v32022Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation
Linfeng Zhang, Jiebo Song, Anni Gao +3
cs.LGstat.MLarXiv:1905.08094v12019Multi-scale Attributed Node Embedding
Benedek Rozemberczki, Carl Allen, Rik Sarkar
cs.LGcs.NIcs.SIarXiv:1909.13021v32019When Retrieval Fails Before It Begins: Structurally Indirect Prerequisite Eviction as a Retention Failure in Agentic Memory
Minkyu Song
cs.AIcs.CLcs.LGarXiv:2608.20400v12026World models of environment, agent and joint agent-environment systems
Manuel Baltieri, Filippo Torresan, Yivan Zhang +2
cs.AIcs.LGarXiv:2608.20401v12026A Hybrid Approach to Privacy-Preserving Federated Learning
Stacey Truex, Nathalie Baracaldo, Ali Anwar +4
cs.LGstat.MLarXiv:1812.03224v22018Inf-Net: Automatic COVID-19 Lung Infection Segmentation from CT Images
Deng-Ping Fan, Tao Zhou, Ge-Peng Ji +5
eess.IVcs.CVcs.LGarXiv:2004.14133v42020A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications
Jie Gui, Zhenan Sun, Yonggang Wen +2
cs.LGstat.MLarXiv:2001.06937v12020Bayesian Active Learning for Classification and Preference Learning
Neil Houlsby, Ferenc Huszár, Zoubin Ghahramani +1
stat.MLcs.LGarXiv:1112.5745v12011A Hierarchical Latent Variable Encoder-Decoder Model for Generating Dialogues
Iulian Vlad Serban, Alessandro Sordoni, Ryan Lowe +4
cs.CLcs.AIcs.LGarXiv:1605.06069v32016Deep Forest
Zhi-Hua Zhou, Ji Feng
cs.LGstat.MLarXiv:1702.08835v42017Agnostic Federated Learning
Mehryar Mohri, Gary Sivek, Ananda Theertha Suresh
cs.LGstat.MLarXiv:1902.00146v12019MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs
Alistair E. W. Johnson, Tom J. Pollard, Nathaniel R. Greenbaum +7
cs.CVcs.LGeess.IVarXiv:1901.07042v52019Florence: A New Foundation Model for Computer Vision
Lu Yuan, Dongdong Chen, Yi-Ling Chen +20
cs.CVcs.AIcs.LGarXiv:2111.11432v12021Evading Defenses to Transferable Adversarial Examples by Translation-Invariant Attacks
Yinpeng Dong, Tianyu Pang, Hang Su +1
cs.CVcs.CRcs.LGarXiv:1904.02884v12019Tune: A Research Platform for Distributed Model Selection and Training
Richard Liaw, Eric Liang, Robert Nishihara +3
cs.LGcs.DCstat.MLarXiv:1807.05118v12018Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
Mahmoud Assran, Quentin Duval, Ishan Misra +5
cs.CVcs.AIcs.LGarXiv:2301.08243v32023An Explanation of In-context Learning as Implicit Bayesian Inference
Sang Michael Xie, Aditi Raghunathan, Percy Liang +1
cs.CLcs.LGarXiv:2111.02080v62021TUDataset: A collection of benchmark datasets for learning with graphs
Christopher Morris, Nils M. Kriege, Franka Bause +3
cs.LGcs.NEstat.MLarXiv:2007.08663v12020Recipe for a General, Powerful, Scalable Graph Transformer
Ladislav Rampášek, Mikhail Galkin, Vijay Prakash Dwivedi +3
cs.LGarXiv:2205.12454v42022Visualizing and Understanding Recurrent Networks
Andrej Karpathy, Justin Johnson, Li Fei-Fei
cs.LGcs.CLcs.NEarXiv:1506.02078v22015VectorNet: Encoding HD Maps and Agent Dynamics from Vectorized Representation
Jiyang Gao, Chen Sun, Hang Zhao +4
cs.CVcs.LGstat.MLarXiv:2005.04259v12020Towards Understanding the Robustness of Sparse Autoencoders
Ahson Saiyed, Sabrina Sadiekh, Chirag Agarwal
cs.LGcs.AIcs.CLarXiv:2604.18756v12026DeblurGAN-v2: Deblurring (Orders-of-Magnitude) Faster and Better
Orest Kupyn, Tetiana Martyniuk, Junru Wu +1
cs.CVcs.LGarXiv:1908.03826v12019Directional Message Passing for Molecular Graphs
Johannes Gasteiger, Janek Groß, Stephan Günnemann
cs.LGphysics.comp-phstat.MLarXiv:2003.03123v22020Don't Decay the Learning Rate, Increase the Batch Size
Samuel L. Smith, Pieter-Jan Kindermans, Chris Ying +1
cs.LGcs.CVcs.DCarXiv:1711.00489v22017GCC: Graph Contrastive Coding for Graph Neural Network Pre-Training
Jiezhong Qiu, Qibin Chen, Yuxiao Dong +5
cs.LGcs.SIstat.MLarXiv:2006.09963v32020Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy
Jiehui Xu, Haixu Wu, Jianmin Wang +1
cs.LGarXiv:2110.02642v52021Dynamic Network Surgery for Efficient DNNs
Yiwen Guo, Anbang Yao, Yurong Chen
cs.NEcs.CVcs.LGarXiv:1608.04493v22016A Generalization of Transformer Networks to Graphs
Vijay Prakash Dwivedi, Xavier Bresson
cs.LGarXiv:2012.09699v22020Meta Networks
Tsendsuren Munkhdalai, Hong Yu
cs.LGstat.MLarXiv:1703.00837v22017MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere +6
eess.AScs.CLcs.LGarXiv:1910.06711v32019Scaling Vision with Sparse Mixture of Experts
Carlos Riquelme, Joan Puigcerver, Basil Mustafa +5
cs.CVcs.LGstat.MLarXiv:2106.05974v12021MixHop: Higher-Order Graph Convolutional Architectures via Sparsified Neighborhood Mixing
Sami Abu-El-Haija, Bryan Perozzi, Amol Kapoor +5
cs.LGcs.SIstat.MLarXiv:1905.00067v32019DETR3D: 3D Object Detection from Multi-view Images via 3D-to-2D Queries
Yue Wang, Vitor Guizilini, Tianyuan Zhang +3
cs.CVcs.AIcs.LGarXiv:2110.06922v12021Towards Robust Interpretability with Self-Explaining Neural Networks
David Alvarez-Melis, Tommi S. Jaakkola
cs.LGstat.MLarXiv:1806.07538v22018Temporal Graph Networks for Deep Learning on Dynamic Graphs
Emanuele Rossi, Ben Chamberlain, Fabrizio Frasca +3
cs.LGstat.MLarXiv:2006.10637v32020Robots that can adapt like animals
Antoine Cully, Jeff Clune, Danesh Tarapore +1
cs.ROcs.AIcs.LGarXiv:1407.3501v42014CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
Wenyi Hong, Ming Ding, Wendi Zheng +2
cs.CVcs.CLcs.LGarXiv:2205.15868v12022Editing Models with Task Arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman +4
cs.LGcs.CLcs.CVarXiv:2212.04089v32022WaveGlow: A Flow-based Generative Network for Speech Synthesis
Ryan Prenger, Rafael Valle, Bryan Catanzaro
cs.SDcs.AIcs.LGarXiv:1811.00002v12018Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling
Xumin Yu, Lulu Tang, Yongming Rao +3
cs.CVcs.AIcs.LGarXiv:2111.14819v22021Regularizing and Optimizing LSTM Language Models
Stephen Merity, Nitish Shirish Keskar, Richard Socher
cs.CLcs.LGcs.NEarXiv:1708.02182v12017Adversarial Training Methods for Semi-Supervised Text Classification
Takeru Miyato, Andrew M. Dai, Ian Goodfellow
stat.MLcs.LGarXiv:1605.07725v42016Depth-supervised NeRF: Fewer Views and Faster Training for Free
Kangle Deng, Andrew Liu, Jun-Yan Zhu +1
cs.CVcs.GRcs.LGarXiv:2107.02791v32021Semi-supervised Knowledge Transfer for Deep Learning from Private Training Data
Nicolas Papernot, Martín Abadi, Úlfar Erlingsson +2
stat.MLcs.CRcs.LGarXiv:1610.05755v42016CoroNet: A deep neural network for detection and diagnosis of COVID-19 from chest x-ray images
Asif Iqbal Khan, Junaid Latief Shah, Mudasir Bhat
eess.IVcs.LGstat.MLarXiv:2004.04931v32020Mass-Editing Memory in a Transformer
Kevin Meng, Arnab Sen Sharma, Alex Andonian +2
cs.CLcs.LGarXiv:2210.07229v22022Learning the solution operator of parametric partial differential equations with physics-informed DeepOnets
Sifan Wang, Hanwen Wang, Paris Perdikaris
cs.LGmath.NAstat.MLarXiv:2103.10974v12021Knowledge Graph Convolutional Networks for Recommender Systems
Hongwei Wang, Miao Zhao, Xing Xie +2
cs.IRcs.LGstat.MLarXiv:1904.12575v12019