Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

17,941 to 18,000 of 20,188

  1. SimPO: Simple Preference Optimization with a Reference-Free Reward

    Yu Meng, Mengzhou Xia, Danqi Chen

    cs.CLcs.LGarXiv:2405.14734v32024
  2. Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models

    Yue Zhang, Yafu Li, Leyang Cui +13

    cs.CLcs.AIcs.CYarXiv:2309.01219v32023
  3. It's Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners

    Timo Schick, Hinrich Schütze

    cs.CLcs.AIcs.LGarXiv:2009.07118v22020
  4. Low-rank Matrix Completion using Alternating Minimization

    Prateek Jain, Praneeth Netrapalli, Sujay Sanghavi

    stat.MLcs.LGmath.OCarXiv:1212.0467v12012
  5. Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting

    Benjamin Wilson, William Qi, Tanmay Agarwal +10

    cs.CVcs.AIcs.LGarXiv:2301.00493v12023
  6. Contrastive Learning of Medical Visual Representations from Paired Images and Text

    Yuhao Zhang, Hang Jiang, Yasuhide Miura +2

    cs.CVcs.CLcs.LGarXiv:2010.00747v22020
  7. Joint Deep Modeling of Users and Items Using Reviews for Recommendation

    Lei Zheng, Vahid Noroozi, Philip S. Yu

    cs.LGcs.IRarXiv:1701.04783v12017
  8. QANet: Combining Local Convolution with Global Self-Attention for Reading Comprehension

    Adams Wei Yu, David Dohan, Minh-Thang Luong +4

    cs.CLcs.AIcs.LGarXiv:1804.09541v12018
  9. BERTweet: A pre-trained language model for English Tweets

    Dat Quoc Nguyen, Thanh Vu, Anh Tuan Nguyen

    cs.CLcs.LGarXiv:2005.10200v22020
  10. StateSMix: Online Lossless Compression via Mamba State Space Models and Sparse N-gram Context Mixing

    Roberto Tacconelli

    cs.LGcs.ITarXiv:2605.02904v12026
  11. A Generalist Agent

    Scott Reed, Konrad Zolna, Emilio Parisotto +17

    cs.AIcs.CLcs.LGarXiv:2205.06175v32022
  12. Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation

    Linfeng Zhang, Jiebo Song, Anni Gao +3

    cs.LGstat.MLarXiv:1905.08094v12019
  13. Multi-scale Attributed Node Embedding

    Benedek Rozemberczki, Carl Allen, Rik Sarkar

    cs.LGcs.NIcs.SIarXiv:1909.13021v32019
  14. When Retrieval Fails Before It Begins: Structurally Indirect Prerequisite Eviction as a Retention Failure in Agentic Memory

    Minkyu Song

    cs.AIcs.CLcs.LGarXiv:2608.20400v12026
  15. World models of environment, agent and joint agent-environment systems

    Manuel Baltieri, Filippo Torresan, Yivan Zhang +2

    cs.AIcs.LGarXiv:2608.20401v12026
  16. A Hybrid Approach to Privacy-Preserving Federated Learning

    Stacey Truex, Nathalie Baracaldo, Ali Anwar +4

    cs.LGstat.MLarXiv:1812.03224v22018
  17. Inf-Net: Automatic COVID-19 Lung Infection Segmentation from CT Images

    Deng-Ping Fan, Tao Zhou, Ge-Peng Ji +5

    eess.IVcs.CVcs.LGarXiv:2004.14133v42020
  18. A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications

    Jie Gui, Zhenan Sun, Yonggang Wen +2

    cs.LGstat.MLarXiv:2001.06937v12020
  19. Bayesian Active Learning for Classification and Preference Learning

    Neil Houlsby, Ferenc Huszár, Zoubin Ghahramani +1

    stat.MLcs.LGarXiv:1112.5745v12011
  20. A Hierarchical Latent Variable Encoder-Decoder Model for Generating Dialogues

    Iulian Vlad Serban, Alessandro Sordoni, Ryan Lowe +4

    cs.CLcs.AIcs.LGarXiv:1605.06069v32016
  21. Deep Forest

    Zhi-Hua Zhou, Ji Feng

    cs.LGstat.MLarXiv:1702.08835v42017
  22. Agnostic Federated Learning

    Mehryar Mohri, Gary Sivek, Ananda Theertha Suresh

    cs.LGstat.MLarXiv:1902.00146v12019
  23. MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs

    Alistair E. W. Johnson, Tom J. Pollard, Nathaniel R. Greenbaum +7

    cs.CVcs.LGeess.IVarXiv:1901.07042v52019
  24. Florence: A New Foundation Model for Computer Vision

    Lu Yuan, Dongdong Chen, Yi-Ling Chen +20

    cs.CVcs.AIcs.LGarXiv:2111.11432v12021
  25. Evading Defenses to Transferable Adversarial Examples by Translation-Invariant Attacks

    Yinpeng Dong, Tianyu Pang, Hang Su +1

    cs.CVcs.CRcs.LGarXiv:1904.02884v12019
  26. Tune: A Research Platform for Distributed Model Selection and Training

    Richard Liaw, Eric Liang, Robert Nishihara +3

    cs.LGcs.DCstat.MLarXiv:1807.05118v12018
  27. Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture

    Mahmoud Assran, Quentin Duval, Ishan Misra +5

    cs.CVcs.AIcs.LGarXiv:2301.08243v32023
  28. An Explanation of In-context Learning as Implicit Bayesian Inference

    Sang Michael Xie, Aditi Raghunathan, Percy Liang +1

    cs.CLcs.LGarXiv:2111.02080v62021
  29. TUDataset: A collection of benchmark datasets for learning with graphs

    Christopher Morris, Nils M. Kriege, Franka Bause +3

    cs.LGcs.NEstat.MLarXiv:2007.08663v12020
  30. Recipe for a General, Powerful, Scalable Graph Transformer

    Ladislav Rampášek, Mikhail Galkin, Vijay Prakash Dwivedi +3

    cs.LGarXiv:2205.12454v42022
  31. Visualizing and Understanding Recurrent Networks

    Andrej Karpathy, Justin Johnson, Li Fei-Fei

    cs.LGcs.CLcs.NEarXiv:1506.02078v22015
  32. VectorNet: Encoding HD Maps and Agent Dynamics from Vectorized Representation

    Jiyang Gao, Chen Sun, Hang Zhao +4

    cs.CVcs.LGstat.MLarXiv:2005.04259v12020
  33. Towards Understanding the Robustness of Sparse Autoencoders

    Ahson Saiyed, Sabrina Sadiekh, Chirag Agarwal

    cs.LGcs.AIcs.CLarXiv:2604.18756v12026
  34. DeblurGAN-v2: Deblurring (Orders-of-Magnitude) Faster and Better

    Orest Kupyn, Tetiana Martyniuk, Junru Wu +1

    cs.CVcs.LGarXiv:1908.03826v12019
  35. Directional Message Passing for Molecular Graphs

    Johannes Gasteiger, Janek Groß, Stephan Günnemann

    cs.LGphysics.comp-phstat.MLarXiv:2003.03123v22020
  36. Don't Decay the Learning Rate, Increase the Batch Size

    Samuel L. Smith, Pieter-Jan Kindermans, Chris Ying +1

    cs.LGcs.CVcs.DCarXiv:1711.00489v22017
  37. GCC: Graph Contrastive Coding for Graph Neural Network Pre-Training

    Jiezhong Qiu, Qibin Chen, Yuxiao Dong +5

    cs.LGcs.SIstat.MLarXiv:2006.09963v32020
  38. Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy

    Jiehui Xu, Haixu Wu, Jianmin Wang +1

    cs.LGarXiv:2110.02642v52021
  39. Dynamic Network Surgery for Efficient DNNs

    Yiwen Guo, Anbang Yao, Yurong Chen

    cs.NEcs.CVcs.LGarXiv:1608.04493v22016
  40. A Generalization of Transformer Networks to Graphs

    Vijay Prakash Dwivedi, Xavier Bresson

    cs.LGarXiv:2012.09699v22020
  41. Meta Networks

    Tsendsuren Munkhdalai, Hong Yu

    cs.LGstat.MLarXiv:1703.00837v22017
  42. MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis

    Kundan Kumar, Rithesh Kumar, Thibault de Boissiere +6

    eess.AScs.CLcs.LGarXiv:1910.06711v32019
  43. Scaling Vision with Sparse Mixture of Experts

    Carlos Riquelme, Joan Puigcerver, Basil Mustafa +5

    cs.CVcs.LGstat.MLarXiv:2106.05974v12021
  44. MixHop: Higher-Order Graph Convolutional Architectures via Sparsified Neighborhood Mixing

    Sami Abu-El-Haija, Bryan Perozzi, Amol Kapoor +5

    cs.LGcs.SIstat.MLarXiv:1905.00067v32019
  45. DETR3D: 3D Object Detection from Multi-view Images via 3D-to-2D Queries

    Yue Wang, Vitor Guizilini, Tianyuan Zhang +3

    cs.CVcs.AIcs.LGarXiv:2110.06922v12021
  46. Towards Robust Interpretability with Self-Explaining Neural Networks

    David Alvarez-Melis, Tommi S. Jaakkola

    cs.LGstat.MLarXiv:1806.07538v22018
  47. Temporal Graph Networks for Deep Learning on Dynamic Graphs

    Emanuele Rossi, Ben Chamberlain, Fabrizio Frasca +3

    cs.LGstat.MLarXiv:2006.10637v32020
  48. Robots that can adapt like animals

    Antoine Cully, Jeff Clune, Danesh Tarapore +1

    cs.ROcs.AIcs.LGarXiv:1407.3501v42014
  49. CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

    Wenyi Hong, Ming Ding, Wendi Zheng +2

    cs.CVcs.CLcs.LGarXiv:2205.15868v12022
  50. Editing Models with Task Arithmetic

    Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman +4

    cs.LGcs.CLcs.CVarXiv:2212.04089v32022
  51. WaveGlow: A Flow-based Generative Network for Speech Synthesis

    Ryan Prenger, Rafael Valle, Bryan Catanzaro

    cs.SDcs.AIcs.LGarXiv:1811.00002v12018
  52. Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling

    Xumin Yu, Lulu Tang, Yongming Rao +3

    cs.CVcs.AIcs.LGarXiv:2111.14819v22021
  53. Regularizing and Optimizing LSTM Language Models

    Stephen Merity, Nitish Shirish Keskar, Richard Socher

    cs.CLcs.LGcs.NEarXiv:1708.02182v12017
  54. Adversarial Training Methods for Semi-Supervised Text Classification

    Takeru Miyato, Andrew M. Dai, Ian Goodfellow

    stat.MLcs.LGarXiv:1605.07725v42016
  55. Depth-supervised NeRF: Fewer Views and Faster Training for Free

    Kangle Deng, Andrew Liu, Jun-Yan Zhu +1

    cs.CVcs.GRcs.LGarXiv:2107.02791v32021
  56. Semi-supervised Knowledge Transfer for Deep Learning from Private Training Data

    Nicolas Papernot, Martín Abadi, Úlfar Erlingsson +2

    stat.MLcs.CRcs.LGarXiv:1610.05755v42016
  57. CoroNet: A deep neural network for detection and diagnosis of COVID-19 from chest x-ray images

    Asif Iqbal Khan, Junaid Latief Shah, Mudasir Bhat

    eess.IVcs.LGstat.MLarXiv:2004.04931v32020
  58. Mass-Editing Memory in a Transformer

    Kevin Meng, Arnab Sen Sharma, Alex Andonian +2

    cs.CLcs.LGarXiv:2210.07229v22022
  59. Learning the solution operator of parametric partial differential equations with physics-informed DeepOnets

    Sifan Wang, Hanwen Wang, Paris Perdikaris

    cs.LGmath.NAstat.MLarXiv:2103.10974v12021
  60. Knowledge Graph Convolutional Networks for Recommender Systems

    Hongwei Wang, Miao Zhao, Xing Xie +2

    cs.IRcs.LGstat.MLarXiv:1904.12575v12019