Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
19,141 to 19,200 of 20,205
BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer
Fei Sun, Jun Liu, Jian Wu +4
cs.IRcs.LGarXiv:1904.06690v22019Summaries:한국어Adversarial Machine Learning at Scale
Alexey Kurakin, Ian Goodfellow, Samy Bengio
cs.CVcs.CRcs.LGarXiv:1611.01236v22016A Contextual-Bandit Approach to Personalized News Article Recommendation
Lihong Li, Wei Chu, John Langford +1
cs.LGcs.AIcs.IRarXiv:1003.0146v22010Local Gains and Fixed-Assignment Set Losses in Shared Set Decoders
Ze Zhang, Yang Zhang
cs.CVcs.LGarXiv:2608.14717v12026MnasNet: Platform-Aware Neural Architecture Search for Mobile
Mingxing Tan, Bo Chen, Ruoming Pang +4
cs.CVcs.LGarXiv:1807.11626v32018FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series Forecasting
Tian Zhou, Ziqing Ma, Qingsong Wen +3
cs.LGstat.MLarXiv:2201.12740v32022CFR without Unbiasedness: Deterministic Guarantees for Persistent Public-Chance Schedules
Jiaxing Guo, Lei Ye
cs.GTcs.LGarXiv:2608.14761v12026The Note-Chord-Voice Framework: Structured Source Separation and Causal Inference for EV Charging Data
Jiajie Chen, Jinfeng Li
eess.SPcs.LGcs.SDarXiv:2608.14756v12026Can Neural Networks Learn by Experimenting on Themselves? Self-Interventional Learning from Functional Consequences to Predictive Self-Knowledge
Michał Tomaszewski
cs.LGarXiv:2608.14894v12026Is Grokking a Loss of Normal Hyperbolicity of the Interpolation Manifold?
Suvinava Basak
cs.LGarXiv:2608.14803v12026Universal and Transferable Adversarial Attacks on Aligned Language Models
Andy Zou, Zifan Wang, Nicholas Carlini +3
cs.CLcs.AIcs.CRarXiv:2307.15043v22023ER-KANs: Efficient and Robust Kolmogorov-Arnold Networks for Data-Scarce Scientific Machine Learning
Harshil Lodhiya
cs.LGcs.AIarXiv:2608.14773v12026A Constant-Competitive Algorithm for Dynamic Mixture-of-Experts Serving
Ian D'Ambrosio
cs.DScs.LGarXiv:2608.16947v12026Every Expert Counts: ExactMoE for Memory-Efficient W4A16 Inference
Amjad Saab
cs.LGarXiv:2608.15383v12026Deeper Insights into Graph Convolutional Networks for Semi-Supervised Learning
Qimai Li, Zhichao Han, Xiao-Ming Wu
cs.LGstat.MLarXiv:1801.07606v12018Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
Xingchao Liu, Chengyue Gong, Qiang Liu
cs.LGarXiv:2209.03003v12022Convolutional Networks on Graphs for Learning Molecular Fingerprints
David Duvenaud, Dougal Maclaurin, Jorge Aguilera-Iparraguirre +4
cs.LGcs.NEstat.MLarXiv:1509.09292v22015CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison
Jeremy Irvin, Pranav Rajpurkar, Michael Ko +17
cs.CVcs.AIcs.LGarXiv:1901.07031v12019MLP-Mixer: An all-MLP Architecture for Vision
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov +9
cs.CVcs.AIcs.LGarXiv:2105.01601v42021Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples
Anish Athalye, Nicholas Carlini, David Wagner
cs.LGcs.AIcs.CRarXiv:1802.00420v42018Do Geometry-Aware Positional Encodings Help Transformers in Spatial Imperfect-Information Games?
Wenji Fu
cs.LGcs.AIstat.MLarXiv:2608.14982v12026Automatic chemical design using a data-driven continuous representation of molecules
Rafael Gómez-Bombarelli, Jennifer N. Wei, David Duvenaud +7
cs.LGphysics.chem-pharXiv:1610.02415v32016MixMatch: A Holistic Approach to Semi-Supervised Learning
David Berthelot, Nicholas Carlini, Ian Goodfellow +3
cs.LGcs.AIcs.CVarXiv:1905.02249v22019Unsupervised Representation Learning by Predicting Image Rotations
Spyros Gidaris, Praveer Singh, Nikos Komodakis
cs.CVcs.LGarXiv:1803.07728v12018Training Compute-Optimal Large Language Models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch +19
cs.CLcs.LGarXiv:2203.15556v12022Gradient Episodic Memory for Continual Learning
David Lopez-Paz, Marc'Aurelio Ranzato
cs.LGcs.AIarXiv:1706.08840v62017The Price of Thinking: Reasoning Effort as a Model-Specific API Contract
Yeabin Moon
cs.AIcs.CLcs.CYarXiv:2608.16956v12026Continual Lifelong Learning with Neural Networks: A Review
German I. Parisi, Ronald Kemker, Jose L. Part +2
cs.LGq-bio.NCstat.MLarXiv:1802.07569v42018Relational inductive biases, deep learning, and graph networks
Peter W. Battaglia, Jessica B. Hamrick, Victor Bapst +24
cs.LGcs.AIstat.MLarXiv:1806.01261v32018Dynamic Entanglement-Weighted Pruning for Quantum Federated Unlearning in Supply-Chain Risk Prediction
Aditya Kumar, Sumit Chongder
quant-phcs.LGarXiv:2608.17069v12026Reinforcement Learning as (Discrete) Potential Theory
Christopher Connolly
cs.LGcs.GTarXiv:2608.17181v12026Complex Embeddings for Simple Link Prediction
Théo Trouillon, Johannes Welbl, Sebastian Riedel +2
cs.AIcs.LGstat.MLarXiv:1606.06357v12016rl-triton: High-Performance Triton GPU Kernels for Reinforcement Learning Credit Assignment
Lars Simon Zehnder
cs.LGcs.DCcs.PFarXiv:2608.17641v22026Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
Philipp Krähenbühl, Vladlen Koltun
cs.CVcs.AIcs.LGarXiv:1210.5644v12012Transformers in Vision: A Survey
Salman Khan, Muzammal Naseer, Munawar Hayat +3
cs.CVcs.AIcs.LGarXiv:2101.01169v52021Understanding Black-box Predictions via Influence Functions
Pang Wei Koh, Percy Liang
stat.MLcs.AIcs.LGarXiv:1703.04730v32017Implicit Neural Representations with Periodic Activation Functions
Vincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman +2
cs.CVcs.LGeess.IVarXiv:2006.09661v12020SAM 2: Segment Anything in Images and Videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu +15
cs.CVcs.AIcs.LGarXiv:2408.00714v22024DreamFusion: Text-to-3D using 2D Diffusion
Ben Poole, Ajay Jain, Jonathan T. Barron +1
cs.CVcs.LGstat.MLarXiv:2209.14988v12022Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains
Matthew Tancik, Pratul P. Srinivasan, Ben Mildenhall +6
cs.CVcs.LGarXiv:2006.10739v12020Mistral 7B
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch +15
cs.CLcs.AIcs.LGarXiv:2310.06825v12023Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Zou +2
cs.CLcs.AIcs.LGarXiv:1607.06520v12016BEiT: BERT Pre-Training of Image Transformers
Hangbo Bao, Li Dong, Songhao Piao +1
cs.CVcs.LGarXiv:2106.08254v22021InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Wenliang Dai, Junnan Li, Dongxu Li +6
cs.CVcs.LGarXiv:2305.06500v22023Think Shallow, Solve Deep: Controlling Recurrent Dynamics for Reliable Test-Time Depth
Ivan Viakhirev, Kirill Borodin, Amirah Almutairi +3
cs.LGcs.CLarXiv:2608.18222v12026End-to-End Training of Deep Visuomotor Policies
Sergey Levine, Chelsea Finn, Trevor Darrell +1
cs.LGcs.CVcs.ROarXiv:1504.00702v52015Self-Attentive Sequential Recommendation
Wang-Cheng Kang, Julian McAuley
cs.IRcs.LGarXiv:1808.09781v12018WhiteMatter: All-to-All Cross-Layer Connections via KV Mixing
Wenbo Zhang, Xiang Ren
cs.CLcs.LGarXiv:2608.18486v12026Deep CORAL: Correlation Alignment for Deep Domain Adaptation
Baochen Sun, Kate Saenko
cs.CVcs.AIcs.LGarXiv:1607.01719v12016Stability-Aware Feature Design for Robust Watermark Detection in Machine-Generated Text
Sina Mansouri, Mohit Marvania, Abolfazl Safikhani
cs.CLcs.LGstat.MLarXiv:2608.18102v12026Let's Verify Step by Step
Hunter Lightman, Vineet Kosaraju, Yura Burda +7
cs.LGcs.AIcs.CLarXiv:2305.20050v12023Self-Attention Generative Adversarial Networks
Han Zhang, Ian Goodfellow, Dimitris Metaxas +1
stat.MLcs.LGarXiv:1805.08318v22018Conformer: Convolution-augmented Transformer for Speech Recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu +8
eess.AScs.LGcs.SDarXiv:2005.08100v12020Neural Graph Collaborative Filtering
Xiang Wang, Xiangnan He, Meng Wang +2
cs.IRcs.LGcs.SIarXiv:1905.08108v22019Unsupervised Feature Learning via Non-Parametric Instance-level Discrimination
Zhirong Wu, Yuanjun Xiong, Stella Yu +1
cs.CVcs.LGarXiv:1805.01978v12018Recurrent Models of Visual Attention
Volodymyr Mnih, Nicolas Heess, Alex Graves +1
cs.LGcs.CVstat.MLarXiv:1406.6247v12014Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
Yoshua Bengio, Nicholas Léonard, Aaron Courville
cs.LGarXiv:1308.3432v12013Practical Black-Box Attacks against Machine Learning
Nicolas Papernot, Patrick McDaniel, Ian Goodfellow +3
cs.CRcs.LGarXiv:1602.02697v42016Efficiently Modeling Long Sequences with Structured State Spaces
Albert Gu, Karan Goel, Christopher Ré
cs.LGarXiv:2111.00396v32021You Are What You Prompt: Prompt Quality, Domain Shift, and Uncertainty in Agrifood Vision-Language Models
Andrea Morales-Garzón, Salvador López-Joya, Miguel López-Pérez +1
cs.CLcs.LGarXiv:2608.18116v12026