Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

15,601 to 15,660 of 20,214

  1. Stochastic Video Generation with a Learned Prior

    Remi Denton, Rob Fergus

    cs.CVcs.AIcs.LGarXiv:1802.07687v22018
  2. Q8BERT: Quantized 8Bit BERT

    Ofir Zafrir, Guy Boudoukh, Peter Izsak +1

    cs.CLcs.LGarXiv:1910.06188v22019
  3. DiD It in 87 Minutes: A Label-Free Softmax-to-Linear Adaptation of Vision Transformers for Object Detection

    Huaiyuan Qin, Gabriel James Goenawan, Zihang Lin +2

    cs.CVcs.LGarXiv:2608.22368v12026
  4. Neural Machine Translation in Linear Time

    Nal Kalchbrenner, Lasse Espeholt, Karen Simonyan +3

    cs.CLcs.LGarXiv:1610.10099v22016
  5. Benchmarking Robustness in Object Detection: Autonomous Driving when Winter is Coming

    Claudio Michaelis, Benjamin Mitzkus, Robert Geirhos +5

    cs.CVcs.LGstat.MLarXiv:1907.07484v22019
  6. Adversarial Agents on Topology Optimization: Understanding the Fragility and Robustness of Deep Learning-based and Physics-Based Design Models under Adversarial Perturbation

    Hoang Anh Nguyen, Yuan Hong, Hongyi Xu

    cs.LGarXiv:2608.22606v12026
  7. Multi-Object Representation Learning with Iterative Variational Inference

    Klaus Greff, Raphaël Lopez Kaufman, Rishabh Kabra +6

    cs.LGcs.CVstat.MLarXiv:1903.00450v32019
  8. Attention in Natural Language Processing

    Andrea Galassi, Marco Lippi, Paolo Torroni

    cs.CLcs.AIcs.LGarXiv:1902.02181v42019
  9. Learning without Memorizing

    Prithviraj Dhar, Rajat Vikram Singh, Kuan-Chuan Peng +2

    cs.CVcs.LGarXiv:1811.08051v22018
  10. LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model

    Ergan Shang, Weijing Tang, Yinqiu He

    cs.CLcs.AIcs.LGarXiv:2608.22295v12026
  11. D$^3$-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation

    Zechen Sun, Zhiwei Zhang, Fei Zhao +7

    cs.LGcs.AIarXiv:2608.24987v12026
  12. Wireless Network Intelligence at the Edge

    Jihong Park, Sumudu Samarakoon, Mehdi Bennis +1

    cs.ITcs.LGcs.NIarXiv:1812.02858v22018
  13. Know-Evolve: Deep Temporal Reasoning for Dynamic Knowledge Graphs

    Rakshit Trivedi, Hanjun Dai, Yichen Wang +1

    cs.AIcs.CLcs.LGarXiv:1705.05742v32017
  14. CD-LoRA: Consistency-Driven Low-Rank Adaptation for Multi-Task Fine-Tuning

    Qian Zha, Jinda Liu, Yuan Wu +1

    cs.LGarXiv:2608.21909v12026
  15. Fast Approximate Nearest Neighbor Search With The Navigating Spreading-out Graph

    Cong Fu, Chao Xiang, Changxu Wang +1

    cs.LGarXiv:1707.00143v102017
  16. On Kernelized Multi-armed Bandits

    Sayak Ray Chowdhury, Aditya Gopalan

    cs.LGarXiv:1704.00445v22017
  17. Attention is Not All You Need: Pure Attention Loses Rank Doubly Exponentially with Depth

    Yihe Dong, Jean-Baptiste Cordonnier, Andreas Loukas

    cs.LGarXiv:2103.03404v22021
  18. Classifying Relations by Ranking with Convolutional Neural Networks

    Cicero Nogueira dos Santos, Bing Xiang, Bowen Zhou

    cs.CLcs.LGcs.NEarXiv:1504.06580v22015
  19. C-RNN-GAN: Continuous recurrent neural networks with adversarial training

    Olof Mogren

    cs.AIcs.LGarXiv:1611.09904v12016
  20. Who Should Teach? Confidence-Aware Dual-Teacher Learning for Few-Shot Node Classification on Text-Attributed Graphs

    Hojin Kim, Sujin Yoon, Sungsu Lim +2

    cs.LGcs.SIarXiv:2608.22127v12026
  21. Conditional validity of inductive conformal predictors

    Vladimir Vovk

    cs.LGarXiv:1209.2673v22012
  22. Spatiotemporal Contrastive Video Representation Learning

    Rui Qian, Tianjian Meng, Boqing Gong +4

    cs.CVcs.LGarXiv:2008.03800v42020
  23. Read, Write, Relax: Why Neural PDE Surrogates Need Both Global and Local Processing

    Anuj Kumar, Heiko Zimmermann, Josiah Bjorgaard +4

    cs.LGcs.AIcs.CEarXiv:2608.21677v12026
  24. CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X

    Qinkai Zheng, Xiao Xia, Xu Zou +10

    cs.LGcs.AIcs.SEarXiv:2303.17568v22023
  25. Revisiting the Calibration of Modern Neural Networks

    Matthias Minderer, Josip Djolonga, Rob Romijnders +5

    cs.LGcs.CVarXiv:2106.07998v22021
  26. DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills

    Xue Bin Peng, Pieter Abbeel, Sergey Levine +1

    cs.GRcs.AIcs.LGarXiv:1804.02717v32018
  27. FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare

    Karim Lekadir, Aasa Feragen, Abdul Joseph Fofanah +117

    cs.CYcs.AIcs.CVarXiv:2309.12325v32023
  28. On the Generalization of Equivariance and Convolution in Neural Networks to the Action of Compact Groups

    Risi Kondor, Shubhendu Trivedi

    stat.MLcs.LGarXiv:1802.03690v32018
  29. Learning from Noisy Labels with Distillation

    Yuncheng Li, Jianchao Yang, Yale Song +3

    cs.CVcs.LGstat.MLarXiv:1703.02391v22017
  30. EDGE: Experience-Distillation for Guided Exploration in Agentic Reinforcement Learning

    Can Xie, Yuyi Zhou, Wen Yang +5

    cs.CLcs.AIcs.LGarXiv:2608.21946v22026
  31. AMOS: A Large-Scale Abdominal Multi-Organ Benchmark for Versatile Medical Image Segmentation

    Yuanfeng Ji, Haotian Bai, Jie Yang +8

    eess.IVcs.CVcs.LGarXiv:2206.08023v32022
  32. CORe50: a New Dataset and Benchmark for Continuous Object Recognition

    Vincenzo Lomonaco, Davide Maltoni

    cs.CVcs.AIcs.LGarXiv:1705.03550v12017
  33. Training generative neural networks via Maximum Mean Discrepancy optimization

    Gintare Karolina Dziugaite, Daniel M. Roy, Zoubin Ghahramani

    stat.MLcs.LGarXiv:1505.03906v12015
  34. ToSCA: Leveraging Hierarchical Reinforcement Learning on Temporal and Strategic Abstractions of Conversational Agents

    Xiaoyu Wang, Qingqing Gu, Yue Zhao +5

    cs.CLcs.HCcs.LGarXiv:2608.21969v12026
  35. Implicit Regularization in Matrix Factorization

    Suriya Gunasekar, Blake Woodworth, Srinadh Bhojanapalli +2

    stat.MLcs.LGarXiv:1705.09280v12017
  36. ALPHABET: A Laplace-Pole History Aggregator with Banked Exponential Transport

    Daehwa Ko, JaeHyeon Kim, Oh Seong Kwon +1

    cs.LGarXiv:2608.24051v12026
  37. Learning Distributed Representations of Sentences from Unlabelled Data

    Felix Hill, Kyunghyun Cho, Anna Korhonen

    cs.CLcs.LGarXiv:1602.03483v12016
  38. Escaping the Big Data Paradigm with Compact Transformers

    Ali Hassani, Steven Walton, Nikhil Shah +3

    cs.CVcs.LGarXiv:2104.05704v42021
  39. Spectral Pre-Filtering for Context-Adaptive Sensor Fusion: A Four-Role FFT-GDCB Integration for High-Stakes Decision Systems

    Oleg Miroshnichenko

    cs.LGarXiv:2608.22023v12026
  40. Bias in Bios: A Case Study of Semantic Representation Bias in a High-Stakes Setting

    Maria De-Arteaga, Alexey Romanov, Hanna Wallach +6

    cs.IRcs.LGstat.MLarXiv:1901.09451v12019
  41. VIMA: General Robot Manipulation with Multimodal Prompts

    Yunfan Jiang, Agrim Gupta, Zichen Zhang +7

    cs.ROcs.AIcs.LGarXiv:2210.03094v22022
  42. MDTE: Minority-Aware Diffusion over Temporal Edge Events for Imbalanced Node Classification

    Zhou Zelong, Zhang Tianming, Yang Zhengyi +4

    cs.LGarXiv:2608.24812v12026
  43. Mode Regularized Generative Adversarial Networks

    Tong Che, Yanran Li, Athul Paul Jacob +2

    cs.LGcs.AIcs.CVarXiv:1612.02136v52016
  44. Steerable CNNs

    Taco S. Cohen, Max Welling

    cs.LGstat.MLarXiv:1612.08498v12016
    Summaries:한국어
  45. Recovering Weighted Tangent Geometry from a Single-Scale Score Field

    Ziqi Zhao, Qingjian Ni

    stat.MLcs.LGarXiv:2608.22334v12026
  46. Width-Independent Compressibility of Deep Neural Networks

    Hong-Yi Wang, Mingze Wang, Liu Ziyin

    cs.LGcond-mat.dis-nncs.ITarXiv:2608.21752v12026
  47. The Blending Ratio Is Not Where the Performance Is: Diagnosing Prototype Blending for Few-Shot Adaptation of Vision-Language Models

    Liangzhi Li, Bowen Wang, Yiming Qian +3

    cs.CVcs.LGarXiv:2608.23634v12026
  48. ChainPrune: Evaluating and Reducing Redundancy in Long Chain-of-Thought Reasoning

    Weihang Pan, Zhengxu Yu, Yuxiang Zhang +5

    cs.LGcs.AIarXiv:2608.21860v12026
  49. BatchEnsemble: An Alternative Approach to Efficient Ensemble and Lifelong Learning

    Yeming Wen, Dustin Tran, Jimmy Ba

    cs.LGstat.MLarXiv:2002.06715v22020
  50. Unsupervised Discovery of Mid-Level Discriminative Patches

    Saurabh Singh, Abhinav Gupta, Alexei A. Efros

    cs.CVcs.AIcs.LGarXiv:1205.3137v22012
  51. clDice -- A Novel Topology-Preserving Loss Function for Tubular Structure Segmentation

    Suprosanna Shit, Johannes C. Paetzold, Anjany Sekuboyina +6

    cs.CVcs.LGeess.IVarXiv:2003.07311v72020
  52. OCGAN: One-class Novelty Detection Using GANs with Constrained Latent Representations

    Pramuditha Perera, Ramesh Nallapati, Bing Xiang

    cs.CVcs.LGarXiv:1903.08550v12019
  53. (More) Efficient Reinforcement Learning via Posterior Sampling

    Ian Osband, Daniel Russo, Benjamin Van Roy

    stat.MLcs.LGarXiv:1306.0940v52013
  54. Multi-Institutional Deep Learning Modeling Without Sharing Patient Data: A Feasibility Study on Brain Tumor Segmentation

    Micah J Sheller, G Anthony Reina, Brandon Edwards +2

    cs.LGstat.MLarXiv:1810.04304v22018
  55. Improving Generalization Performance by Switching from Adam to SGD

    Nitish Shirish Keskar, Richard Socher

    cs.LGmath.OCarXiv:1712.07628v12017
  56. Batch Renormalization: Towards Reducing Minibatch Dependence in Batch-Normalized Models

    Sergey Ioffe

    cs.LGarXiv:1702.03275v22017
  57. Differentially Private Generative Adversarial Network

    Liyang Xie, Kaixiang Lin, Shu Wang +2

    cs.LGcs.CRstat.MLarXiv:1802.06739v12018
  58. Replicable Conformal Prediction

    Marios Papamichalis, Regina Ruane, Theofanis Papamichalis

    stat.MLcs.LGarXiv:2608.23638v12026
  59. Non-convex learning via Stochastic Gradient Langevin Dynamics: a nonasymptotic analysis

    Maxim Raginsky, Alexander Rakhlin, Matus Telgarsky

    cs.LGmath.OCmath.PRarXiv:1702.03849v32017
  60. Lexical Perturbations Disrupt LLM Reasoning: An Empirical Study of Attention Diversion

    Jiaqian Zhu, Yang Zhang, Junhua Ding +1

    cs.CLcs.AIcs.LGarXiv:2608.22140v12026