Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

16,921 to 16,980 of 20,217

  1. Rethinking Graph Transformers with Spectral Attention

    Devin Kreuzer, Dominique Beaini, William L. Hamilton +2

    cs.LGarXiv:2106.03893v32021
  2. Making LLMs Optimize Multi-Scenario CUDA Kernels Like Experts

    Yuxuan Han, Meng-Hao Guo, Zhengning Liu +2

    cs.LGstat.MLarXiv:2603.07169v12026
  3. Efficient RLVR Training via Weighted Mutual Information Data Selection

    Xinyu Zhou, Boyu Zhu, Haotian Zhang +2

    cs.LGcs.CLarXiv:2603.01907v12026
  4. FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance

    Quanhao Li, Zhen Xing, Rui Wang +4

    cs.CVcs.AIcs.LGarXiv:2603.12146v12026
  5. Interactive Benchmarks

    Baoqing Yue, Zihan Zhu, Yutong Han +6

    cs.AIcs.CLcs.LGarXiv:2603.04737v42026
  6. Neural Boltzmann Equations

    Jonas Spinner, Jack Shergold

    hep-phcs.LGstat.MLarXiv:2608.23022v12026
  7. Balanced Meta-Softmax for Long-Tailed Visual Recognition

    Jiawei Ren, Cunjun Yu, Shunan Sheng +4

    cs.LGcs.CVstat.MLarXiv:2007.10740v32020
  8. A Survey on Data Collection for Machine Learning: a Big Data -- AI Integration Perspective

    Yuji Roh, Geon Heo, Steven Euijong Whang

    cs.LGstat.MLarXiv:1811.03402v22018
  9. Independently Recurrent Neural Network (IndRNN): Building A Longer and Deeper RNN

    Shuai Li, Wanqing Li, Chris Cook +2

    cs.CVcs.LGarXiv:1803.04831v32018
  10. Teaching Models to Express Their Uncertainty in Words

    Stephanie Lin, Jacob Hilton, Owain Evans

    cs.CLcs.AIcs.LGarXiv:2205.14334v22022
  11. Sample Efficient Actor-Critic with Experience Replay

    Ziyu Wang, Victor Bapst, Nicolas Heess +4

    cs.LGarXiv:1611.01224v22016
  12. RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

    Songming Liu, Lingxuan Wu, Bangguo Li +6

    cs.ROcs.AIcs.CVarXiv:2410.07864v22024
  13. Credal Large Language Models for Semantic Commitment under Uncertainty

    Shireen Kudukkil Manchingal, Sofiia Nikolenko, Fabio Cuzzolin

    cs.CLcs.AIcs.LGarXiv:2608.23244v12026
  14. Federated Optimization:Distributed Optimization Beyond the Datacenter

    Jakub Konečný, Brendan McMahan, Daniel Ramage

    cs.LGmath.OCarXiv:1511.03575v12015
  15. Deep Semantic Segmentation of Natural and Medical Images: A Review

    Saeid Asgari Taghanaki, Kumar Abhishek, Joseph Paul Cohen +2

    cs.CVcs.LGeess.IVarXiv:1910.07655v42019
  16. Sum-Product Networks: A New Deep Architecture

    Hoifung Poon, Pedro Domingos

    cs.LGcs.AIstat.MLarXiv:1202.3732v12012
  17. A Note on the Inception Score

    Shane Barratt, Rishi Sharma

    stat.MLcs.LGarXiv:1801.01973v22018
  18. Learning Self-Correction in Vision-Language Models via Rollout Augmentation

    Yi Ding, Ziliang Qiu, Bolian Li +1

    cs.CVcs.CLcs.LGarXiv:2602.08503v22026
  19. Discovering Latent Knowledge in Language Models Without Supervision

    Collin Burns, Haotian Ye, Dan Klein +1

    cs.CLcs.AIcs.LGarXiv:2212.03827v22022
  20. DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

    DeepSeek-AI, :, Xiao Bi +85

    cs.CLcs.AIcs.LGarXiv:2401.02954v12024
  21. GradMem: Learning to Write Context into Memory with Test-Time Gradient Descent

    Yuri Kuratov, Matvey Kairov, Aydar Bulatov +2

    cs.CLcs.LGarXiv:2603.13875v22026
  22. Bayesian Convolutional Neural Networks with Bernoulli Approximate Variational Inference

    Yarin Gal, Zoubin Ghahramani

    stat.MLcs.LGarXiv:1506.02158v62015
  23. Bayesian Deep Learning and a Probabilistic Perspective of Generalization

    Andrew Gordon Wilson, Pavel Izmailov

    cs.LGstat.MLarXiv:2002.08791v42020
  24. cGANs with Projection Discriminator

    Takeru Miyato, Masanori Koyama

    cs.LGcs.CVstat.MLarXiv:1802.05637v22018
  25. Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors

    Zhiwei Zhang, Fei Zhao, Rui Wang +6

    cs.LGcs.AIarXiv:2601.15625v22026
  26. Rumor Detection on Social Media with Bi-Directional Graph Convolutional Networks

    Tian Bian, Xi Xiao, Tingyang Xu +4

    cs.SIcs.LGarXiv:2001.06362v12020
  27. Learning Scheduling Algorithms for Data Processing Clusters

    Hongzi Mao, Malte Schwarzkopf, Shaileshh Bojja Venkatakrishnan +2

    cs.LGstat.MLarXiv:1810.01963v42018
  28. Distilling Conversations: Abstract Compression of Conversational Audio Context for LLM-based ASR

    Shashi Kumar, Esaú Villatoro-Tello, Sergio Burdisso +7

    cs.CLcs.AIcs.LGarXiv:2603.26246v12026
  29. Doubly Robust Policy Evaluation and Learning

    Miroslav Dudik, John Langford, Lihong Li

    cs.LGcs.AIcs.ROarXiv:1103.4601v22011
  30. Graph neural networks for materials science and chemistry

    Patrick Reiser, Marlen Neubert, André Eberhard +8

    physics.chem-phcond-mat.mtrl-scics.LGarXiv:2208.09481v12022
  31. TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data

    Pengcheng Yin, Graham Neubig, Wen-tau Yih +1

    cs.CLcs.LGarXiv:2005.08314v12020
  32. Effective Distillation to Hybrid xLSTM Architectures

    Lukas Hauzenberger, Niklas Schmidinger, Thomas Schmied +7

    cs.LGarXiv:2603.15590v22026
  33. Style Transfer from Non-Parallel Text by Cross-Alignment

    Tianxiao Shen, Tao Lei, Regina Barzilay +1

    cs.CLcs.LGarXiv:1705.09655v22017
  34. End-to-End Safe Reinforcement Learning through Barrier Functions for Safety-Critical Continuous Control Tasks

    Richard Cheng, Gabor Orosz, Richard M. Murray +1

    cs.LGeess.SYstat.MLarXiv:1903.08792v12019
  35. Insight-V++: Towards Advanced Long-Chain Visual Reasoning with Multimodal Large Language Models

    Yuhao Dong, Zuyan Liu, Shulin Tian +2

    cs.CVcs.AIcs.LGarXiv:2603.18118v12026
  36. BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning

    Eric Jang, Alex Irpan, Mohi Khansari +5

    cs.ROcs.LGarXiv:2202.02005v12022
  37. Large Language Models Are Zero-Shot Time Series Forecasters

    Nate Gruver, Marc Finzi, Shikai Qiu +1

    cs.LGarXiv:2310.07820v32023
  38. A decoder-only foundation model for time-series forecasting

    Abhimanyu Das, Weihao Kong, Rajat Sen +1

    cs.CLcs.AIcs.LGarXiv:2310.10688v42023
  39. Flow-based Extremal Mathematical Structure Discovery

    Gergely Bérczi, Baran Hashemi, Jonas Klüver

    math.COcs.LGarXiv:2601.18005v12026
  40. Neural Fields in Visual Computing and Beyond

    Yiheng Xie, Towaki Takikawa, Shunsuke Saito +7

    cs.CVcs.GRcs.LGarXiv:2111.11426v42021
  41. FROST: Filtering Reasoning Outliers with Attention for Efficient Reasoning

    Haozheng Luo, Zhuolin Jiang, Md Zahid Hasan +2

    cs.CLcs.AIcs.LGarXiv:2601.19001v22026
  42. Self-Improving Pretraining: using post-trained models to pretrain better models

    Ellen Xiaoqing Tan, Jack Lanchantin, Shehzaad Dhuliawala +9

    cs.CLcs.AIcs.LGarXiv:2601.21343v32026
  43. HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose Estimation

    Bowen Cheng, Bin Xiao, Jingdong Wang +3

    cs.CVcs.LGeess.IVarXiv:1908.10357v32019
  44. Contrastive Representation-Guided Genetic Minority Oversampling for Imbalanced Time-Series Classification

    Wenbin Pei, Yunrong Hao, Zhen Liu +4

    cs.LGarXiv:2608.22804v12026
  45. PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

    Yanli Zhao, Andrew Gu, Rohan Varma +15

    cs.DCcs.AIcs.LGarXiv:2304.11277v22023
  46. Improving Unsupervised Defect Segmentation by Applying Structural Similarity to Autoencoders

    Paul Bergmann, Sindy Löwe, Michael Fauser +2

    cs.CVcs.LGarXiv:1807.02011v32018
  47. NAS-Bench-101: Towards Reproducible Neural Architecture Search

    Chris Ying, Aaron Klein, Esteban Real +3

    cs.LGstat.MLarXiv:1902.09635v22019
  48. Multi-task Sequence to Sequence Learning

    Minh-Thang Luong, Quoc V. Le, Ilya Sutskever +2

    cs.LGcs.CLstat.MLarXiv:1511.06114v42015
  49. Variable Rate Image Compression with Recurrent Neural Networks

    George Toderici, Sean M. O'Malley, Sung Jin Hwang +5

    cs.CVcs.LGcs.NEarXiv:1511.06085v52015
  50. Beyond Low-frequency Information in Graph Convolutional Networks

    Deyu Bo, Xiao Wang, Chuan Shi +1

    cs.LGcs.SIarXiv:2101.00797v12021
  51. The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus

    Amartya Roy, Rasul Tutunov, Xiaotong Ji +2

    cs.LGcs.AIarXiv:2603.20105v12026
  52. WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models

    Runjie Zhou, Youbo Shao, Haoyu Lu +16

    cs.CVcs.LGarXiv:2602.02537v12026
  53. Benchmarking Large Language Models for News Summarization

    Tianyi Zhang, Faisal Ladhak, Esin Durmus +3

    cs.CLcs.AIcs.LGarXiv:2301.13848v12023
  54. Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation

    Pingzhi Tang, Yiding Wang, Muhan Zhang

    cs.LGcs.AIcs.CLarXiv:2601.11258v22026
  55. VoxServe: Streaming-Centric Serving System for Speech Language Models

    Keisuke Kamahori, Wei-Tzu Lee, Atindra Jha +4

    cs.LGcs.AIcs.DCarXiv:2602.00269v12026
  56. QuantLRM: Quantization of Large Reasoning Models via Fine-Tuning Signals

    Nan Zhang, Eugene Kwek, Yusen Zhang +4

    cs.LGcs.AIarXiv:2602.02581v12026
  57. Automatic Differentiation Variational Inference

    Alp Kucukelbir, Dustin Tran, Rajesh Ranganath +2

    stat.MLcs.AIcs.LGarXiv:1603.00788v12016
  58. Aligning Agentic World Models via Knowledgeable Experience Learning

    Baochang Ren, Yunzhi Yao, Rui Sun +3

    cs.CLcs.AIcs.CVarXiv:2601.13247v12026
  59. Spiking Neural Networks for Continuous Control: Neuromorphic Reinforcement Learning in Conventional Computing

    Jessica Hunter, Md Maruf Hossain Shuvo, Krishna Roy

    cs.LGcs.NEarXiv:2608.22729v12026
  60. SEVerA: Verified Synthesis of Self-Evolving Agents

    Debangshu Banerjee, Changming Xu, Eugene Ie +4

    cs.LGcs.PLcs.SEarXiv:2603.25111v22026