Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

5,101 to 5,160 of 20,192

  1. LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism

    Bingyang Wu, Shengyu Liu, Yinmin Zhong +3

    cs.DCcs.LGarXiv:2404.09526v22024
  2. Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning

    Jie Cheng, Gang Xiong, Ruixi Qiao +5

    cs.AIcs.LGarXiv:2504.15275v32025
  3. Focused Transformer: Contrastive Training for Context Scaling

    Szymon Tworkowski, Konrad Staniszewski, Mikołaj Pacek +3

    cs.CLcs.AIcs.LGarXiv:2307.03170v22023
  4. Chemception: A Deep Neural Network with Minimal Chemistry Knowledge Matches the Performance of Expert-developed QSAR/QSPR Models

    Garrett B. Goh, Charles Siegel, Abhinav Vishnu +2

    stat.MLcs.AIcs.CEarXiv:1706.06689v12017
  5. Do Large Language Model Benchmarks Test Reliability?

    Joshua Vendrow, Edward Vendrow, Sara Beery +1

    cs.LGcs.CLarXiv:2502.03461v12025
  6. Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

    Weichen Fan, Chenyang Si, Junhao Song +16

    cs.CVcs.LGarXiv:2501.08453v12025
  7. Error Detection for PET/CT Radiology Reports: Domain-Specific vs Large Language Models

    Hermione Warr, Harry Anthony, Lilli J Freischem +3

    cs.LGcs.AIarXiv:2608.30021v12026
  8. On the Instance Hardness as a Decision Criterion in TinyML Systems

    Tobiasz Puslecki, Krzysztof Walkowiak

    cs.AIcs.LGarXiv:2608.29913v12026
  9. Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations

    Katie Matton, Robert Osazuwa Ness, John Guttag +1

    cs.CLcs.AIcs.LGarXiv:2504.14150v22025
  10. Adversarial Dropout for Supervised and Semi-supervised Learning

    Sungrae Park, Jun-Keon Park, Su-Jin Shin +1

    cs.LGcs.CVarXiv:1707.03631v22017
  11. On Vanishing Gradients, Over-Smoothing, and Over-Squashing in GNNs: Bridging Recurrent and Graph Learning

    Álvaro Arroyo, Alessio Gravina, Benjamin Gutteridge +5

    cs.LGcs.AIarXiv:2502.10818v22025
  12. Text-to-Image Diffusion Models are Zero-Shot Classifiers

    Kevin Clark, Priyank Jaini

    cs.CVcs.AIcs.LGarXiv:2303.15233v22023
  13. Dataset Pruning: Reducing Training Data by Examining Generalization Influence

    Shuo Yang, Zeke Xie, Hanyu Peng +3

    cs.LGarXiv:2205.09329v22022
  14. The Intervention Gap in Latent World Models

    Donna Vakalis

    cs.LGarXiv:2608.29998v12026
  15. Sparse-Interest Network for Sequential Recommendation

    Qiaoyu Tan, Jianwei Zhang, Jiangchao Yao +4

    cs.IRcs.LGarXiv:2102.09267v12021
  16. Adversarial Attacks on Machine Learning Cybersecurity Defences in Industrial Control Systems

    Eirini Anthi, Lowri Williams, Matilda Rhode +2

    cs.LGcs.CReess.SParXiv:2004.05005v12020
  17. Joint Spatiotemporal Spectral Neural Operators for Learning PDEs on Irregular Domains

    Abdolmehdi Behroozi, Chaopeng Shen

    cs.LGarXiv:2608.29892v12026
  18. Graphon Neural Networks and the Transferability of Graph Neural Networks

    Luana Ruiz, Luiz F. O. Chamon, Alejandro Ribeiro

    cs.LGstat.MLarXiv:2006.03548v22020
  19. The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems

    Richard Ren, Arunim Agarwal, Mantas Mazeika +13

    cs.LGcs.AIcs.CLarXiv:2503.03750v32025
  20. HyperDiffusion: Generating Implicit Neural Fields with Weight-Space Diffusion

    Ziya Erkoç, Fangchang Ma, Qi Shan +2

    cs.CVcs.LGarXiv:2303.17015v12023
  21. On the Resilience of Text-to-Video Diffusion Models to Hardware Faults

    Zachary Coalson, A M Aahad, Stella Doehring +2

    cs.LGarXiv:2608.29598v12026
  22. GNN-RAG: Graph Neural Retrieval for Large Language Model Reasoning

    Costas Mavromatis, George Karypis

    cs.CLcs.AIcs.LGarXiv:2405.20139v12024
  23. MedCache: Efficient and Temporally Valid Memory for Longitudinal Clinical Agents

    Hei Ting, Chan, Chenwei Wu +5

    cs.LGcs.DCcs.MAarXiv:2608.29528v12026
  24. LEMUR 2: Unlocking Neural Network Diversity for AI

    Tolgay Atinc Uzun, Waleed Khalid, Saif U Din +17

    cs.LGcs.CVarXiv:2607.06839v12026
  25. SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization

    Xianzhi Du, Tsung-Yi Lin, Pengchong Jin +5

    cs.CVcs.LGeess.IVarXiv:1912.05027v32019
  26. The Curse of Depth in Large Language Models

    Wenfang Sun, Xinyuan Song, Pengxiang Li +3

    cs.LGcs.AIarXiv:2502.05795v62025
  27. VICRegL: Self-Supervised Learning of Local Visual Features

    Adrien Bardes, Jean Ponce, Yann LeCun

    cs.CVcs.AIcs.LGarXiv:2210.01571v12022
  28. Conservative Hybrid Graph Networks for Process Systems with Learned Routing

    Paolo Guida

    cs.LGarXiv:2608.28896v12026
  29. Latent Forcing: Reordering the Diffusion Trajectory for Pixel-Space Image Generation

    Alan Baade, Eric Ryan Chan, Kyle Sargent +4

    cs.CVcs.LGarXiv:2602.11401v12026
  30. ERR+: Sequential Entropy Resolution for Efficient and Decisive LLM Reasoning

    Xin Jiang, Minhao Wang, Wen Wu +4

    cs.LGarXiv:2608.28771v12026
  31. Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior

    Daniel Wurgaft, Can Rager, Matthew Kowal +13

    cs.LGarXiv:2605.05115v12026
  32. MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

    Yixing Jiang, Kameron C. Black, Gloria Geng +4

    cs.LGcs.AIcs.MAarXiv:2501.14654v22025
  33. MACE-POLAR-1: A Polarisable Electrostatic Foundation Model for Molecular Chemistry

    Ilyes Batatia, William J. Baldwin, Domantas Kuryla +10

    physics.chem-phcs.LGarXiv:2602.19411v12026
  34. Pushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction Tuning

    Ted Zadouri, Ahmet Üstün, Arash Ahmadian +3

    cs.CLcs.LGarXiv:2309.05444v12023
  35. RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

    Leyi Pan, Shuchang Tao, Yunpeng Zhai +5

    cs.LGcs.CLarXiv:2606.11709v12026
  36. SGTR: End-to-end Scene Graph Generation with Transformer

    Rongjie Li, Songyang Zhang, Xuming He

    cs.CVcs.LGarXiv:2112.12970v32021
  37. Dynamic Trend Fusion Module for Traffic Flow Prediction

    Jing Chen, Haocheng Ye, Zhian Ying +2

    cs.LGarXiv:2501.10796v12025
  38. Interpreting CLIP with Hierarchical Sparse Autoencoders

    Vladimir Zaigrajew, Hubert Baniecki, Przemyslaw Biecek

    cs.CVcs.AIcs.LGarXiv:2502.20578v22025
  39. LSHTC: A Benchmark for Large-Scale Text Classification

    Ioannis Partalas, Aris Kosmopoulos, Nicolas Baskiotis +6

    cs.IRcs.CLcs.LGarXiv:1503.08581v12015
  40. Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling

    Liliang Ren, Yang Liu, Yadong Lu +3

    cs.CLcs.LGarXiv:2406.07522v32024
  41. A Comprehensive Survey of Deep Learning for Multivariate Time Series Forecasting: A Channel Strategy Perspective

    Xiangfei Qiu, Hanyin Cheng, Xingjian Wu +5

    cs.LGarXiv:2502.10721v32025
  42. Degree-Quant: Quantization-Aware Training for Graph Neural Networks

    Shyam A. Tailor, Javier Fernandez-Marques, Nicholas D. Lane

    cs.LGstat.MLarXiv:2008.05000v32020
  43. A User Simulator for Task-Completion Dialogues

    Xiujun Li, Zachary C. Lipton, Bhuwan Dhingra +3

    cs.LGcs.AIcs.CLarXiv:1612.05688v32016
  44. Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models

    Zihan Qiu, Zeyu Huang, Bo Zheng +7

    cs.LGcs.CLarXiv:2501.11873v22025
  45. Proportionally Fair Clustering

    Xingyu Chen, Brandon Fain, Liang Lyu +1

    cs.LGcs.DScs.GTarXiv:1905.03674v32019
  46. Unconstrained Monotonic Neural Networks

    Antoine Wehenkel, Gilles Louppe

    cs.LGcs.NEstat.MLarXiv:1908.05164v32019
  47. APIFlow-Bench: Measuring Whether Agents Survive Long, Dependent API Workflows

    Zelin Wan, Arash Nourian, Xiaoxiao Li +2

    cs.AIcs.LGcs.SEarXiv:2608.29128v12026
  48. MNN: A Universal and Efficient Inference Engine

    Xiaotang Jiang, Huan Wang, Yiliu Chen +9

    cs.CVcs.DCcs.LGarXiv:2002.12418v12020
  49. D2A: A Dataset Built for AI-Based Vulnerability Detection Methods Using Differential Analysis

    Yunhui Zheng, Saurabh Pujar, Burn Lewis +6

    cs.SEcs.AIcs.LGarXiv:2102.07995v12021
  50. Contextual Agent Security: A Policy for Every Purpose

    Lillian Tsai, Eugene Bagdasarian

    cs.CRcs.CLcs.LGarXiv:2501.17070v32025
  51. Unsupervised Latent Space Alignment with Hyperspherical Geodesic Matching

    Cameron Ryan, Vivek Sivaraman Narayanaswamy, Kowshik Thopalli +1

    cs.LGarXiv:2608.28840v12026
  52. VideoRAG: Retrieval-Augmented Generation over Video Corpus

    Soyeong Jeong, Kangsan Kim, Jinheon Baek +1

    cs.CVcs.AIcs.CLarXiv:2501.05874v32025
  53. Randomized Nonlinear Component Analysis

    David Lopez-Paz, Suvrit Sra, Alex Smola +2

    stat.MLcs.LGarXiv:1402.0119v22014
  54. A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models

    Dong Shu, Xuansheng Wu, Haiyan Zhao +4

    cs.LGcs.AIcs.CLarXiv:2503.05613v32025
  55. A Graph to Graphs Framework for Retrosynthesis Prediction

    Chence Shi, Minkai Xu, Hongyu Guo +2

    cs.LGstat.MLarXiv:2003.12725v32020
  56. Language Models Use Trigonometry to Do Addition

    Subhash Kantamneni, Max Tegmark

    cs.AIcs.CLcs.LGarXiv:2502.00873v12025
  57. PathBridger: Subgoal Bridges for Offline Goal-Conditioned Reinforcement Learning

    Soohyun Choi, Seonvin Cho, Songnam Hong

    cs.LGcs.ROarXiv:2608.29061v12026
  58. Generation of High-Level Concepts in 3D Scene Graphs via Autoregressive Diffusion

    Jose Andres Millan-Romera, Samuel Cognolato, Holger Voos +2

    cs.ROcs.LGarXiv:2608.28733v12026
  59. Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities

    Alexander Nikitin, Jannik Kossen, Yarin Gal +1

    cs.LGcs.AIcs.CLarXiv:2405.20003v12024
  60. Flows for simultaneous manifold learning and density estimation

    Johann Brehmer, Kyle Cranmer

    stat.MLcs.LGarXiv:2003.13913v32020