Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
5,101 to 5,160 of 20,192
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
Bingyang Wu, Shengyu Liu, Yinmin Zhong +3
cs.DCcs.LGarXiv:2404.09526v22024Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning
Jie Cheng, Gang Xiong, Ruixi Qiao +5
cs.AIcs.LGarXiv:2504.15275v32025Focused Transformer: Contrastive Training for Context Scaling
Szymon Tworkowski, Konrad Staniszewski, Mikołaj Pacek +3
cs.CLcs.AIcs.LGarXiv:2307.03170v22023Chemception: A Deep Neural Network with Minimal Chemistry Knowledge Matches the Performance of Expert-developed QSAR/QSPR Models
Garrett B. Goh, Charles Siegel, Abhinav Vishnu +2
stat.MLcs.AIcs.CEarXiv:1706.06689v12017Do Large Language Model Benchmarks Test Reliability?
Joshua Vendrow, Edward Vendrow, Sara Beery +1
cs.LGcs.CLarXiv:2502.03461v12025Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models
Weichen Fan, Chenyang Si, Junhao Song +16
cs.CVcs.LGarXiv:2501.08453v12025Error Detection for PET/CT Radiology Reports: Domain-Specific vs Large Language Models
Hermione Warr, Harry Anthony, Lilli J Freischem +3
cs.LGcs.AIarXiv:2608.30021v12026On the Instance Hardness as a Decision Criterion in TinyML Systems
Tobiasz Puslecki, Krzysztof Walkowiak
cs.AIcs.LGarXiv:2608.29913v12026Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
Katie Matton, Robert Osazuwa Ness, John Guttag +1
cs.CLcs.AIcs.LGarXiv:2504.14150v22025Adversarial Dropout for Supervised and Semi-supervised Learning
Sungrae Park, Jun-Keon Park, Su-Jin Shin +1
cs.LGcs.CVarXiv:1707.03631v22017On Vanishing Gradients, Over-Smoothing, and Over-Squashing in GNNs: Bridging Recurrent and Graph Learning
Álvaro Arroyo, Alessio Gravina, Benjamin Gutteridge +5
cs.LGcs.AIarXiv:2502.10818v22025Text-to-Image Diffusion Models are Zero-Shot Classifiers
Kevin Clark, Priyank Jaini
cs.CVcs.AIcs.LGarXiv:2303.15233v22023Dataset Pruning: Reducing Training Data by Examining Generalization Influence
Shuo Yang, Zeke Xie, Hanyu Peng +3
cs.LGarXiv:2205.09329v22022The Intervention Gap in Latent World Models
Donna Vakalis
cs.LGarXiv:2608.29998v12026Sparse-Interest Network for Sequential Recommendation
Qiaoyu Tan, Jianwei Zhang, Jiangchao Yao +4
cs.IRcs.LGarXiv:2102.09267v12021Adversarial Attacks on Machine Learning Cybersecurity Defences in Industrial Control Systems
Eirini Anthi, Lowri Williams, Matilda Rhode +2
cs.LGcs.CReess.SParXiv:2004.05005v12020Joint Spatiotemporal Spectral Neural Operators for Learning PDEs on Irregular Domains
Abdolmehdi Behroozi, Chaopeng Shen
cs.LGarXiv:2608.29892v12026Graphon Neural Networks and the Transferability of Graph Neural Networks
Luana Ruiz, Luiz F. O. Chamon, Alejandro Ribeiro
cs.LGstat.MLarXiv:2006.03548v22020The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
Richard Ren, Arunim Agarwal, Mantas Mazeika +13
cs.LGcs.AIcs.CLarXiv:2503.03750v32025HyperDiffusion: Generating Implicit Neural Fields with Weight-Space Diffusion
Ziya Erkoç, Fangchang Ma, Qi Shan +2
cs.CVcs.LGarXiv:2303.17015v12023On the Resilience of Text-to-Video Diffusion Models to Hardware Faults
Zachary Coalson, A M Aahad, Stella Doehring +2
cs.LGarXiv:2608.29598v12026GNN-RAG: Graph Neural Retrieval for Large Language Model Reasoning
Costas Mavromatis, George Karypis
cs.CLcs.AIcs.LGarXiv:2405.20139v12024MedCache: Efficient and Temporally Valid Memory for Longitudinal Clinical Agents
Hei Ting, Chan, Chenwei Wu +5
cs.LGcs.DCcs.MAarXiv:2608.29528v12026LEMUR 2: Unlocking Neural Network Diversity for AI
Tolgay Atinc Uzun, Waleed Khalid, Saif U Din +17
cs.LGcs.CVarXiv:2607.06839v12026SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization
Xianzhi Du, Tsung-Yi Lin, Pengchong Jin +5
cs.CVcs.LGeess.IVarXiv:1912.05027v32019The Curse of Depth in Large Language Models
Wenfang Sun, Xinyuan Song, Pengxiang Li +3
cs.LGcs.AIarXiv:2502.05795v62025VICRegL: Self-Supervised Learning of Local Visual Features
Adrien Bardes, Jean Ponce, Yann LeCun
cs.CVcs.AIcs.LGarXiv:2210.01571v12022Conservative Hybrid Graph Networks for Process Systems with Learned Routing
Paolo Guida
cs.LGarXiv:2608.28896v12026Latent Forcing: Reordering the Diffusion Trajectory for Pixel-Space Image Generation
Alan Baade, Eric Ryan Chan, Kyle Sargent +4
cs.CVcs.LGarXiv:2602.11401v12026ERR+: Sequential Entropy Resolution for Efficient and Decisive LLM Reasoning
Xin Jiang, Minhao Wang, Wen Wu +4
cs.LGarXiv:2608.28771v12026Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
Daniel Wurgaft, Can Rager, Matthew Kowal +13
cs.LGarXiv:2605.05115v12026MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Yixing Jiang, Kameron C. Black, Gloria Geng +4
cs.LGcs.AIcs.MAarXiv:2501.14654v22025MACE-POLAR-1: A Polarisable Electrostatic Foundation Model for Molecular Chemistry
Ilyes Batatia, William J. Baldwin, Domantas Kuryla +10
physics.chem-phcs.LGarXiv:2602.19411v12026Pushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction Tuning
Ted Zadouri, Ahmet Üstün, Arash Ahmadian +3
cs.CLcs.LGarXiv:2309.05444v12023RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation
Leyi Pan, Shuchang Tao, Yunpeng Zhai +5
cs.LGcs.CLarXiv:2606.11709v12026SGTR: End-to-end Scene Graph Generation with Transformer
Rongjie Li, Songyang Zhang, Xuming He
cs.CVcs.LGarXiv:2112.12970v32021Dynamic Trend Fusion Module for Traffic Flow Prediction
Jing Chen, Haocheng Ye, Zhian Ying +2
cs.LGarXiv:2501.10796v12025Interpreting CLIP with Hierarchical Sparse Autoencoders
Vladimir Zaigrajew, Hubert Baniecki, Przemyslaw Biecek
cs.CVcs.AIcs.LGarXiv:2502.20578v22025LSHTC: A Benchmark for Large-Scale Text Classification
Ioannis Partalas, Aris Kosmopoulos, Nicolas Baskiotis +6
cs.IRcs.CLcs.LGarXiv:1503.08581v12015Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling
Liliang Ren, Yang Liu, Yadong Lu +3
cs.CLcs.LGarXiv:2406.07522v32024A Comprehensive Survey of Deep Learning for Multivariate Time Series Forecasting: A Channel Strategy Perspective
Xiangfei Qiu, Hanyin Cheng, Xingjian Wu +5
cs.LGarXiv:2502.10721v32025Degree-Quant: Quantization-Aware Training for Graph Neural Networks
Shyam A. Tailor, Javier Fernandez-Marques, Nicholas D. Lane
cs.LGstat.MLarXiv:2008.05000v32020A User Simulator for Task-Completion Dialogues
Xiujun Li, Zachary C. Lipton, Bhuwan Dhingra +3
cs.LGcs.AIcs.CLarXiv:1612.05688v32016Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models
Zihan Qiu, Zeyu Huang, Bo Zheng +7
cs.LGcs.CLarXiv:2501.11873v22025Proportionally Fair Clustering
Xingyu Chen, Brandon Fain, Liang Lyu +1
cs.LGcs.DScs.GTarXiv:1905.03674v32019Unconstrained Monotonic Neural Networks
Antoine Wehenkel, Gilles Louppe
cs.LGcs.NEstat.MLarXiv:1908.05164v32019APIFlow-Bench: Measuring Whether Agents Survive Long, Dependent API Workflows
Zelin Wan, Arash Nourian, Xiaoxiao Li +2
cs.AIcs.LGcs.SEarXiv:2608.29128v12026MNN: A Universal and Efficient Inference Engine
Xiaotang Jiang, Huan Wang, Yiliu Chen +9
cs.CVcs.DCcs.LGarXiv:2002.12418v12020D2A: A Dataset Built for AI-Based Vulnerability Detection Methods Using Differential Analysis
Yunhui Zheng, Saurabh Pujar, Burn Lewis +6
cs.SEcs.AIcs.LGarXiv:2102.07995v12021Contextual Agent Security: A Policy for Every Purpose
Lillian Tsai, Eugene Bagdasarian
cs.CRcs.CLcs.LGarXiv:2501.17070v32025Unsupervised Latent Space Alignment with Hyperspherical Geodesic Matching
Cameron Ryan, Vivek Sivaraman Narayanaswamy, Kowshik Thopalli +1
cs.LGarXiv:2608.28840v12026VideoRAG: Retrieval-Augmented Generation over Video Corpus
Soyeong Jeong, Kangsan Kim, Jinheon Baek +1
cs.CVcs.AIcs.CLarXiv:2501.05874v32025Randomized Nonlinear Component Analysis
David Lopez-Paz, Suvrit Sra, Alex Smola +2
stat.MLcs.LGarXiv:1402.0119v22014A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
Dong Shu, Xuansheng Wu, Haiyan Zhao +4
cs.LGcs.AIcs.CLarXiv:2503.05613v32025A Graph to Graphs Framework for Retrosynthesis Prediction
Chence Shi, Minkai Xu, Hongyu Guo +2
cs.LGstat.MLarXiv:2003.12725v32020Language Models Use Trigonometry to Do Addition
Subhash Kantamneni, Max Tegmark
cs.AIcs.CLcs.LGarXiv:2502.00873v12025PathBridger: Subgoal Bridges for Offline Goal-Conditioned Reinforcement Learning
Soohyun Choi, Seonvin Cho, Songnam Hong
cs.LGcs.ROarXiv:2608.29061v12026Generation of High-Level Concepts in 3D Scene Graphs via Autoregressive Diffusion
Jose Andres Millan-Romera, Samuel Cognolato, Holger Voos +2
cs.ROcs.LGarXiv:2608.28733v12026Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
Alexander Nikitin, Jannik Kossen, Yarin Gal +1
cs.LGcs.AIcs.CLarXiv:2405.20003v12024Flows for simultaneous manifold learning and density estimation
Johann Brehmer, Kyle Cranmer
stat.MLcs.LGarXiv:2003.13913v32020