Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
16,921 to 16,980 of 20,217
Rethinking Graph Transformers with Spectral Attention
Devin Kreuzer, Dominique Beaini, William L. Hamilton +2
cs.LGarXiv:2106.03893v32021Making LLMs Optimize Multi-Scenario CUDA Kernels Like Experts
Yuxuan Han, Meng-Hao Guo, Zhengning Liu +2
cs.LGstat.MLarXiv:2603.07169v12026Efficient RLVR Training via Weighted Mutual Information Data Selection
Xinyu Zhou, Boyu Zhu, Haotian Zhang +2
cs.LGcs.CLarXiv:2603.01907v12026FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance
Quanhao Li, Zhen Xing, Rui Wang +4
cs.CVcs.AIcs.LGarXiv:2603.12146v12026Interactive Benchmarks
Baoqing Yue, Zihan Zhu, Yutong Han +6
cs.AIcs.CLcs.LGarXiv:2603.04737v42026Neural Boltzmann Equations
Jonas Spinner, Jack Shergold
hep-phcs.LGstat.MLarXiv:2608.23022v12026Balanced Meta-Softmax for Long-Tailed Visual Recognition
Jiawei Ren, Cunjun Yu, Shunan Sheng +4
cs.LGcs.CVstat.MLarXiv:2007.10740v32020A Survey on Data Collection for Machine Learning: a Big Data -- AI Integration Perspective
Yuji Roh, Geon Heo, Steven Euijong Whang
cs.LGstat.MLarXiv:1811.03402v22018Independently Recurrent Neural Network (IndRNN): Building A Longer and Deeper RNN
Shuai Li, Wanqing Li, Chris Cook +2
cs.CVcs.LGarXiv:1803.04831v32018Teaching Models to Express Their Uncertainty in Words
Stephanie Lin, Jacob Hilton, Owain Evans
cs.CLcs.AIcs.LGarXiv:2205.14334v22022Sample Efficient Actor-Critic with Experience Replay
Ziyu Wang, Victor Bapst, Nicolas Heess +4
cs.LGarXiv:1611.01224v22016RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
Songming Liu, Lingxuan Wu, Bangguo Li +6
cs.ROcs.AIcs.CVarXiv:2410.07864v22024Credal Large Language Models for Semantic Commitment under Uncertainty
Shireen Kudukkil Manchingal, Sofiia Nikolenko, Fabio Cuzzolin
cs.CLcs.AIcs.LGarXiv:2608.23244v12026Federated Optimization:Distributed Optimization Beyond the Datacenter
Jakub Konečný, Brendan McMahan, Daniel Ramage
cs.LGmath.OCarXiv:1511.03575v12015Deep Semantic Segmentation of Natural and Medical Images: A Review
Saeid Asgari Taghanaki, Kumar Abhishek, Joseph Paul Cohen +2
cs.CVcs.LGeess.IVarXiv:1910.07655v42019Sum-Product Networks: A New Deep Architecture
Hoifung Poon, Pedro Domingos
cs.LGcs.AIstat.MLarXiv:1202.3732v12012A Note on the Inception Score
Shane Barratt, Rishi Sharma
stat.MLcs.LGarXiv:1801.01973v22018Learning Self-Correction in Vision-Language Models via Rollout Augmentation
Yi Ding, Ziliang Qiu, Bolian Li +1
cs.CVcs.CLcs.LGarXiv:2602.08503v22026Discovering Latent Knowledge in Language Models Without Supervision
Collin Burns, Haotian Ye, Dan Klein +1
cs.CLcs.AIcs.LGarXiv:2212.03827v22022DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
DeepSeek-AI, :, Xiao Bi +85
cs.CLcs.AIcs.LGarXiv:2401.02954v12024GradMem: Learning to Write Context into Memory with Test-Time Gradient Descent
Yuri Kuratov, Matvey Kairov, Aydar Bulatov +2
cs.CLcs.LGarXiv:2603.13875v22026Bayesian Convolutional Neural Networks with Bernoulli Approximate Variational Inference
Yarin Gal, Zoubin Ghahramani
stat.MLcs.LGarXiv:1506.02158v62015Bayesian Deep Learning and a Probabilistic Perspective of Generalization
Andrew Gordon Wilson, Pavel Izmailov
cs.LGstat.MLarXiv:2002.08791v42020cGANs with Projection Discriminator
Takeru Miyato, Masanori Koyama
cs.LGcs.CVstat.MLarXiv:1802.05637v22018Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors
Zhiwei Zhang, Fei Zhao, Rui Wang +6
cs.LGcs.AIarXiv:2601.15625v22026Rumor Detection on Social Media with Bi-Directional Graph Convolutional Networks
Tian Bian, Xi Xiao, Tingyang Xu +4
cs.SIcs.LGarXiv:2001.06362v12020Learning Scheduling Algorithms for Data Processing Clusters
Hongzi Mao, Malte Schwarzkopf, Shaileshh Bojja Venkatakrishnan +2
cs.LGstat.MLarXiv:1810.01963v42018Distilling Conversations: Abstract Compression of Conversational Audio Context for LLM-based ASR
Shashi Kumar, Esaú Villatoro-Tello, Sergio Burdisso +7
cs.CLcs.AIcs.LGarXiv:2603.26246v12026Doubly Robust Policy Evaluation and Learning
Miroslav Dudik, John Langford, Lihong Li
cs.LGcs.AIcs.ROarXiv:1103.4601v22011Graph neural networks for materials science and chemistry
Patrick Reiser, Marlen Neubert, André Eberhard +8
physics.chem-phcond-mat.mtrl-scics.LGarXiv:2208.09481v12022TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data
Pengcheng Yin, Graham Neubig, Wen-tau Yih +1
cs.CLcs.LGarXiv:2005.08314v12020Effective Distillation to Hybrid xLSTM Architectures
Lukas Hauzenberger, Niklas Schmidinger, Thomas Schmied +7
cs.LGarXiv:2603.15590v22026Style Transfer from Non-Parallel Text by Cross-Alignment
Tianxiao Shen, Tao Lei, Regina Barzilay +1
cs.CLcs.LGarXiv:1705.09655v22017End-to-End Safe Reinforcement Learning through Barrier Functions for Safety-Critical Continuous Control Tasks
Richard Cheng, Gabor Orosz, Richard M. Murray +1
cs.LGeess.SYstat.MLarXiv:1903.08792v12019Insight-V++: Towards Advanced Long-Chain Visual Reasoning with Multimodal Large Language Models
Yuhao Dong, Zuyan Liu, Shulin Tian +2
cs.CVcs.AIcs.LGarXiv:2603.18118v12026BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning
Eric Jang, Alex Irpan, Mohi Khansari +5
cs.ROcs.LGarXiv:2202.02005v12022Large Language Models Are Zero-Shot Time Series Forecasters
Nate Gruver, Marc Finzi, Shikai Qiu +1
cs.LGarXiv:2310.07820v32023A decoder-only foundation model for time-series forecasting
Abhimanyu Das, Weihao Kong, Rajat Sen +1
cs.CLcs.AIcs.LGarXiv:2310.10688v42023Flow-based Extremal Mathematical Structure Discovery
Gergely Bérczi, Baran Hashemi, Jonas Klüver
math.COcs.LGarXiv:2601.18005v12026Neural Fields in Visual Computing and Beyond
Yiheng Xie, Towaki Takikawa, Shunsuke Saito +7
cs.CVcs.GRcs.LGarXiv:2111.11426v42021FROST: Filtering Reasoning Outliers with Attention for Efficient Reasoning
Haozheng Luo, Zhuolin Jiang, Md Zahid Hasan +2
cs.CLcs.AIcs.LGarXiv:2601.19001v22026Self-Improving Pretraining: using post-trained models to pretrain better models
Ellen Xiaoqing Tan, Jack Lanchantin, Shehzaad Dhuliawala +9
cs.CLcs.AIcs.LGarXiv:2601.21343v32026HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose Estimation
Bowen Cheng, Bin Xiao, Jingdong Wang +3
cs.CVcs.LGeess.IVarXiv:1908.10357v32019Contrastive Representation-Guided Genetic Minority Oversampling for Imbalanced Time-Series Classification
Wenbin Pei, Yunrong Hao, Zhen Liu +4
cs.LGarXiv:2608.22804v12026PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Yanli Zhao, Andrew Gu, Rohan Varma +15
cs.DCcs.AIcs.LGarXiv:2304.11277v22023Improving Unsupervised Defect Segmentation by Applying Structural Similarity to Autoencoders
Paul Bergmann, Sindy Löwe, Michael Fauser +2
cs.CVcs.LGarXiv:1807.02011v32018NAS-Bench-101: Towards Reproducible Neural Architecture Search
Chris Ying, Aaron Klein, Esteban Real +3
cs.LGstat.MLarXiv:1902.09635v22019Multi-task Sequence to Sequence Learning
Minh-Thang Luong, Quoc V. Le, Ilya Sutskever +2
cs.LGcs.CLstat.MLarXiv:1511.06114v42015Variable Rate Image Compression with Recurrent Neural Networks
George Toderici, Sean M. O'Malley, Sung Jin Hwang +5
cs.CVcs.LGcs.NEarXiv:1511.06085v52015Beyond Low-frequency Information in Graph Convolutional Networks
Deyu Bo, Xiao Wang, Chuan Shi +1
cs.LGcs.SIarXiv:2101.00797v12021The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus
Amartya Roy, Rasul Tutunov, Xiaotong Ji +2
cs.LGcs.AIarXiv:2603.20105v12026WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models
Runjie Zhou, Youbo Shao, Haoyu Lu +16
cs.CVcs.LGarXiv:2602.02537v12026Benchmarking Large Language Models for News Summarization
Tianyi Zhang, Faisal Ladhak, Esin Durmus +3
cs.CLcs.AIcs.LGarXiv:2301.13848v12023Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation
Pingzhi Tang, Yiding Wang, Muhan Zhang
cs.LGcs.AIcs.CLarXiv:2601.11258v22026VoxServe: Streaming-Centric Serving System for Speech Language Models
Keisuke Kamahori, Wei-Tzu Lee, Atindra Jha +4
cs.LGcs.AIcs.DCarXiv:2602.00269v12026QuantLRM: Quantization of Large Reasoning Models via Fine-Tuning Signals
Nan Zhang, Eugene Kwek, Yusen Zhang +4
cs.LGcs.AIarXiv:2602.02581v12026Automatic Differentiation Variational Inference
Alp Kucukelbir, Dustin Tran, Rajesh Ranganath +2
stat.MLcs.AIcs.LGarXiv:1603.00788v12016Aligning Agentic World Models via Knowledgeable Experience Learning
Baochang Ren, Yunzhi Yao, Rui Sun +3
cs.CLcs.AIcs.CVarXiv:2601.13247v12026Spiking Neural Networks for Continuous Control: Neuromorphic Reinforcement Learning in Conventional Computing
Jessica Hunter, Md Maruf Hossain Shuvo, Krishna Roy
cs.LGcs.NEarXiv:2608.22729v12026SEVerA: Verified Synthesis of Self-Evolving Agents
Debangshu Banerjee, Changming Xu, Eugene Ie +4
cs.LGcs.PLcs.SEarXiv:2603.25111v22026