Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
8,401 to 8,460 of 20,218
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Tianzhe Chu, Yuexiang Zhai, Jihan Yang +6
cs.AIcs.CVcs.LGarXiv:2501.17161v22025Motion-Aware Feature for Improved Video Anomaly Detection
Yi Zhu, Shawn Newsam
cs.CVcs.LGeess.IVarXiv:1907.10211v12019Mean Flows for One-step Generative Modeling
Zhengyang Geng, Mingyang Deng, Xingjian Bai +2
cs.LGcs.CVarXiv:2505.13447v12025V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Mido Assran, Adrien Bardes, David Fan +27
cs.AIcs.CVcs.LGarXiv:2506.09985v12025Group Sequence Policy Optimization
Chujie Zheng, Shixuan Liu, Mingze Li +9
cs.LGcs.AIcs.CLarXiv:2507.18071v22025Self-Training: A Survey
Massih-Reza Amini, Vasilii Feofanov, Loic Pauletto +3
cs.LGarXiv:2202.12040v62022Human-AI Collaboration via Conditional Delegation: A Case Study of Content Moderation
Vivian Lai, Samuel Carton, Rajat Bhatnagar +3
cs.AIcs.HCcs.LGarXiv:2204.11788v12022Abnormal respiratory patterns classifier may contribute to large-scale screening of people infected with COVID-19 in an accurate and unobtrusive manner
Yunlu Wang, Menghan Hu, Qingli Li +3
cs.LGcs.CVeess.SParXiv:2002.05534v22020Conditional Generative Neural System for Probabilistic Trajectory Prediction
Jiachen Li, Hengbo Ma, Masayoshi Tomizuka
cs.CVcs.AIcs.LGarXiv:1905.01631v22019Generative Adversarial Active Learning
Jia-Jie Zhu, José Bento
cs.LGstat.MLarXiv:1702.07956v52017Beyond Blur: A Semantic Tri-view Pipeline for Teledermatology Gradability via Skin Micro-relief
Robert Engel
eess.IVcs.CVcs.HCarXiv:2609.03095v12026Towards Interpretable Semantic Segmentation via Gradient-weighted Class Activation Mapping
Kira Vinogradova, Alexandr Dibrov, Gene Myers
cs.CVcs.LGeess.IVarXiv:2002.11434v12020Supervised Classification Performance of Multispectral Images
K. Perumal, R. Bhaskaran
cs.LGcs.CVarXiv:1002.4046v12010From Euclidean to Graph-Structured Data: A Survey of Collaborative Learning
Rémi Bourgerie, Šarūnas Girdzijauskas, Viktoria Fodor
cs.LGcs.MAcs.SIarXiv:2609.02984v12026Scene Parsing with Multiscale Feature Learning, Purity Trees, and Optimal Covers
Clément Farabet, Camille Couprie, Laurent Najman +1
cs.CVcs.LGarXiv:1202.2160v22012Automated Speed and Lane Change Decision Making using Deep Reinforcement Learning
Carl-Johan Hoel, Krister Wolff, Leo Laine
cs.ROcs.AIcs.LGarXiv:1803.10056v22018Linear model predictive safety certification for learning-based control
Kim P. Wabersich, Melanie N. Zeilinger
eess.SYcs.LGarXiv:1803.08552v62018On the Interaction Between Model Compression and Test-Time Adaptation
Francesco Corti, Dong Wang, Young D. Kwon +2
cs.LGcs.AIarXiv:2609.03604v12026Learning efficient sparse and low rank models
Pablo Sprechmann, Alex M. Bronstein, Guillermo Sapiro
cs.LGarXiv:1212.3631v12012Investigating Linear Probe Robustness to Linguistic Register, Medical Specialty, and Corpus Shifts in Medical QA
Nishant Mishra, Ameen Abu-Hanna, Iacer Calixto
cs.CLcs.LGarXiv:2609.01361v12026Towards Human-Level Bimanual Dexterous Manipulation with Reinforcement Learning
Yuanpei Chen, Tianhao Wu, Shengjie Wang +8
cs.ROcs.AIcs.LGarXiv:2206.08686v22022An Empirical Study of Mamba-based Language Models
Roger Waleffe, Wonmin Byeon, Duncan Riach +13
cs.LGcs.CLarXiv:2406.07887v12024A Physics-Informed Deep Learning Paradigm for Car-Following Models
Zhaobin Mo, Xuan Di, Rongye Shi
cs.LGeess.SParXiv:2012.13376v42020GazeRefine: Expert Gaze as a Test-Time Prompt for Training-Free Medical Image Segmentation
Mohammed Oussama Benyahia, Marouane Tliba, Mohamed Amine Kerkouri +10
eess.IVcs.AIcs.CVarXiv:2609.01310v12026Safe Exploration in Finite Markov Decision Processes with Gaussian Processes
Matteo Turchetta, Felix Berkenkamp, Andreas Krause
cs.LGcs.AIcs.ROarXiv:1606.04753v22016Topological Steering
Benoît Guérand, Tan Minh Nguyen
cs.LGarXiv:2609.00597v12026Boltzmann Exploration Done Right
Nicolò Cesa-Bianchi, Claudio Gentile, Gábor Lugosi +1
cs.LGstat.MLarXiv:1705.10257v22017Robust Unsupervised Video Anomaly Detection by Multi-Path Frame Prediction
Xuanzhao Wang, Zhengping Che, Bo Jiang +6
cs.CVcs.LGarXiv:2011.02763v22020Learning to select data for transfer learning with Bayesian Optimization
Sebastian Ruder, Barbara Plank
cs.CLcs.LGarXiv:1707.05246v12017Deep Learning Inference in Facebook Data Centers: Characterization, Performance Optimizations and Hardware Implications
Jongsoo Park, Maxim Naumov, Protonu Basu +25
cs.LGstat.MLarXiv:1811.09886v22018TinyLLaVA: A Framework of Small-scale Large Multimodal Models
Baichuan Zhou, Ying Hu, Xi Weng +5
cs.LGcs.CLarXiv:2402.14289v12024Contraction Theory for Nonlinear Stability Analysis and Learning-based Control: A Tutorial Overview
Hiroyasu Tsukamoto, Soon-Jo Chung, Jean-Jacques E. Slotine
cs.LGcs.ROeess.SYarXiv:2110.00675v82021Class-Incremental Continual Learning into the eXtended DER-verse
Matteo Boschini, Lorenzo Bonicelli, Pietro Buzzega +2
cs.LGstat.MLarXiv:2201.00766v22022Large Language Diffusion Models
Shen Nie, Fengqi Zhu, Zebin You +7
cs.CLcs.LGarXiv:2502.09992v32025Kimi k1.5: Scaling Reinforcement Learning with LLMs
Kimi Team, Angang Du, Bofei Gao +93
cs.AIcs.LGarXiv:2501.12599v42025Private Learning and Sanitization: Pure vs. Approximate Differential Privacy
Amos Beimel, Kobbi Nissim, Uri Stemmer
cs.LGcs.CRstat.MLarXiv:1407.2674v12014Noise Flow: Noise Modeling with Conditional Normalizing Flows
Abdelrahman Abdelhamed, Marcus A. Brubaker, Michael S. Brown
cs.CVcs.LGeess.IVarXiv:1908.08453v12019Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts
Lean Wang, Huazuo Gao, Chenggang Zhao +2
cs.LGcs.CLarXiv:2408.15664v12024Summarizing Opinions: Aspect Extraction Meets Sentiment Prediction and They Are Both Weakly Supervised
Stefanos Angelidis, Mirella Lapata
cs.CLcs.AIcs.LGarXiv:1808.08858v12018Convolutional Neural Network Pruning with Structural Redundancy Reduction
Zi Wang, Chengcheng Li, Xiangyang Wang
cs.CVcs.LGarXiv:2104.03438v12021Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
Xun Huang, Zhengqi Li, Guande He +2
cs.CVcs.AIcs.LGarXiv:2506.08009v22025Graph-based, Self-Supervised Program Repair from Diagnostic Feedback
Michihiro Yasunaga, Percy Liang
cs.SEcs.CLcs.LGarXiv:2005.10636v22020Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
Moo Jin Kim, Chelsea Finn, Percy Liang
cs.ROcs.AIcs.CVarXiv:2502.19645v22025Neighborhood Contrastive Learning for Novel Class Discovery
Zhun Zhong, Enrico Fini, Subhankar Roy +3
cs.CVcs.AIcs.LGarXiv:2106.10731v12021Efficient Low Rank Tensor Ring Completion
Wenqi Wang, Vaneet Aggarwal, Shuchin Aeron
cs.LGcs.ITarXiv:1707.08184v12017Lightweight Probabilistic Deep Networks
Jochen Gast, Stefan Roth
cs.CVcs.LGstat.MLarXiv:1805.11327v12018GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
NVIDIA, :, Johan Bjorck +40
cs.ROcs.AIcs.LGarXiv:2503.14734v22025Understanding R1-Zero-Like Training: A Critical Perspective
Zichen Liu, Changyu Chen, Wenjun Li +5
cs.LGcs.AIcs.CLarXiv:2503.20783v22025MAST: A Memory-Augmented Self-supervised Tracker
Zihang Lai, Erika Lu, Weidi Xie
cs.CVcs.LGarXiv:2002.07793v22020FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations
Ziyao Wang, Zheyu Shen, Yexiao He +4
cs.LGcs.DCarXiv:2409.05976v12024Deep Metric Learning for Few-Shot Image Classification: A Review of Recent Developments
Xiaoxu Li, Xiaochen Yang, Zhanyu Ma +1
cs.CVcs.LGarXiv:2105.08149v22021Pooling and Drift in Delayed Bandits
Melika Baghi
stat.MLcs.LGarXiv:2609.01761v12026OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
Shubham Toshniwal, Wei Du, Ivan Moshkov +3
cs.CLcs.AIcs.LGarXiv:2410.01560v22024FurnitureBench: Reproducible Real-World Benchmark for Long-Horizon Complex Manipulation
Minho Heo, Youngwoon Lee, Doohyun Lee +1
cs.ROcs.AIcs.LGarXiv:2305.12821v12023Policy Finetuning: Bridging Sample-Efficient Offline and Online Reinforcement Learning
Tengyang Xie, Nan Jiang, Huan Wang +2
cs.LGstat.MLarXiv:2106.04895v22021MACER: Attack-free and Scalable Robust Training via Maximizing Certified Radius
Runtian Zhai, Chen Dan, Di He +5
cs.LGcs.CRstat.MLarXiv:2001.02378v42020Manifold Elastic Net: A Unified Framework for Sparse Dimension Reduction
Tianyi Zhou, Dacheng Tao, Xindong Wu
cs.LGstat.MLarXiv:1007.3564v32010Molecular graph generation with Graph Neural Networks
Pietro Bongini, Monica Bianchini, Franco Scarselli
stat.MLcs.LGq-bio.BMarXiv:2012.07397v22020A Review on Methods and Applications in Multimodal Deep Learning
Jabeen Summaira, Xi Li, Amin Muhammad Shoib +1
cs.LGcs.MMarXiv:2202.09195v12022Illuminating Generalization in Deep Reinforcement Learning through Procedural Level Generation
Niels Justesen, Ruben Rodriguez Torrado, Philip Bontrager +3
cs.LGcs.AIstat.MLarXiv:1806.10729v52018