Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

8,641 to 8,700 of 20,454

  1. Online Non-Monotone DR-Submodular Maximization Matching the Offline $0.401$ Factor

    Vaneet Aggarwal, Yiyang Lu

    cs.LGcs.AIcs.CCarXiv:2609.02145v12026
  2. Open-ended Learning in Symmetric Zero-sum Games

    David Balduzzi, Marta Garnelo, Yoram Bachrach +4

    cs.LGcs.GTcs.MAarXiv:1901.08106v22019
  3. Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

    Siyan Zhao, Zhihui Xie, Mengchen Liu +4

    cs.LGcs.CLarXiv:2601.18734v32026
  4. Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems

    Jinxi Yu, Yubei Li, Eric Hanchen Jiang +6

    cs.AIcs.LGcs.MAarXiv:2609.02264v12026
  5. Prediction Poisoning: Towards Defenses Against DNN Model Stealing Attacks

    Tribhuvanesh Orekondy, Bernt Schiele, Mario Fritz

    cs.LGcs.CRcs.CVarXiv:1906.10908v22019
  6. SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

    Tianzhe Chu, Yuexiang Zhai, Jihan Yang +6

    cs.AIcs.CVcs.LGarXiv:2501.17161v22025
  7. Motion-Aware Feature for Improved Video Anomaly Detection

    Yi Zhu, Shawn Newsam

    cs.CVcs.LGeess.IVarXiv:1907.10211v12019
  8. Mean Flows for One-step Generative Modeling

    Zhengyang Geng, Mingyang Deng, Xingjian Bai +2

    cs.LGcs.CVarXiv:2505.13447v12025
  9. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

    Mido Assran, Adrien Bardes, David Fan +27

    cs.AIcs.CVcs.LGarXiv:2506.09985v12025
  10. Group Sequence Policy Optimization

    Chujie Zheng, Shixuan Liu, Mingze Li +9

    cs.LGcs.AIcs.CLarXiv:2507.18071v22025
  11. Self-Training: A Survey

    Massih-Reza Amini, Vasilii Feofanov, Loic Pauletto +3

    cs.LGarXiv:2202.12040v62022
  12. Human-AI Collaboration via Conditional Delegation: A Case Study of Content Moderation

    Vivian Lai, Samuel Carton, Rajat Bhatnagar +3

    cs.AIcs.HCcs.LGarXiv:2204.11788v12022
  13. Abnormal respiratory patterns classifier may contribute to large-scale screening of people infected with COVID-19 in an accurate and unobtrusive manner

    Yunlu Wang, Menghan Hu, Qingli Li +3

    cs.LGcs.CVeess.SParXiv:2002.05534v22020
  14. Conditional Generative Neural System for Probabilistic Trajectory Prediction

    Jiachen Li, Hengbo Ma, Masayoshi Tomizuka

    cs.CVcs.AIcs.LGarXiv:1905.01631v22019
  15. Generative Adversarial Active Learning

    Jia-Jie Zhu, José Bento

    cs.LGstat.MLarXiv:1702.07956v52017
  16. Beyond Blur: A Semantic Tri-view Pipeline for Teledermatology Gradability via Skin Micro-relief

    Robert Engel

    eess.IVcs.CVcs.HCarXiv:2609.03095v12026
  17. Towards Interpretable Semantic Segmentation via Gradient-weighted Class Activation Mapping

    Kira Vinogradova, Alexandr Dibrov, Gene Myers

    cs.CVcs.LGeess.IVarXiv:2002.11434v12020
  18. Supervised Classification Performance of Multispectral Images

    K. Perumal, R. Bhaskaran

    cs.LGcs.CVarXiv:1002.4046v12010
  19. From Euclidean to Graph-Structured Data: A Survey of Collaborative Learning

    Rémi Bourgerie, Šarūnas Girdzijauskas, Viktoria Fodor

    cs.LGcs.MAcs.SIarXiv:2609.02984v12026
  20. Scene Parsing with Multiscale Feature Learning, Purity Trees, and Optimal Covers

    Clément Farabet, Camille Couprie, Laurent Najman +1

    cs.CVcs.LGarXiv:1202.2160v22012
  21. Automated Speed and Lane Change Decision Making using Deep Reinforcement Learning

    Carl-Johan Hoel, Krister Wolff, Leo Laine

    cs.ROcs.AIcs.LGarXiv:1803.10056v22018
  22. Linear model predictive safety certification for learning-based control

    Kim P. Wabersich, Melanie N. Zeilinger

    eess.SYcs.LGarXiv:1803.08552v62018
  23. On the Interaction Between Model Compression and Test-Time Adaptation

    Francesco Corti, Dong Wang, Young D. Kwon +2

    cs.LGcs.AIarXiv:2609.03604v12026
  24. Learning efficient sparse and low rank models

    Pablo Sprechmann, Alex M. Bronstein, Guillermo Sapiro

    cs.LGarXiv:1212.3631v12012
  25. Investigating Linear Probe Robustness to Linguistic Register, Medical Specialty, and Corpus Shifts in Medical QA

    Nishant Mishra, Ameen Abu-Hanna, Iacer Calixto

    cs.CLcs.LGarXiv:2609.01361v12026
  26. Towards Human-Level Bimanual Dexterous Manipulation with Reinforcement Learning

    Yuanpei Chen, Tianhao Wu, Shengjie Wang +8

    cs.ROcs.AIcs.LGarXiv:2206.08686v22022
  27. An Empirical Study of Mamba-based Language Models

    Roger Waleffe, Wonmin Byeon, Duncan Riach +13

    cs.LGcs.CLarXiv:2406.07887v12024
  28. A Physics-Informed Deep Learning Paradigm for Car-Following Models

    Zhaobin Mo, Xuan Di, Rongye Shi

    cs.LGeess.SParXiv:2012.13376v42020
  29. GazeRefine: Expert Gaze as a Test-Time Prompt for Training-Free Medical Image Segmentation

    Mohammed Oussama Benyahia, Marouane Tliba, Mohamed Amine Kerkouri +10

    eess.IVcs.AIcs.CVarXiv:2609.01310v12026
  30. Safe Exploration in Finite Markov Decision Processes with Gaussian Processes

    Matteo Turchetta, Felix Berkenkamp, Andreas Krause

    cs.LGcs.AIcs.ROarXiv:1606.04753v22016
  31. Topological Steering

    Benoît Guérand, Tan Minh Nguyen

    cs.LGarXiv:2609.00597v12026
  32. Boltzmann Exploration Done Right

    Nicolò Cesa-Bianchi, Claudio Gentile, Gábor Lugosi +1

    cs.LGstat.MLarXiv:1705.10257v22017
  33. Robust Unsupervised Video Anomaly Detection by Multi-Path Frame Prediction

    Xuanzhao Wang, Zhengping Che, Bo Jiang +6

    cs.CVcs.LGarXiv:2011.02763v22020
  34. Learning to select data for transfer learning with Bayesian Optimization

    Sebastian Ruder, Barbara Plank

    cs.CLcs.LGarXiv:1707.05246v12017
  35. Deep Learning Inference in Facebook Data Centers: Characterization, Performance Optimizations and Hardware Implications

    Jongsoo Park, Maxim Naumov, Protonu Basu +25

    cs.LGstat.MLarXiv:1811.09886v22018
  36. TinyLLaVA: A Framework of Small-scale Large Multimodal Models

    Baichuan Zhou, Ying Hu, Xi Weng +5

    cs.LGcs.CLarXiv:2402.14289v12024
  37. Contraction Theory for Nonlinear Stability Analysis and Learning-based Control: A Tutorial Overview

    Hiroyasu Tsukamoto, Soon-Jo Chung, Jean-Jacques E. Slotine

    cs.LGcs.ROeess.SYarXiv:2110.00675v82021
  38. Class-Incremental Continual Learning into the eXtended DER-verse

    Matteo Boschini, Lorenzo Bonicelli, Pietro Buzzega +2

    cs.LGstat.MLarXiv:2201.00766v22022
  39. Large Language Diffusion Models

    Shen Nie, Fengqi Zhu, Zebin You +7

    cs.CLcs.LGarXiv:2502.09992v32025
  40. Kimi k1.5: Scaling Reinforcement Learning with LLMs

    Kimi Team, Angang Du, Bofei Gao +93

    cs.AIcs.LGarXiv:2501.12599v42025
  41. Private Learning and Sanitization: Pure vs. Approximate Differential Privacy

    Amos Beimel, Kobbi Nissim, Uri Stemmer

    cs.LGcs.CRstat.MLarXiv:1407.2674v12014
  42. Noise Flow: Noise Modeling with Conditional Normalizing Flows

    Abdelrahman Abdelhamed, Marcus A. Brubaker, Michael S. Brown

    cs.CVcs.LGeess.IVarXiv:1908.08453v12019
  43. Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts

    Lean Wang, Huazuo Gao, Chenggang Zhao +2

    cs.LGcs.CLarXiv:2408.15664v12024
  44. Summarizing Opinions: Aspect Extraction Meets Sentiment Prediction and They Are Both Weakly Supervised

    Stefanos Angelidis, Mirella Lapata

    cs.CLcs.AIcs.LGarXiv:1808.08858v12018
  45. Convolutional Neural Network Pruning with Structural Redundancy Reduction

    Zi Wang, Chengcheng Li, Xiangyang Wang

    cs.CVcs.LGarXiv:2104.03438v12021
  46. Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion

    Xun Huang, Zhengqi Li, Guande He +2

    cs.CVcs.AIcs.LGarXiv:2506.08009v22025
  47. Graph-based, Self-Supervised Program Repair from Diagnostic Feedback

    Michihiro Yasunaga, Percy Liang

    cs.SEcs.CLcs.LGarXiv:2005.10636v22020
  48. Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

    Moo Jin Kim, Chelsea Finn, Percy Liang

    cs.ROcs.AIcs.CVarXiv:2502.19645v22025
  49. Neighborhood Contrastive Learning for Novel Class Discovery

    Zhun Zhong, Enrico Fini, Subhankar Roy +3

    cs.CVcs.AIcs.LGarXiv:2106.10731v12021
  50. Efficient Low Rank Tensor Ring Completion

    Wenqi Wang, Vaneet Aggarwal, Shuchin Aeron

    cs.LGcs.ITarXiv:1707.08184v12017
  51. Lightweight Probabilistic Deep Networks

    Jochen Gast, Stefan Roth

    cs.CVcs.LGstat.MLarXiv:1805.11327v12018
  52. GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

    NVIDIA, :, Johan Bjorck +40

    cs.ROcs.AIcs.LGarXiv:2503.14734v22025
  53. Understanding R1-Zero-Like Training: A Critical Perspective

    Zichen Liu, Changyu Chen, Wenjun Li +5

    cs.LGcs.AIcs.CLarXiv:2503.20783v22025
  54. MAST: A Memory-Augmented Self-supervised Tracker

    Zihang Lai, Erika Lu, Weidi Xie

    cs.CVcs.LGarXiv:2002.07793v22020
  55. FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations

    Ziyao Wang, Zheyu Shen, Yexiao He +4

    cs.LGcs.DCarXiv:2409.05976v12024
  56. Deep Metric Learning for Few-Shot Image Classification: A Review of Recent Developments

    Xiaoxu Li, Xiaochen Yang, Zhanyu Ma +1

    cs.CVcs.LGarXiv:2105.08149v22021
  57. Pooling and Drift in Delayed Bandits

    Melika Baghi

    stat.MLcs.LGarXiv:2609.01761v12026
  58. OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data

    Shubham Toshniwal, Wei Du, Ivan Moshkov +3

    cs.CLcs.AIcs.LGarXiv:2410.01560v22024
  59. FurnitureBench: Reproducible Real-World Benchmark for Long-Horizon Complex Manipulation

    Minho Heo, Youngwoon Lee, Doohyun Lee +1

    cs.ROcs.AIcs.LGarXiv:2305.12821v12023
  60. Policy Finetuning: Bridging Sample-Efficient Offline and Online Reinforcement Learning

    Tengyang Xie, Nan Jiang, Huan Wang +2

    cs.LGstat.MLarXiv:2106.04895v22021