Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

6,001 to 6,060 of 20,454

  1. Cartridges: Lightweight and general-purpose long context representations via self-study

    Sabri Eyuboglu, Ryan Ehrlich, Simran Arora +8

    cs.CLcs.AIcs.LGarXiv:2506.06266v32025
  2. The Invisible Leash: Why RLVR May or May Not Escape Its Origin

    Fang Wu, Weihao Xuan, Ximing Lu +4

    cs.LGcs.AIcs.CLarXiv:2507.14843v42025
  3. Linear Convergence in Federated Learning: Tackling Client Heterogeneity and Sparse Gradients

    Aritra Mitra, Rayana Jaafar, George J. Pappas +1

    cs.LGcs.DCeess.SYarXiv:2102.07053v22021
  4. Design Patterns for Securing LLM Agents against Prompt Injections

    Luca Beurer-Kellner, Beat Buesser, Ana-Maria Creţu +11

    cs.LGcs.CRarXiv:2506.08837v32025
  5. Towards Understanding Camera Motions in Any Video

    Zhiqiu Lin, Siyuan Cen, Daniel Jiang +12

    cs.CVcs.AIcs.CLarXiv:2504.15376v22025
  6. Urban Driver: Learning to Drive from Real-world Demonstrations Using Policy Gradients

    Oliver Scheel, Luca Bergamini, Maciej Wołczyk +2

    cs.ROcs.AIcs.CVarXiv:2109.13333v12021
  7. Efficient Online Reinforcement Learning for Diffusion Policy

    Haitong Ma, Tianyi Chen, Kai Wang +2

    cs.LGarXiv:2502.00361v42025
  8. Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

    Tianbao Xie, Siheng Zhao, Chen Henry Wu +5

    cs.LGcs.AIcs.CLarXiv:2309.11489v32023
  9. A feature agnostic approach for glaucoma detection in OCT volumes

    Stefan Maetschke, Bhavna Antony, Hiroshi Ishikawa +3

    cs.CVcs.LGstat.MLarXiv:1807.04855v42018
  10. STEm-Seg: Spatio-temporal Embeddings for Instance Segmentation in Videos

    Ali Athar, Sabarinath Mahadevan, Aljoša Ošep +2

    cs.CVcs.LGeess.IVarXiv:2003.08429v42020
  11. Parametrized quantum policies for reinforcement learning

    Sofiene Jerbi, Casper Gyurik, Simon C. Marshall +2

    quant-phcs.AIcs.LGarXiv:2103.05577v22021
  12. MedRAX: Medical Reasoning Agent for Chest X-ray

    Adibvafa Fallahpour, Jun Ma, Alif Munim +2

    cs.LGcs.AIcs.MAarXiv:2502.02673v22025
  13. TimeFilter: Patch-Specific Spatial-Temporal Graph Filtration for Time Series Forecasting

    Yifan Hu, Guibin Zhang, Peiyuan Liu +6

    cs.LGarXiv:2501.13041v22025
  14. Character-Level Question Answering with Attention

    David Golub, Xiaodong He

    cs.CLcs.AIcs.LGarXiv:1604.00727v42016
  15. Dense Extreme Inception Network for Edge Detection

    Xavier Soria, Angel Sappa, Patricio Humanante +1

    cs.CVcs.LGarXiv:2112.02250v22021
  16. Outcome-based Exploration for LLM Reasoning

    Yuda Song, Julia Kempe, Remi Munos

    cs.LGcs.CLarXiv:2509.06941v12025
  17. Data Banzhaf: A Robust Data Valuation Framework for Machine Learning

    Jiachen T. Wang, Ruoxi Jia

    cs.LGcs.GTstat.MLarXiv:2205.15466v72022
  18. Generating 3D Molecules for Target Protein Binding

    Meng Liu, Youzhi Luo, Kanji Uchino +2

    q-bio.BMcs.AIcs.LGarXiv:2204.09410v22022
  19. SS-ESOAP: Self-Scaled Adaptive Preconditioning for Physics-Informed Learning

    Guangyuan Wang, Mads Toftrup, Sebastian Loeschcke +2

    cs.LGcs.AImath.OCarXiv:2608.29448v12026
  20. AdaMatch: A Unified Approach to Semi-Supervised Learning and Domain Adaptation

    David Berthelot, Rebecca Roelofs, Kihyuk Sohn +2

    cs.LGcs.AIcs.CVarXiv:2106.04732v22021
  21. Position Focused Attention Network for Image-Text Matching

    Yaxiong Wang, Hao Yang, Xueming Qian +4

    cs.CLcs.IRcs.LGarXiv:1907.09748v12019
  22. EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis

    Xiaoshuai Song, Haofei Chang, Guanting Dong +3

    cs.CLcs.AIcs.LGarXiv:2601.05808v22026
  23. Self-Adaptive Hierarchical Sentence Model

    Han Zhao, Zhengdong Lu, Pascal Poupart

    cs.CLcs.LGcs.NEarXiv:1504.05070v22015
  24. Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series Forecasting

    Siru Zhong, Weilin Ruan, Ming Jin +3

    cs.CVcs.LGarXiv:2502.04395v22025
  25. Reducing Overestimation Bias in Multi-Agent Domains Using Double Centralized Critics

    Johannes Ackermann, Volker Gabler, Takayuki Osa +1

    cs.LGcs.AIcs.MAarXiv:1910.01465v22019
  26. Temporal Query Network for Efficient Multivariate Time Series Forecasting

    Shengsheng Lin, Haojun Chen, Haijie Wu +2

    cs.LGarXiv:2505.12917v22025
  27. On the convergence of single-call stochastic extra-gradient methods

    Yu-Guan Hsieh, Franck Iutzeler, Jérôme Malick +1

    math.OCcs.GTcs.LGarXiv:1908.08465v22019
  28. Are DeepSeek R1 And Other Reasoning Models More Faithful?

    James Chua, Owain Evans

    cs.LGarXiv:2501.08156v52025
  29. Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought

    Hanlin Zhu, Shibo Hao, Zhiting Hu +3

    cs.LGarXiv:2505.12514v32025
  30. Instance Credibility Inference for Few-Shot Learning

    Yikai Wang, Chengming Xu, Chen Liu +2

    cs.CVcs.LGarXiv:2003.11853v22020
  31. AgentProv: Auditing Agentic LLM API Providers via Tool-use Policy Probes

    Xun Wang, Bihe Zhao, Michael Backes +2

    cs.CRcs.CLcs.LGarXiv:2609.00052v12026
  32. Regularizing Generative Adversarial Networks under Limited Data

    Hung-Yu Tseng, Lu Jiang, Ce Liu +2

    cs.LGcs.CVarXiv:2104.03310v12021
  33. Textless Speech-to-Speech Translation on Real Data

    Ann Lee, Hongyu Gong, Paul-Ambroise Duquenne +8

    cs.CLcs.AIcs.LGarXiv:2112.08352v22021
  34. A mixed formulation for physics-informed neural networks as a potential solver for engineering problems in heterogeneous domains: comparison with finite element method

    Shahed Rezaei, Ali Harandi, Ahmad Moeineddin +2

    cs.CEcs.LGarXiv:2206.13103v12022
  35. Deep Learning with Label Differential Privacy

    Badih Ghazi, Noah Golowich, Ravi Kumar +2

    cs.LGcs.DSarXiv:2102.06062v22021
  36. Fast Algorithms for Online Stochastic Convex Programming

    Shipra Agrawal, Nikhil R. Devanur

    cs.LGcs.DSmath.OCarXiv:1410.7596v12014
  37. Leveraging the Feature Distribution in Transfer-based Few-Shot Learning

    Yuqing Hu, Vincent Gripon, Stéphane Pateux

    cs.LGstat.MLarXiv:2006.03806v32020
  38. Scientific Machine Learning Benchmarks

    Jeyan Thiyagalingam, Mallikarjun Shankar, Geoffrey Fox +1

    cs.LGphysics.comp-pharXiv:2110.12773v12021
  39. Radiological images and machine learning: trends, perspectives, and prospects

    Zhenwei Zhang, Ervin Sejdic

    eess.IVcs.LGarXiv:1903.11726v12019
  40. Do LLMs Know Your Neighborhood? Auditing LLM Priors for Neighborhood-Level Mobility Prediction and Structural Alignment

    Saad Mohammad Abrar, Eesha Kurella, Arnav Dadarya +3

    cs.LGcs.CYarXiv:2609.00345v12026
  41. Global disease monitoring and forecasting with Wikipedia

    Nicholas Generous, Geoffrey Fairchild, Alina Deshpande +2

    cs.SIcs.LGphysics.soc-pharXiv:1405.3612v22014
  42. Deep Learning for Source Code Modeling and Generation: Models, Applications and Challenges

    Triet H. M. Le, Hao Chen, M. Ali Babar

    cs.SEcs.AIcs.LGarXiv:2002.05442v12020
  43. Near Optimal Behavior via Approximate State Abstraction

    David Abel, D. Ellis Hershkowitz, Michael L. Littman

    cs.LGcs.AIarXiv:1701.04113v12017
  44. Reinforcing General Reasoning without Verifiers

    Xiangxin Zhou, Zichen Liu, Anya Sims +6

    cs.LGcs.CLarXiv:2505.21493v12025
  45. Iterative Normalization: Beyond Standardization towards Efficient Whitening

    Lei Huang, Yi Zhou, Fan Zhu +2

    cs.CVcs.LGarXiv:1904.03441v12019
  46. Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement Learning

    Haozhen Zhang, Tao Feng, Jiaxuan You

    cs.CLcs.AIcs.LGarXiv:2506.09033v32025
  47. Whoever Started the Interference Should End It: Guiding Data-Free Model Merging via Task Vectors

    Runxi Cheng, Feng Xiong, Yongxian Wei +2

    cs.LGarXiv:2503.08099v22025
  48. Convergence of Gradient Descent on Separable Data

    Mor Shpigel Nacson, Jason D. Lee, Suriya Gunasekar +3

    stat.MLcs.LGarXiv:1803.01905v32018
  49. Rethinking embedding coupling in pre-trained language models

    Hyung Won Chung, Thibault Févry, Henry Tsai +2

    cs.CLcs.LGarXiv:2010.12821v12020
  50. NeuroPriv: Adversarial Representation Learning for Privacy in Wearable EEG Systems

    Sarmistha Sarna Gomasta, Bhawana Chhaglani, Prashant Shenoy

    cs.CRcs.HCcs.LGarXiv:2609.00390v12026
  51. SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models

    Thinh Pham, Nguyen Nguyen, Pratibha Zunjare +3

    cs.CLcs.AIcs.LGarXiv:2506.01062v42025
  52. Incorporating Domain Knowledge into Deep Neural Networks

    Tirtharaj Dash, Sharad Chitlangia, Aditya Ahuja +1

    cs.NEcs.AIcs.LGarXiv:2103.00180v22021
  53. Active Learning on a Budget: Opposite Strategies Suit High and Low Budgets

    Guy Hacohen, Avihu Dekel, Daphna Weinshall

    cs.LGarXiv:2202.02794v42022
  54. Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models

    NVIDIA, :, Aaron Blakeman +198

    cs.CLcs.AIcs.LGarXiv:2504.03624v42025
  55. Neural Prompt Search

    Yuanhan Zhang, Kaiyang Zhou, Ziwei Liu

    cs.CVcs.AIcs.LGarXiv:2206.04673v22022
  56. Multi-fidelity Bayesian Neural Networks: Algorithms and Applications

    Xuhui Meng, Hessam Babaee, George Em Karniadakis

    cs.LGphysics.comp-pharXiv:2012.13294v12020
  57. Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't

    Quy-Anh Dang, Chris Ngo

    cs.LGcs.CLarXiv:2503.16219v22025
  58. Benchmarking Reinforcement Learning Algorithms on Real-World Robots

    A. Rupam Mahmood, Dmytro Korenkevych, Gautham Vasan +2

    cs.LGcs.AIcs.ROarXiv:1809.07731v12018
  59. DISTAL: Distillation and Self-Supervised Pretraining for Structure-Agnostic Materials Property Prediction

    Weiran Wang, Xintong Huo, Yueying Wang +7

    cs.LGcs.AIcs.ETarXiv:2609.00059v12026
  60. Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection

    Bertie Vidgen, Tristan Thrush, Zeerak Waseem +1

    cs.CLcs.LGarXiv:2012.15761v22020