Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
6,001 to 6,060 of 20,454
Cartridges: Lightweight and general-purpose long context representations via self-study
Sabri Eyuboglu, Ryan Ehrlich, Simran Arora +8
cs.CLcs.AIcs.LGarXiv:2506.06266v32025The Invisible Leash: Why RLVR May or May Not Escape Its Origin
Fang Wu, Weihao Xuan, Ximing Lu +4
cs.LGcs.AIcs.CLarXiv:2507.14843v42025Linear Convergence in Federated Learning: Tackling Client Heterogeneity and Sparse Gradients
Aritra Mitra, Rayana Jaafar, George J. Pappas +1
cs.LGcs.DCeess.SYarXiv:2102.07053v22021Design Patterns for Securing LLM Agents against Prompt Injections
Luca Beurer-Kellner, Beat Buesser, Ana-Maria Creţu +11
cs.LGcs.CRarXiv:2506.08837v32025Towards Understanding Camera Motions in Any Video
Zhiqiu Lin, Siyuan Cen, Daniel Jiang +12
cs.CVcs.AIcs.CLarXiv:2504.15376v22025Urban Driver: Learning to Drive from Real-world Demonstrations Using Policy Gradients
Oliver Scheel, Luca Bergamini, Maciej Wołczyk +2
cs.ROcs.AIcs.CVarXiv:2109.13333v12021Efficient Online Reinforcement Learning for Diffusion Policy
Haitong Ma, Tianyi Chen, Kai Wang +2
cs.LGarXiv:2502.00361v42025Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
Tianbao Xie, Siheng Zhao, Chen Henry Wu +5
cs.LGcs.AIcs.CLarXiv:2309.11489v32023A feature agnostic approach for glaucoma detection in OCT volumes
Stefan Maetschke, Bhavna Antony, Hiroshi Ishikawa +3
cs.CVcs.LGstat.MLarXiv:1807.04855v42018STEm-Seg: Spatio-temporal Embeddings for Instance Segmentation in Videos
Ali Athar, Sabarinath Mahadevan, Aljoša Ošep +2
cs.CVcs.LGeess.IVarXiv:2003.08429v42020Parametrized quantum policies for reinforcement learning
Sofiene Jerbi, Casper Gyurik, Simon C. Marshall +2
quant-phcs.AIcs.LGarXiv:2103.05577v22021MedRAX: Medical Reasoning Agent for Chest X-ray
Adibvafa Fallahpour, Jun Ma, Alif Munim +2
cs.LGcs.AIcs.MAarXiv:2502.02673v22025TimeFilter: Patch-Specific Spatial-Temporal Graph Filtration for Time Series Forecasting
Yifan Hu, Guibin Zhang, Peiyuan Liu +6
cs.LGarXiv:2501.13041v22025Character-Level Question Answering with Attention
David Golub, Xiaodong He
cs.CLcs.AIcs.LGarXiv:1604.00727v42016Dense Extreme Inception Network for Edge Detection
Xavier Soria, Angel Sappa, Patricio Humanante +1
cs.CVcs.LGarXiv:2112.02250v22021Outcome-based Exploration for LLM Reasoning
Yuda Song, Julia Kempe, Remi Munos
cs.LGcs.CLarXiv:2509.06941v12025Data Banzhaf: A Robust Data Valuation Framework for Machine Learning
Jiachen T. Wang, Ruoxi Jia
cs.LGcs.GTstat.MLarXiv:2205.15466v72022Generating 3D Molecules for Target Protein Binding
Meng Liu, Youzhi Luo, Kanji Uchino +2
q-bio.BMcs.AIcs.LGarXiv:2204.09410v22022SS-ESOAP: Self-Scaled Adaptive Preconditioning for Physics-Informed Learning
Guangyuan Wang, Mads Toftrup, Sebastian Loeschcke +2
cs.LGcs.AImath.OCarXiv:2608.29448v12026AdaMatch: A Unified Approach to Semi-Supervised Learning and Domain Adaptation
David Berthelot, Rebecca Roelofs, Kihyuk Sohn +2
cs.LGcs.AIcs.CVarXiv:2106.04732v22021Position Focused Attention Network for Image-Text Matching
Yaxiong Wang, Hao Yang, Xueming Qian +4
cs.CLcs.IRcs.LGarXiv:1907.09748v12019EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis
Xiaoshuai Song, Haofei Chang, Guanting Dong +3
cs.CLcs.AIcs.LGarXiv:2601.05808v22026Self-Adaptive Hierarchical Sentence Model
Han Zhao, Zhengdong Lu, Pascal Poupart
cs.CLcs.LGcs.NEarXiv:1504.05070v22015Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series Forecasting
Siru Zhong, Weilin Ruan, Ming Jin +3
cs.CVcs.LGarXiv:2502.04395v22025Reducing Overestimation Bias in Multi-Agent Domains Using Double Centralized Critics
Johannes Ackermann, Volker Gabler, Takayuki Osa +1
cs.LGcs.AIcs.MAarXiv:1910.01465v22019Temporal Query Network for Efficient Multivariate Time Series Forecasting
Shengsheng Lin, Haojun Chen, Haijie Wu +2
cs.LGarXiv:2505.12917v22025On the convergence of single-call stochastic extra-gradient methods
Yu-Guan Hsieh, Franck Iutzeler, Jérôme Malick +1
math.OCcs.GTcs.LGarXiv:1908.08465v22019Are DeepSeek R1 And Other Reasoning Models More Faithful?
James Chua, Owain Evans
cs.LGarXiv:2501.08156v52025Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought
Hanlin Zhu, Shibo Hao, Zhiting Hu +3
cs.LGarXiv:2505.12514v32025Instance Credibility Inference for Few-Shot Learning
Yikai Wang, Chengming Xu, Chen Liu +2
cs.CVcs.LGarXiv:2003.11853v22020AgentProv: Auditing Agentic LLM API Providers via Tool-use Policy Probes
Xun Wang, Bihe Zhao, Michael Backes +2
cs.CRcs.CLcs.LGarXiv:2609.00052v12026Regularizing Generative Adversarial Networks under Limited Data
Hung-Yu Tseng, Lu Jiang, Ce Liu +2
cs.LGcs.CVarXiv:2104.03310v12021Textless Speech-to-Speech Translation on Real Data
Ann Lee, Hongyu Gong, Paul-Ambroise Duquenne +8
cs.CLcs.AIcs.LGarXiv:2112.08352v22021A mixed formulation for physics-informed neural networks as a potential solver for engineering problems in heterogeneous domains: comparison with finite element method
Shahed Rezaei, Ali Harandi, Ahmad Moeineddin +2
cs.CEcs.LGarXiv:2206.13103v12022Deep Learning with Label Differential Privacy
Badih Ghazi, Noah Golowich, Ravi Kumar +2
cs.LGcs.DSarXiv:2102.06062v22021Fast Algorithms for Online Stochastic Convex Programming
Shipra Agrawal, Nikhil R. Devanur
cs.LGcs.DSmath.OCarXiv:1410.7596v12014Leveraging the Feature Distribution in Transfer-based Few-Shot Learning
Yuqing Hu, Vincent Gripon, Stéphane Pateux
cs.LGstat.MLarXiv:2006.03806v32020Scientific Machine Learning Benchmarks
Jeyan Thiyagalingam, Mallikarjun Shankar, Geoffrey Fox +1
cs.LGphysics.comp-pharXiv:2110.12773v12021Radiological images and machine learning: trends, perspectives, and prospects
Zhenwei Zhang, Ervin Sejdic
eess.IVcs.LGarXiv:1903.11726v12019Do LLMs Know Your Neighborhood? Auditing LLM Priors for Neighborhood-Level Mobility Prediction and Structural Alignment
Saad Mohammad Abrar, Eesha Kurella, Arnav Dadarya +3
cs.LGcs.CYarXiv:2609.00345v12026Global disease monitoring and forecasting with Wikipedia
Nicholas Generous, Geoffrey Fairchild, Alina Deshpande +2
cs.SIcs.LGphysics.soc-pharXiv:1405.3612v22014Deep Learning for Source Code Modeling and Generation: Models, Applications and Challenges
Triet H. M. Le, Hao Chen, M. Ali Babar
cs.SEcs.AIcs.LGarXiv:2002.05442v12020Near Optimal Behavior via Approximate State Abstraction
David Abel, D. Ellis Hershkowitz, Michael L. Littman
cs.LGcs.AIarXiv:1701.04113v12017Reinforcing General Reasoning without Verifiers
Xiangxin Zhou, Zichen Liu, Anya Sims +6
cs.LGcs.CLarXiv:2505.21493v12025Iterative Normalization: Beyond Standardization towards Efficient Whitening
Lei Huang, Yi Zhou, Fan Zhu +2
cs.CVcs.LGarXiv:1904.03441v12019Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement Learning
Haozhen Zhang, Tao Feng, Jiaxuan You
cs.CLcs.AIcs.LGarXiv:2506.09033v32025Whoever Started the Interference Should End It: Guiding Data-Free Model Merging via Task Vectors
Runxi Cheng, Feng Xiong, Yongxian Wei +2
cs.LGarXiv:2503.08099v22025Convergence of Gradient Descent on Separable Data
Mor Shpigel Nacson, Jason D. Lee, Suriya Gunasekar +3
stat.MLcs.LGarXiv:1803.01905v32018Rethinking embedding coupling in pre-trained language models
Hyung Won Chung, Thibault Févry, Henry Tsai +2
cs.CLcs.LGarXiv:2010.12821v12020NeuroPriv: Adversarial Representation Learning for Privacy in Wearable EEG Systems
Sarmistha Sarna Gomasta, Bhawana Chhaglani, Prashant Shenoy
cs.CRcs.HCcs.LGarXiv:2609.00390v12026SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models
Thinh Pham, Nguyen Nguyen, Pratibha Zunjare +3
cs.CLcs.AIcs.LGarXiv:2506.01062v42025Incorporating Domain Knowledge into Deep Neural Networks
Tirtharaj Dash, Sharad Chitlangia, Aditya Ahuja +1
cs.NEcs.AIcs.LGarXiv:2103.00180v22021Active Learning on a Budget: Opposite Strategies Suit High and Low Budgets
Guy Hacohen, Avihu Dekel, Daphna Weinshall
cs.LGarXiv:2202.02794v42022Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
NVIDIA, :, Aaron Blakeman +198
cs.CLcs.AIcs.LGarXiv:2504.03624v42025Neural Prompt Search
Yuanhan Zhang, Kaiyang Zhou, Ziwei Liu
cs.CVcs.AIcs.LGarXiv:2206.04673v22022Multi-fidelity Bayesian Neural Networks: Algorithms and Applications
Xuhui Meng, Hessam Babaee, George Em Karniadakis
cs.LGphysics.comp-pharXiv:2012.13294v12020Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't
Quy-Anh Dang, Chris Ngo
cs.LGcs.CLarXiv:2503.16219v22025Benchmarking Reinforcement Learning Algorithms on Real-World Robots
A. Rupam Mahmood, Dmytro Korenkevych, Gautham Vasan +2
cs.LGcs.AIcs.ROarXiv:1809.07731v12018DISTAL: Distillation and Self-Supervised Pretraining for Structure-Agnostic Materials Property Prediction
Weiran Wang, Xintong Huo, Yueying Wang +7
cs.LGcs.AIcs.ETarXiv:2609.00059v12026Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection
Bertie Vidgen, Tristan Thrush, Zeerak Waseem +1
cs.CLcs.LGarXiv:2012.15761v22020