Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
5,881 to 5,940 of 20,193
Vision-Language-Guided Pseudo-Labels for Unsupervised Domain Adaptation in Semantic Segmentation for Waste Sorting
Udo Schlegel, Shubhangi, Gabriel Dax +3
cs.CVcs.AIcs.LGarXiv:2609.00898v12026Missing Data Imputation using Optimal Transport
Boris Muzellec, Julie Josse, Claire Boyer +1
stat.MLcs.LGarXiv:2002.03860v32020PaSa: An LLM Agent for Comprehensive Academic Paper Search
Yichen He, Guanhua Huang, Peiyuan Feng +4
cs.IRcs.LGarXiv:2501.10120v22025Solving Schrödinger Bridges via Maximum Likelihood
Francisco Vargas, Pierre Thodoroff, Neil D. Lawrence +1
stat.MLcs.LGarXiv:2106.02081v92021FlowDPS: Flow-Driven Posterior Sampling for Inverse Problems
Jeongsol Kim, Bryan Sangwoo Kim, Jong Chul Ye
cs.CVcs.AIcs.LGarXiv:2503.08136v12025Transformers are Graph Neural Networks
Chaitanya K. Joshi
cs.LGcs.AIarXiv:2506.22084v12025Learning with Good Feature Representations in Bandits and in RL with a Generative Model
Tor Lattimore, Csaba Szepesvari, Gellert Weisz
stat.MLcs.LGarXiv:1911.07676v22019AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
Xing Han Lù, Amirhossein Kazemnejad, Nicholas Meade +7
cs.LGcs.AIcs.CLarXiv:2504.08942v22025Real-time Traffic Accident Risk Prediction based on Frequent Pattern Tree
Lei Lin, Qian Wang, Adel W. Sadek
stat.APcs.LGarXiv:1701.05691v22017RL-100: Performant Robotic Manipulation with Real-World Reinforcement Learning
Kun Lei, Huanyu Li, Dongjie Yu +6
cs.ROcs.AIcs.LGarXiv:2510.14830v42025Machine Learning with Knowledge Constraints for Process Optimization of Open-Air Perovskite Solar Cell Manufacturing
Zhe Liu, Nicholas Rolston, Austin C. Flick +4
cs.LGphysics.app-pharXiv:2110.01387v42021Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
Howard Chen, Noam Razin, Karthik Narasimhan +1
cs.LGcs.CLarXiv:2510.18874v32025Learning to Predict the Cosmological Structure Formation
Siyu He, Yin Li, Yu Feng +4
astro-ph.COcs.AIcs.LGarXiv:1811.06533v22018Reasoning-SQL: Reinforcement Learning with SQL Tailored Partial Rewards for Reasoning-Enhanced Text-to-SQL
Mohammadreza Pourreza, Shayan Talaei, Ruoxi Sun +5
cs.LGcs.AIcs.DBarXiv:2503.23157v22025GPT Takes the Bar Exam
Michael Bommarito, Daniel Martin Katz
cs.CLcs.AIcs.LGarXiv:2212.14402v12022From Pixels to Torques: Policy Learning with Deep Dynamical Models
Niklas Wahlström, Thomas B. Schön, Marc Peter Deisenroth
stat.MLcs.LGcs.ROarXiv:1502.02251v32015LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs
Yunhui Xia, Wei Shen, Yan Wang +5
cs.LGcs.CLcs.SEarXiv:2504.14655v12025A PSO and Pattern Search based Memetic Algorithm for SVMs Parameters Optimization
Yukun Bao, Zhongyi Hu, Tao Xiong
cs.LGcs.AIcs.NEarXiv:1401.1926v12014StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
Yinmin Zhong, Zili Zhang, Xiaoniu Song +11
cs.LGcs.DCarXiv:2504.15930v12025CAMEL: A Weakly Supervised Learning Framework for Histopathology Image Segmentation
Gang Xu, Zhigang Song, Zhuo Sun +6
eess.IVcs.CVcs.LGarXiv:1908.10555v12019Quad-networks: unsupervised learning to rank for interest point detection
Nikolay Savinov, Akihito Seki, Lubor Ladicky +2
cs.CVcs.LGcs.NEarXiv:1611.07571v22016TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators
Jianling Li, Shangzhan Li, Zhenye Gao +9
cs.CLcs.LGarXiv:2502.14752v12025Federated Neural Collaborative Filtering
Vasileios Perifanis, Pavlos S. Efraimidis
cs.IRcs.CRcs.LGarXiv:2106.04405v22021Knowledge Graph Embedding: A Survey from the Perspective of Representation Spaces
Jiahang Cao, Jinyuan Fang, Zaiqiao Meng +1
cs.LGcs.AIcs.CLarXiv:2211.03536v22022On the Adversarial Robustness of Multi-Modal Foundation Models
Christian Schlarmann, Matthias Hein
cs.LGcs.AIcs.CRarXiv:2308.10741v12023Reward Shaping to Mitigate Reward Hacking in RLHF
Jiayi Fu, Xuandong Zhao, Chengyuan Yao +2
cs.LGcs.AIcs.CLarXiv:2502.18770v72025CAFE: Catastrophic Data Leakage in Vertical Federated Learning
Xiao Jin, Pin-Yu Chen, Chia-Yi Hsu +2
cs.LGcs.AIarXiv:2110.15122v42021Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Weixun Wang, Shaopan Xiong, Gengru Chen +38
cs.LGcs.DCarXiv:2506.06122v12025Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing
Yusu Qian, Eli Bocek-Rivele, Liangchen Song +5
cs.CVcs.CLcs.LGarXiv:2510.19808v12025Expanding the Capabilities of Reinforcement Learning via Text Feedback
Yuda Song, Lili Chen, Fahim Tajwar +5
cs.LGarXiv:2602.02482v22026Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks
Shuo He, Lang Feng, Qi Wei +3
cs.LGcs.AIarXiv:2602.22817v12026Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond
Chongyu Fan, Jinghan Jia, Yihua Zhang +3
cs.LGcs.CLarXiv:2502.05374v42025Attending to Characters in Neural Sequence Labeling Models
Marek Rei, Gamal K. O. Crichton, Sampo Pyysalo
cs.CLcs.LGcs.NEarXiv:1611.04361v12016Towards Causal VQA: Revealing and Reducing Spurious Correlations by Invariant and Covariant Semantic Editing
Vedika Agarwal, Rakshith Shetty, Mario Fritz
cs.CVcs.CLcs.LGarXiv:1912.07538v32019Differentiable Ranks and Sorting using Optimal Transport
Marco Cuturi, Olivier Teboul, Jean-Philippe Vert
cs.LGstat.MLarXiv:1905.11885v22019Fathom: Reference Workloads for Modern Deep Learning Methods
Robert Adolf, Saketh Rama, Brandon Reagen +2
cs.LGarXiv:1608.06581v12016The Implications of Linguistic Illegibility for LLM Security
James Mickens
cs.LGcs.CRarXiv:2609.02852v12026Gold-medalist Performance in Solving Olympiad Geometry with AlphaGeometry2
Yuri Chervonyi, Trieu H. Trinh, Miroslav Olšák +8
cs.AIcs.LGarXiv:2502.03544v32025Learning Language-Conditioned Robot Behavior from Offline Data and Crowd-Sourced Annotation
Suraj Nair, Eric Mitchell, Kevin Chen +3
cs.ROcs.AIcs.LGarXiv:2109.01115v22021Deep Neural Networks with Random Gaussian Weights: A Universal Classification Strategy?
Raja Giryes, Guillermo Sapiro, Alex M. Bronstein
cs.NEcs.LGstat.MLarXiv:1504.08291v52015A proof that artificial neural networks overcome the curse of dimensionality in the numerical approximation of Black-Scholes partial differential equations
Philipp Grohs, Fabian Hornung, Arnulf Jentzen +1
math.NAcs.LGmath.PRarXiv:1809.02362v32018Learning Spectral-Like Mesh-Free Discretisations
Lucas Gerken Starepravo, Henry Broadley, Steven Lind +1
physics.comp-phcs.LGarXiv:2609.02833v12026Recent Advances in Imitation Learning from Observation
Faraz Torabi, Garrett Warnell, Peter Stone
cs.ROcs.AIcs.LGarXiv:1905.13566v22019Dimension Dependent Correlation Gap Bounds under Restricted Independence
Arjun Ramachandra
math.PRcs.LGmath.COarXiv:2609.02659v12026HyperMC: Multi-Fidelity Hyperparameter Tuning for Stochastic Gradient MCMC
Ming Tan, Xiyun Jiao
stat.MLcs.LGstat.COarXiv:2609.02138v12026Influence Function based Data Poisoning Attacks to Top-N Recommender Systems
Minghong Fang, Neil Zhenqiang Gong, Jia Liu
cs.CRcs.IRcs.LGarXiv:2002.08025v32020Similarity-Aware Personalized Federated Learning in Heterogeneous Environments
Arun Kumar A, Sunil Gupta, Dang Ngyuen +2
cs.LGarXiv:2609.02241v12026CAHR-Net: Condition-Adaptive Hysteresis Reconstruction for Compact and Interpretable Magnetic Core Loss Modeling
Chunye Gong, Cong Yao
cs.LGarXiv:2609.01991v12026Algorithms and Theory for Multiple-Source Adaptation
Judy Hoffman, Mehryar Mohri, Ningshan Zhang
cs.LGstat.MLarXiv:1805.08727v12018A Comparative Study of Graph Representations for GNN-Based Power Grid Control in L2RPN
Adrian Degenkolb, Qiong Huang, Benjamin Schäfer
cs.LGarXiv:2609.02538v12026SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
Hwiwon Lee, Ziqi Zhang, Hanxiao Lu +1
cs.LGcs.CRarXiv:2506.11791v22025Reconstructing subclonal composition and evolution from whole genome sequencing of tumors
Amit G. Deshwar, Shankar Vembu, Christina K. Yung +3
q-bio.PEcs.LGstat.MLarXiv:1406.7250v32014SOFTS: Efficient Multivariate Time Series Forecasting with Series-Core Fusion
Lu Han, Xu-Yang Chen, Han-Jia Ye +1
cs.LGarXiv:2404.14197v32024Robust Federated Learning: The Case of Affine Distribution Shifts
Amirhossein Reisizadeh, Farzan Farnia, Ramtin Pedarsani +1
cs.LGmath.OCstat.MLarXiv:2006.08907v12020Airbert: In-domain Pretraining for Vision-and-Language Navigation
Pierre-Louis Guhur, Makarand Tapaswi, Shizhe Chen +2
cs.CVcs.AIcs.CLarXiv:2108.09105v12021Gossip Learning with Linear Models on Fully Distributed Data
Róbert Ormándi, István Hegedüs, Márk Jelasity
cs.LGcs.DCarXiv:1109.1396v32011AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy
Zihan Liu, Zhuolin Yang, Yang Chen +4
cs.CLcs.AIcs.LGarXiv:2506.13284v12025When Models Manipulate Manifolds: The Geometry of a Counting Task
Wes Gurnee, Emmanuel Ameisen, Isaac Kauvar +4
cs.LGarXiv:2601.04480v12026Kernel-based Reconstruction of Graph Signals
Daniel Romero, Meng Ma, Georgios B. Giannakis
stat.MLcs.LGarXiv:1605.07174v12016When Decodability Is Not Enough: Logical Validity Representations, Behavioral Dissociation, and Causal Tests in Language Models
Smitha Muthya Sudheendra, Jaideep Srivastava
cs.CLcs.LGarXiv:2609.02438v12026