Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
6,121 to 6,180 of 20,193
SafeArena: Evaluating the Safety of Autonomous Web Agents
Ada Defne Tur, Nicholas Meade, Xing Han Lù +6
cs.LGcs.AIcs.CLarXiv:2503.04957v12025Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning
DiJia Su, Hanlin Zhu, Yingchen Xu +3
cs.CLcs.AIcs.LGarXiv:2502.03275v22025Do RNN and LSTM have Long Memory?
Jingyu Zhao, Feiqing Huang, Jia Lv +4
stat.MLcs.LGarXiv:2006.03860v22020Natural Backdoor Attacks on Speech Recognition Models
Jinwen Xin, Xixiang Lyu, Jing Ma
cs.CRcs.LGcs.SDarXiv:2607.15724v12026Generalisation error in learning with random features and the hidden manifold model
Federica Gerace, Bruno Loureiro, Florent Krzakala +2
math.STcs.LGmath.PRarXiv:2002.09339v22020S*: Test Time Scaling for Code Generation
Dacheng Li, Shiyi Cao, Chengkun Cao +6
cs.LGcs.AIarXiv:2502.14382v12025BeamDojo: Learning Agile Humanoid Locomotion on Sparse Footholds
Huayi Wang, Zirui Wang, Junli Ren +4
cs.ROcs.AIcs.LGarXiv:2502.10363v32025Learning Linear-Quadratic Regulators Efficiently with only $\sqrt{T}$ Regret
Alon Cohen, Tomer Koren, Yishay Mansour
cs.LGstat.MLarXiv:1902.06223v22019Reinforcement Learning for Long-Horizon Interactive LLM Agents
Kevin Chen, Marco Cusumano-Towner, Brody Huval +4
cs.LGcs.AIarXiv:2502.01600v32025Improved Speech Enhancement with the Wave-U-Net
Craig Macartney, Tillman Weyde
cs.SDcs.LGcs.NEarXiv:1811.11307v12018Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Chaoqi Wang, Yibo Jiang, Chenghao Yang +2
cs.LGcs.AIstat.MLarXiv:2309.16240v12023DK-GBMKKM: Dynamic Kernel-Space Granular-Ball Multiple Kernel $k$-Means Clustering
Xiaoyu Lian, Yuchao Zhang, Shuyin Xia +2
cs.LGarXiv:2609.00647v12026Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Vaishnavi Shrivastava, Ahmed Awadallah, Vidhisha Balachandran +3
cs.CLcs.LGarXiv:2508.09726v12025AgentEvolver: Towards Efficient Self-Evolving Agent System
Yunpeng Zhai, Shuchang Tao, Cheng Chen +10
cs.LGcs.AIcs.CLarXiv:2511.10395v12025Deep Neural Network Fingerprinting by Conferrable Adversarial Examples
Nils Lukas, Yuxuan Zhang, Florian Kerschbaum
cs.LGcs.CRstat.MLarXiv:1912.00888v42019A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
Andreas Hochlehnert, Hardik Bhatnagar, Vishaal Udandarao +3
cs.LGcs.CLarXiv:2504.07086v22025Poisson-Gamma Dynamical Systems with Time-varying Transition Dynamics
Jiahao Wang, Yijun Wang, Nan Fang +1
cs.LGarXiv:2609.00896v12026Variational Gaussian Process State-Space Models
Roger Frigola, Yutian Chen, Carl E. Rasmussen
cs.LGcs.ROeess.SYarXiv:1406.4905v22014Multi-Agent Design: Optimizing Agents with Better Prompts and Topologies
Han Zhou, Xingchen Wan, Ruoxi Sun +5
cs.LGcs.AIcs.CLarXiv:2502.02533v22025Verdict Instability of OOD Scores under Reference Resampling
Donghoon Lee, Shinjin Kang
cs.LGstat.MLarXiv:2609.00691v12026Optimization with Non-Differentiable Constraints with Applications to Fairness, Recall, Churn, and Other Goals
Andrew Cotter, Heinrich Jiang, Serena Wang +4
cs.LGcs.AIcs.GTarXiv:1809.04198v12018MLGym: A New Framework and Benchmark for Advancing AI Research Agents
Deepak Nathani, Lovish Madaan, Nicholas Roberts +14
cs.CLcs.AIcs.LGarXiv:2502.14499v12025Tree Search for LLM Agent Reinforcement Learning
Yuxiang Ji, Ziyu Ma, Yong Wang +3
cs.LGcs.AIarXiv:2509.21240v32025Improving Adversarial Transferability via Neuron Attribution-Based Attacks
Jianping Zhang, Weibin Wu, Jen-tse Huang +4
cs.LGcs.CRarXiv:2204.00008v12022The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
Tom Wollschläger, Jannes Elstner, Simon Geisler +3
cs.LGcs.AIcs.CLarXiv:2502.17420v22025Memp: Exploring Agent Procedural Memory
Runnan Fang, Yuan Liang, Xiaobin Wang +6
cs.CLcs.AIcs.LGarXiv:2508.06433v42025Diffusion Beats Autoregressive in Data-Constrained Settings
Mihir Prabhudesai, Mengning Wu, Amir Zadeh +2
cs.LGcs.AIcs.CVarXiv:2507.15857v72025Explainable $k$-Means and $k$-Medians Clustering
Sanjoy Dasgupta, Nave Frost, Michal Moshkovitz +1
cs.LGcs.CGcs.DSarXiv:2002.12538v22020Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains
Vighnesh Subramaniam, Yilun Du, Joshua B. Tenenbaum +3
cs.CLcs.AIcs.LGarXiv:2501.05707v22025Follow the Leader If You Can, Hedge If You Must
Steven de Rooij, Tim van Erven, Peter D. Grünwald +1
cs.LGstat.MLarXiv:1301.0534v22013TraveL: Transformer-based Multi-view Path Distributional Representation Learning
Fang He, Tao-yang Fu, Wang-chien Lee
cs.LGcs.AIarXiv:2609.03427v12026RW-LoRA: Communication-Efficient Decentralized LoRA Fine-Tuning via Random Walks
Xingran Chen, Rohit Bhagat, Ghadir Ayache +3
cs.LGcs.AIarXiv:2609.00078v12026Policy Mirror Descent for Reinforcement Learning: Linear Convergence, New Sampling Complexity, and Generalized Problem Classes
Guanghui Lan
cs.LGcs.AImath.OCarXiv:2102.00135v62021Quasar: Datasets for Question Answering by Search and Reading
Bhuwan Dhingra, Kathryn Mazaitis, William W. Cohen
cs.CLcs.IRcs.LGarXiv:1707.03904v22017Convergent Linear Representations of Emergent Misalignment
Anna Soligo, Edward Turner, Senthooran Rajamanoharan +1
cs.LGcs.AIarXiv:2506.11618v22025Foundation models for electricity price forecasting and battery arbitrage: Can they replace market-specific forecasting models?
Arkadiusz Lipiecki, Rafał Weron
cs.LGecon.EMarXiv:2609.00089v12026The PGM-index: a multicriteria, compressed and learned approach to data indexing
Paolo Ferragina, Giorgio Vinciguerra
cs.DScs.DBcs.IRarXiv:1910.06169v12019Curriculum Loss: Robust Learning and Generalization against Label Corruption
Yueming Lyu, Ivor W. Tsang
cs.LGstat.MLarXiv:1905.10045v32019Multifunctional Metasurface Design with a Generative Adversarial Network
Sensong An, Bowen Zheng, Hong Tang +7
physics.opticscs.LGarXiv:1908.04851v22019Flawed in Nature, Perfect through Evolution
J. M. Diederik Kruijssen
cs.LGcs.AIcs.NEarXiv:2609.00129v12026DeepPicker: a Deep Learning Approach for Fully Automated Particle Picking in Cryo-EM
Feng Wang, Huichao Gong, Gaochao liu +5
q-bio.QMcs.LGarXiv:1605.01838v12016Rank1: Test-Time Compute for Reranking in Information Retrieval
Orion Weller, Kathryn Ricci, Eugene Yang +3
cs.IRcs.CLcs.LGarXiv:2502.18418v22025Unveiling COVID-19 from Chest X-ray with deep learning: a hurdles race with small data
Enzo Tartaglione, Carlo Alberto Barbano, Claudio Berzovini +2
eess.IVcs.CVcs.LGarXiv:2004.05405v12020Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
Qixiu Li, Yu Deng, Yaobo Liang +14
cs.ROcs.AIcs.CVarXiv:2510.21571v12025RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
Pranav Atreya, Karl Pertsch, Tony Lee +29
cs.ROcs.LGarXiv:2506.18123v22025Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency
Kaiwen Zheng, Yuji Wang, Qianli Ma +7
cs.CVcs.LGarXiv:2510.08431v32025Persona Features Control Emergent Misalignment
Miles Wang, Tom Dupré la Tour, Olivia Watkins +8
cs.LGcs.AIarXiv:2506.19823v22025A Comprehensive Survey of Dataset Distillation
Shiye Lei, Dacheng Tao
cs.LGarXiv:2301.05603v42023Federated Learning in the Sky: Aerial-Ground Air Quality Sensing Framework with UAV Swarms
Yi Liu, Jiangtian Nie, Xuandi Li +3
eess.SPcs.LGarXiv:2007.12004v12020Artificial Intelligence in the Battle against Coronavirus (COVID-19): A Survey and Future Research Directions
Thanh Thi Nguyen, Quoc Viet Hung Nguyen, Dung Tien Nguyen +7
cs.CYcs.AIcs.LGarXiv:2008.07343v42020Localizing Model Behavior with Path Patching
Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato +1
cs.LGarXiv:2304.05969v22023Reinforcement Learning from Human Feedback
Nathan Lambert
cs.LGarXiv:2504.12501v112025FLARE: Robot Learning with Implicit World Modeling
Ruijie Zheng, Jing Wang, Scott Reed +18
cs.ROcs.LGarXiv:2505.15659v12025k-Space Deep Learning for Accelerated MRI
Yoseob Han, Leonard Sunwoo, Jong Chul Ye
cs.CVcs.LGstat.MLarXiv:1805.03779v32018Muon Optimizes Under Spectral Norm Constraints
Lizhang Chen, Jonathan Li, Qiang Liu
cs.LGmath.OCstat.MLarXiv:2506.15054v22025Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation Models
Xiyuan Zhang, Danielle C. Maddix, Junming Yin +11
cs.LGarXiv:2510.21204v12025Distilled One-Shot Federated Learning
Yanlin Zhou, George Pu, Xiyao Ma +2
cs.LGcs.AIstat.MLarXiv:2009.07999v32020Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
Audrey Huang, Adam Block, Qinghua Liu +3
cs.AIcs.LGstat.MLarXiv:2503.21878v22025Message-Aware Graph Attention Networks for Large-Scale Multi-Robot Path Planning
Qingbiao Li, Weizhe Lin, Zhe Liu +1
cs.ROcs.DCcs.LGarXiv:2011.13219v22020hLLM: Single Pass Decoding for Generative Reranking
Emil Laftchiev, Prachi Agrawal, Moe Kayali +7
cs.LGcs.AIcs.IRarXiv:2609.01807v12026