Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

6,121 to 6,180 of 20,193

  1. SafeArena: Evaluating the Safety of Autonomous Web Agents

    Ada Defne Tur, Nicholas Meade, Xing Han Lù +6

    cs.LGcs.AIcs.CLarXiv:2503.04957v12025
  2. Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning

    DiJia Su, Hanlin Zhu, Yingchen Xu +3

    cs.CLcs.AIcs.LGarXiv:2502.03275v22025
  3. Do RNN and LSTM have Long Memory?

    Jingyu Zhao, Feiqing Huang, Jia Lv +4

    stat.MLcs.LGarXiv:2006.03860v22020
  4. Natural Backdoor Attacks on Speech Recognition Models

    Jinwen Xin, Xixiang Lyu, Jing Ma

    cs.CRcs.LGcs.SDarXiv:2607.15724v12026
  5. Generalisation error in learning with random features and the hidden manifold model

    Federica Gerace, Bruno Loureiro, Florent Krzakala +2

    math.STcs.LGmath.PRarXiv:2002.09339v22020
  6. S*: Test Time Scaling for Code Generation

    Dacheng Li, Shiyi Cao, Chengkun Cao +6

    cs.LGcs.AIarXiv:2502.14382v12025
  7. BeamDojo: Learning Agile Humanoid Locomotion on Sparse Footholds

    Huayi Wang, Zirui Wang, Junli Ren +4

    cs.ROcs.AIcs.LGarXiv:2502.10363v32025
  8. Learning Linear-Quadratic Regulators Efficiently with only $\sqrt{T}$ Regret

    Alon Cohen, Tomer Koren, Yishay Mansour

    cs.LGstat.MLarXiv:1902.06223v22019
  9. Reinforcement Learning for Long-Horizon Interactive LLM Agents

    Kevin Chen, Marco Cusumano-Towner, Brody Huval +4

    cs.LGcs.AIarXiv:2502.01600v32025
  10. Improved Speech Enhancement with the Wave-U-Net

    Craig Macartney, Tillman Weyde

    cs.SDcs.LGcs.NEarXiv:1811.11307v12018
  11. Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

    Chaoqi Wang, Yibo Jiang, Chenghao Yang +2

    cs.LGcs.AIstat.MLarXiv:2309.16240v12023
  12. DK-GBMKKM: Dynamic Kernel-Space Granular-Ball Multiple Kernel $k$-Means Clustering

    Xiaoyu Lian, Yuchao Zhang, Shuyin Xia +2

    cs.LGarXiv:2609.00647v12026
  13. Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

    Vaishnavi Shrivastava, Ahmed Awadallah, Vidhisha Balachandran +3

    cs.CLcs.LGarXiv:2508.09726v12025
  14. AgentEvolver: Towards Efficient Self-Evolving Agent System

    Yunpeng Zhai, Shuchang Tao, Cheng Chen +10

    cs.LGcs.AIcs.CLarXiv:2511.10395v12025
  15. Deep Neural Network Fingerprinting by Conferrable Adversarial Examples

    Nils Lukas, Yuxuan Zhang, Florian Kerschbaum

    cs.LGcs.CRstat.MLarXiv:1912.00888v42019
  16. A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility

    Andreas Hochlehnert, Hardik Bhatnagar, Vishaal Udandarao +3

    cs.LGcs.CLarXiv:2504.07086v22025
  17. Poisson-Gamma Dynamical Systems with Time-varying Transition Dynamics

    Jiahao Wang, Yijun Wang, Nan Fang +1

    cs.LGarXiv:2609.00896v12026
  18. Variational Gaussian Process State-Space Models

    Roger Frigola, Yutian Chen, Carl E. Rasmussen

    cs.LGcs.ROeess.SYarXiv:1406.4905v22014
  19. Multi-Agent Design: Optimizing Agents with Better Prompts and Topologies

    Han Zhou, Xingchen Wan, Ruoxi Sun +5

    cs.LGcs.AIcs.CLarXiv:2502.02533v22025
  20. Verdict Instability of OOD Scores under Reference Resampling

    Donghoon Lee, Shinjin Kang

    cs.LGstat.MLarXiv:2609.00691v12026
  21. Optimization with Non-Differentiable Constraints with Applications to Fairness, Recall, Churn, and Other Goals

    Andrew Cotter, Heinrich Jiang, Serena Wang +4

    cs.LGcs.AIcs.GTarXiv:1809.04198v12018
  22. MLGym: A New Framework and Benchmark for Advancing AI Research Agents

    Deepak Nathani, Lovish Madaan, Nicholas Roberts +14

    cs.CLcs.AIcs.LGarXiv:2502.14499v12025
  23. Tree Search for LLM Agent Reinforcement Learning

    Yuxiang Ji, Ziyu Ma, Yong Wang +3

    cs.LGcs.AIarXiv:2509.21240v32025
  24. Improving Adversarial Transferability via Neuron Attribution-Based Attacks

    Jianping Zhang, Weibin Wu, Jen-tse Huang +4

    cs.LGcs.CRarXiv:2204.00008v12022
  25. The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence

    Tom Wollschläger, Jannes Elstner, Simon Geisler +3

    cs.LGcs.AIcs.CLarXiv:2502.17420v22025
  26. Memp: Exploring Agent Procedural Memory

    Runnan Fang, Yuan Liang, Xiaobin Wang +6

    cs.CLcs.AIcs.LGarXiv:2508.06433v42025
  27. Diffusion Beats Autoregressive in Data-Constrained Settings

    Mihir Prabhudesai, Mengning Wu, Amir Zadeh +2

    cs.LGcs.AIcs.CVarXiv:2507.15857v72025
  28. Explainable $k$-Means and $k$-Medians Clustering

    Sanjoy Dasgupta, Nave Frost, Michal Moshkovitz +1

    cs.LGcs.CGcs.DSarXiv:2002.12538v22020
  29. Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains

    Vighnesh Subramaniam, Yilun Du, Joshua B. Tenenbaum +3

    cs.CLcs.AIcs.LGarXiv:2501.05707v22025
  30. Follow the Leader If You Can, Hedge If You Must

    Steven de Rooij, Tim van Erven, Peter D. Grünwald +1

    cs.LGstat.MLarXiv:1301.0534v22013
  31. TraveL: Transformer-based Multi-view Path Distributional Representation Learning

    Fang He, Tao-yang Fu, Wang-chien Lee

    cs.LGcs.AIarXiv:2609.03427v12026
  32. RW-LoRA: Communication-Efficient Decentralized LoRA Fine-Tuning via Random Walks

    Xingran Chen, Rohit Bhagat, Ghadir Ayache +3

    cs.LGcs.AIarXiv:2609.00078v12026
  33. Policy Mirror Descent for Reinforcement Learning: Linear Convergence, New Sampling Complexity, and Generalized Problem Classes

    Guanghui Lan

    cs.LGcs.AImath.OCarXiv:2102.00135v62021
  34. Quasar: Datasets for Question Answering by Search and Reading

    Bhuwan Dhingra, Kathryn Mazaitis, William W. Cohen

    cs.CLcs.IRcs.LGarXiv:1707.03904v22017
  35. Convergent Linear Representations of Emergent Misalignment

    Anna Soligo, Edward Turner, Senthooran Rajamanoharan +1

    cs.LGcs.AIarXiv:2506.11618v22025
  36. Foundation models for electricity price forecasting and battery arbitrage: Can they replace market-specific forecasting models?

    Arkadiusz Lipiecki, Rafał Weron

    cs.LGecon.EMarXiv:2609.00089v12026
  37. The PGM-index: a multicriteria, compressed and learned approach to data indexing

    Paolo Ferragina, Giorgio Vinciguerra

    cs.DScs.DBcs.IRarXiv:1910.06169v12019
  38. Curriculum Loss: Robust Learning and Generalization against Label Corruption

    Yueming Lyu, Ivor W. Tsang

    cs.LGstat.MLarXiv:1905.10045v32019
  39. Multifunctional Metasurface Design with a Generative Adversarial Network

    Sensong An, Bowen Zheng, Hong Tang +7

    physics.opticscs.LGarXiv:1908.04851v22019
  40. Flawed in Nature, Perfect through Evolution

    J. M. Diederik Kruijssen

    cs.LGcs.AIcs.NEarXiv:2609.00129v12026
  41. DeepPicker: a Deep Learning Approach for Fully Automated Particle Picking in Cryo-EM

    Feng Wang, Huichao Gong, Gaochao liu +5

    q-bio.QMcs.LGarXiv:1605.01838v12016
  42. Rank1: Test-Time Compute for Reranking in Information Retrieval

    Orion Weller, Kathryn Ricci, Eugene Yang +3

    cs.IRcs.CLcs.LGarXiv:2502.18418v22025
  43. Unveiling COVID-19 from Chest X-ray with deep learning: a hurdles race with small data

    Enzo Tartaglione, Carlo Alberto Barbano, Claudio Berzovini +2

    eess.IVcs.CVcs.LGarXiv:2004.05405v12020
  44. Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos

    Qixiu Li, Yu Deng, Yaobo Liang +14

    cs.ROcs.AIcs.CVarXiv:2510.21571v12025
  45. RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies

    Pranav Atreya, Karl Pertsch, Tony Lee +29

    cs.ROcs.LGarXiv:2506.18123v22025
  46. Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency

    Kaiwen Zheng, Yuji Wang, Qianli Ma +7

    cs.CVcs.LGarXiv:2510.08431v32025
  47. Persona Features Control Emergent Misalignment

    Miles Wang, Tom Dupré la Tour, Olivia Watkins +8

    cs.LGcs.AIarXiv:2506.19823v22025
  48. A Comprehensive Survey of Dataset Distillation

    Shiye Lei, Dacheng Tao

    cs.LGarXiv:2301.05603v42023
  49. Federated Learning in the Sky: Aerial-Ground Air Quality Sensing Framework with UAV Swarms

    Yi Liu, Jiangtian Nie, Xuandi Li +3

    eess.SPcs.LGarXiv:2007.12004v12020
  50. Artificial Intelligence in the Battle against Coronavirus (COVID-19): A Survey and Future Research Directions

    Thanh Thi Nguyen, Quoc Viet Hung Nguyen, Dung Tien Nguyen +7

    cs.CYcs.AIcs.LGarXiv:2008.07343v42020
  51. Localizing Model Behavior with Path Patching

    Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato +1

    cs.LGarXiv:2304.05969v22023
  52. Reinforcement Learning from Human Feedback

    Nathan Lambert

    cs.LGarXiv:2504.12501v112025
  53. FLARE: Robot Learning with Implicit World Modeling

    Ruijie Zheng, Jing Wang, Scott Reed +18

    cs.ROcs.LGarXiv:2505.15659v12025
  54. k-Space Deep Learning for Accelerated MRI

    Yoseob Han, Leonard Sunwoo, Jong Chul Ye

    cs.CVcs.LGstat.MLarXiv:1805.03779v32018
  55. Muon Optimizes Under Spectral Norm Constraints

    Lizhang Chen, Jonathan Li, Qiang Liu

    cs.LGmath.OCstat.MLarXiv:2506.15054v22025
  56. Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation Models

    Xiyuan Zhang, Danielle C. Maddix, Junming Yin +11

    cs.LGarXiv:2510.21204v12025
  57. Distilled One-Shot Federated Learning

    Yanlin Zhou, George Pu, Xiyao Ma +2

    cs.LGcs.AIstat.MLarXiv:2009.07999v32020
  58. Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment

    Audrey Huang, Adam Block, Qinghua Liu +3

    cs.AIcs.LGstat.MLarXiv:2503.21878v22025
  59. Message-Aware Graph Attention Networks for Large-Scale Multi-Robot Path Planning

    Qingbiao Li, Weizhe Lin, Zhe Liu +1

    cs.ROcs.DCcs.LGarXiv:2011.13219v22020
  60. hLLM: Single Pass Decoding for Generative Reranking

    Emil Laftchiev, Prachi Agrawal, Moe Kayali +7

    cs.LGcs.AIcs.IRarXiv:2609.01807v12026