Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

6,841 to 6,900 of 20,199

  1. VAFL: a Method of Vertical Asynchronous Federated Learning

    Tianyi Chen, Xiao Jin, Yuejiao Sun +1

    cs.LGcs.DCmath.OCarXiv:2007.06081v12020
  2. SecAlign: Defending Against Prompt Injection with Preference Optimization

    Sizhe Chen, Arman Zharmagambetov, Saeed Mahloujifar +3

    cs.CRcs.LGarXiv:2410.05451v32024
  3. Improved Mean Flows: On the Challenges of Fastforward Generative Models

    Zhengyang Geng, Yiyang Lu, Zongze Wu +3

    cs.CVcs.LGarXiv:2512.02012v22025
  4. A Time-Vertex Signal Processing Framework

    Francesco Grassi, Andreas Loukas, Nathanaël Perraudin +1

    cs.LGarXiv:1705.02307v12017
  5. SciFive: a text-to-text transformer model for biomedical literature

    Long N. Phan, James T. Anibal, Hieu Tran +4

    cs.CLcs.AIcs.LGarXiv:2106.03598v12021
  6. prDeep: Robust Phase Retrieval with a Flexible Deep Network

    Christopher A. Metzler, Philip Schniter, Ashok Veeraraghavan +1

    stat.MLcs.LGarXiv:1803.00212v22018
  7. RWKV-7 "Goose" with Expressive Dynamic State Evolution

    Bo Peng, Ruichong Zhang, Daniel Goldstein +15

    cs.CLcs.AIcs.LGarXiv:2503.14456v22025
  8. Independent Policy Gradient Methods for Competitive Reinforcement Learning

    Constantinos Daskalakis, Dylan J. Foster, Noah Golowich

    cs.LGarXiv:2101.04233v12021
  9. Computing stable configurations of confined smectic liquid crystals with a deep variational framework

    Yuchen Xie, Baoming Shi, Yucen Han +1

    cond-mat.softcs.LGarXiv:2609.03389v12026
  10. MambaIRv2: Attentive State Space Restoration

    Hang Guo, Yong Guo, Yaohua Zha +5

    eess.IVcs.CVcs.LGarXiv:2411.15269v22024
  11. From Language to Programs: Bridging Reinforcement Learning and Maximum Marginal Likelihood

    Kelvin Guu, Panupong Pasupat, Evan Zheran Liu +1

    cs.AIcs.LGstat.MLarXiv:1704.07926v12017
  12. On the Importance of Difficulty Calibration in Membership Inference Attacks

    Lauren Watson, Chuan Guo, Graham Cormode +1

    cs.CRcs.LGarXiv:2111.08440v22021
  13. dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching

    Zhiyuan Liu, Yicun Yang, Yaojie Zhang +6

    cs.LGcs.AIcs.CLarXiv:2506.06295v32025
  14. Multiple-output support vector regression with a firefly algorithm for interval-valued stock price index forecasting

    Tao Xiong, Yukun Bao, Zhongyi Hu

    cs.CEcs.LGq-fin.STarXiv:1401.1916v12014
  15. Mutual Information, Neural Networks and the Renormalization Group

    Maciej Koch-Janusz, Zohar Ringel

    cond-mat.dis-nncond-mat.stat-mechcs.ITarXiv:1704.06279v22017
  16. Multiple decision trees

    Suk Wah Kwok, Chris Carter

    cs.LGcs.AIstat.MLarXiv:1304.2363v12013
  17. Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks

    Tilman Räuker, Anson Ho, Stephen Casper +1

    cs.LGcs.AIcs.CLarXiv:2207.13243v62022
  18. Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

    NVIDIA, :, Alisson Azzolini +51

    cs.AIcs.CVcs.LGarXiv:2503.15558v32025
  19. Illiterate DALL-E Learns to Compose

    Gautam Singh, Fei Deng, Sungjin Ahn

    cs.CVcs.LGarXiv:2110.11405v32021
  20. On the Effects of Data Scale on UI Control Agents

    Wei Li, William Bishop, Alice Li +4

    cs.AIcs.LGarXiv:2406.03679v62024
  21. Nested Learning: The Illusion of Deep Learning Architectures

    Ali Behrouz, Meisam Razaviyayn, Peilin Zhong +1

    cs.LGcs.AIarXiv:2512.24695v12025
  22. Making Better Mistakes: Leveraging Class Hierarchies with Deep Networks

    Luca Bertinetto, Romain Mueller, Konstantinos Tertikas +2

    cs.CVcs.LGarXiv:1912.09393v22019
  23. Evaluation and Benchmarking of LLM Agents: A Survey

    Mahmoud Mohammadi, Yipeng Li, Jane Lo +1

    cs.LGcs.AIarXiv:2507.21504v12025
  24. mHC: Manifold-Constrained Hyper-Connections

    Zhenda Xie, Yixuan Wei, Huanqi Cao +17

    cs.CLcs.AIcs.LGarXiv:2512.24880v22025
  25. The Diffusion Duality

    Subham Sekhar Sahoo, Justin Deschenaux, Aaron Gokaslan +3

    cs.LGcs.AIcs.CLarXiv:2506.10892v32025
  26. Foundation Models Meet Agriculture: Challenges Beyond Pretraining

    Vishal Nedungadi, Xingguo Xiong, Marc Rußwurm +1

    cs.LGarXiv:2608.30392v12026
  27. When Can Conditional Flow Matching Replace Pointwise Negative Log-Likelihood?

    Yansen Han, Hongxin Sun, Tao Lin

    cs.LGcs.AIarXiv:2608.28010v12026
  28. Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions

    Jing Gu, Eliana Stefani, Qi Wu +2

    cs.CVcs.AIcs.CLarXiv:2203.12667v32022
  29. Soft Adaptive Policy Optimization

    Chang Gao, Chujie Zheng, Xiong-Hui Chen +7

    cs.LGcs.AIcs.CLarXiv:2511.20347v22025
  30. ARCHER: Amortized cross-specimen pose estimation for cryo-electron microscopy

    Nhan D. Nguyen, Bao Pham

    cs.LGmath-phq-bio.BMarXiv:2608.22029v12026
  31. NeuronGuard: Robust LLM Safety Alignment via Ablation-Aware Safety Signal Redistribution

    Anjun Gao, Yueyang Quan, Yufei Xia +2

    cs.CRcs.AIcs.IRarXiv:2608.23959v12026
  32. Beyond Observed Auxiliary Relations: Environment-Conditioned Modeling for Multi-Behavior Recommendation

    Seunghan Lee, Hyunsik Yoo, Jian Kang +2

    cs.AIcs.LGarXiv:2608.22920v12026
  33. Semi-Supervised Classification with Informative Missing Labels in Weibull Mixture Models

    Jinran Wu, You-Gan Wang, Geoffrey J. McLachlan

    stat.MLcs.LGarXiv:2609.00774v12026
  34. Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning

    Varun Giridhar, Anant Khandelwal, Jeremy A. Collins +2

    cs.ROcs.LGarXiv:2608.21204v12026
  35. MemCatalyst: Amplifying Data Auditing on Vision-Language Models via Data Poisoning

    Xukun Luan, Jinyan Liu, Yuhui Gong +4

    cs.CRcs.LGarXiv:2608.17722v12026
  36. Online Generalized Sparse Regression: How Does Overparametrization Help?

    Shuoguang Yang, Qiang Sun

    stat.MLcs.LGmath.STarXiv:2608.17466v12026
  37. On the Pseudo-Mixing of Kac's Walk

    Natesh S. Pillai, Aaron Smith, Vinod Vaikuntanathan

    math.PRcs.CRcs.LGarXiv:2608.17374v12026
  38. Learning to Price with Persuasion

    Maria-Florina Balcan, Tejas Pagare, Karan Singh

    cs.GTcs.LGecon.THarXiv:2608.16699v12026
  39. Deep Vision in Smart Manufacturing: MODERN Framework for Intelligent Quality Monitoring and Diagnosis

    Yicheng Kang, Yuling Jiao, Xin Geng +1

    stat.MLcs.LGarXiv:2608.13937v12026
  40. Toward Optimal Second-Order Path-Length Guarantee for Adversarial Multi-Armed Bandits

    Mengxiao Zhang

    cs.LGarXiv:2608.15996v12026
  41. KOALA: Koopman Operator Learning for WiFi-Based Anticipatory Hum

    Quang-Anh N. D., Duc Pham Minh, Thao Phuong Pham +3

    cs.LGarXiv:2608.15815v12026
  42. Optimal Lower Bounds for Networked Information Aggregation

    Ambar Pal

    cs.LGcs.AIstat.MLarXiv:2608.15472v12026
  43. Anatomy of a Quantized Agent: VRAM Stability and Forecasting in Code-Synthesis Agentic Workloads

    Anubhab Banerjee

    cs.AIcs.DCcs.LGarXiv:2608.15117v12026
  44. Lipschitz Bandits with Arbitrary Feedback Delays

    Yuhao Liu, Yu Chen, Longbo Huang

    cs.LGarXiv:2608.15036v12026
  45. KV Cache Compression Through the Lens of Transform Coding

    Hannah Laus, Claudio Mayrink Verdun, Hao Wang +2

    cs.LGcs.CLeess.SParXiv:2608.14191v12026
  46. Learning to Run Power Networks: Effective AlphaZero-inspired Topological Control

    Lukas Zetto, Benjamin Schäfer, Qiong Huang

    cs.LGarXiv:2608.14114v12026
  47. Federated Prompt Learning: A Unified Framework, Empirical Analysis, and Future Directions

    Qinglin Yang, Chen Qiu, Hongyuan Zhang +3

    cs.LGcs.AIcs.DCarXiv:2608.13844v12026
  48. PPAPlace: Differentiable Cross-Stage Objectives for Chip Placement Optimization

    Ruogu Chen, Jie Han

    cs.LGcs.AIcs.ARarXiv:2608.13790v12026
  49. Modular TTT: Rethinking Test-Time Training as Composable Modules

    Bohao Tang, Zhen Qin, Yuqi Pan +3

    cs.LGcs.CLarXiv:2608.07110v12026
  50. $β$-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

    Jiawei Xu, Minghui Liu, Juzheng Zhang +2

    cs.LGarXiv:2607.28582v12026
  51. Length Penalties Make Chain-of-Thought Less Monitorable

    Bryce Little

    cs.AIcs.CLcs.LGarXiv:2607.09786v32026
  52. Learning to Trigger: Reinforcement Learning at the Large Hadron Collider

    Zixin Ding, Shaghayegh Emami, Giovanna Salvi +7

    cs.LGcs.AIhep-exarXiv:2606.23993v32026
  53. Comparing Linear Probes with Mahalanobis Cosine Similarity

    Zhuofan Josh Ying, Peter Hase, Nikolaus Kriegeskorte

    cs.LGarXiv:2606.19603v12026
  54. STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability

    Haipeng Luo, Qingfeng Sun, Songli Wu +4

    cs.LGcs.AIcs.CLarXiv:2606.19236v12026
  55. MUGEN: Generating Unlearnable Graph Examples for Multiple Learning Tasks

    Ziyan Liu, Chengshuai Zhao, Huan Liu

    cs.LGarXiv:2609.00696v22026
  56. Can Generalist Agents Automate Data Curation?

    Feiyang Kang, Hanze Li, Adam Nguyen +5

    cs.AIcs.CLcs.CVarXiv:2606.04261v12026
  57. Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention

    Jing Huang, Daniel Wurgaft, Rachit Bansal +6

    cs.LGarXiv:2605.29548v22026
  58. Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models

    Mingze Wang, Shuchen Zhu, Yuxin Fang +3

    cs.LGcs.AIstat.MLarXiv:2605.26895v12026
  59. Resource Management in Wireless Networks via Multi-Agent Deep Reinforcement Learning

    Navid Naderializadeh, Jaroslaw Sydir, Meryem Simsek +1

    cs.LGcs.ITcs.MAarXiv:2002.06215v22020
  60. Good Token Hunting: A Hitchhiker's Guide to Token Selection for Visual Geometry Transformers

    Shuhong Zheng, Michael Oechsle, Erik Sandström +3

    cs.CVcs.AIcs.GRarXiv:2605.23892v12026