Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

17,761 to 17,820 of 20,193

  1. Meta-learning In-Context Enables Training-Free Cross Subject Brain Decoding

    Mu Nan, Muquan Yu, Weijian Mai +12

    cs.LGq-bio.NCarXiv:2604.08537v12026
  2. Envisioning the Future, One Step at a Time

    Stefan Andreas Baumann, Jannik Wiese, Tommaso Martorella +2

    cs.CVcs.AIcs.LGarXiv:2604.09527v12026
  3. MixFlow: Mixed Source Distributions Improve Rectified Flows

    Nazir Nayal, Christopher Wewer, Jan Eric Lenssen

    cs.CVcs.LGarXiv:2604.09181v12026
  4. DMax: Aggressive Parallel Decoding for dLLMs

    Zigeng Chen, Gongfan Fang, Xinyin Ma +2

    cs.LGcs.AIarXiv:2604.08302v32026
  5. Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Video

    Chanhyuk Choi, Taesoo Kim, Donggyu Lee +2

    cs.CVcs.LGarXiv:2604.07786v22026
  6. Efficient RL Training for LLMs with Experience Replay

    Charles Arnal, Vivien Cabannes, Taco Cohen +2

    cs.LGarXiv:2604.08706v12026
  7. 3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding

    Makanjuola Ogunleye, Eman Abdelrahman, Ismini Lourentzou

    cs.CVcs.AIcs.LGarXiv:2604.08645v12026
  8. EquiformerV3: Scaling Efficient, Expressive, and General SE(3)-Equivariant Graph Attention Transformers

    Yi-Lun Liao, Alexander J. Hoffman, Sabrina C. Shen +3

    cs.LGcs.AIphysics.comp-pharXiv:2604.09130v12026
  9. Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types

    Hadas Orgad, Boyi Wei, Kaden Zheng +4

    cs.CLcs.AIcs.LGarXiv:2604.09544v22026
  10. ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion

    Lifeng Chen, Tianqi You, Hao Liu +8

    cs.LGcs.AIeess.IVarXiv:2604.09450v22026
  11. Hierarchical SVG Tokenization: Learning Compact Visual Programs for Scalable Vector Graphics Modeling

    Ximing Xing, Ziteng Xue, Zhenxi Li +8

    cs.LGarXiv:2604.05072v22026
  12. Not All Denoising Steps Are Equal: Model Scheduling for Faster Masked Diffusion Language Models

    Ivan Sedykh, Nikita Sorokin, Valentin Malykh

    cs.LGcs.CLarXiv:2604.02340v22026
  13. A Temporally Augmented Graph Attention Network for Affordance Classification

    Ami Chopra, Supriya Bordoloi, Shyamanta M. Hazarika

    cs.LGcs.AIarXiv:2604.10149v12026
  14. Beyond Perception Errors: Semantic Fixation in Large Vision-Language Models

    Md Tanvirul Alam

    cs.CVcs.LGarXiv:2604.12119v12026
  15. ADD for Multi-Bit Image Watermarking

    An Luo, Jie Ding

    stat.MLcs.AIcs.LGarXiv:2604.11491v12026
  16. How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models

    Gregory N. Frank

    cs.CLcs.AIcs.LGarXiv:2604.04385v52026
  17. Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind

    Hanqi Xiao, Vaidehi Patil, Zaid Khan +3

    cs.CLcs.AIcs.LGarXiv:2604.11666v22026
  18. Diversity Without Fidelity: A Solver-Sampler Mismatch in Multi-Agent LLM Negotiation Simulation

    Sandro Andric

    cs.LGcs.AIcs.CYarXiv:2604.11840v32026
  19. Eliciting Medical Reasoning with Knowledge-enhanced Data Synthesis: A Semi-Supervised Reinforcement Learning Approach

    Haolin Li, Shuyang Jiang, Ruipeng Zhang +3

    cs.LGcs.CLarXiv:2604.11547v12026
  20. Continuous Adversarial Flow Models

    Shanchuan Lin, Ceyuan Yang, Zhijie Lin +2

    cs.LGcs.CVarXiv:2604.11521v12026
  21. Solving Physics Olympiad via Reinforcement Learning on Physics Simulators

    Mihir Prabhudesai, Aryan Satpathy, Yangmin Li +6

    cs.LGcs.AIcs.CVarXiv:2604.11805v12026
  22. IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs

    Yuzhen Mao, Qitong Wang, Martin Ester +1

    cs.LGcs.AIarXiv:2604.10539v12026
  23. Rethinking the Diffusion Model from a Langevin Perspective

    Candi Zheng, Yuan Lan

    cs.LGcs.AIcs.CVarXiv:2604.10465v12026
  24. PokeRL: Reinforcement Learning for Pokemon Red

    Dheeraj Mudireddy, Sai Patibandla

    cs.LGarXiv:2604.10812v12026
  25. Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration

    Zhipeng Chen, Tao Qian, Wayne Xin Zhao +1

    cs.LGcs.AIcs.CLarXiv:2604.11446v12026
  26. The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping

    Yang Liu, Enxi Wang, Yufei Gao +6

    cs.LGcs.AIcs.CLarXiv:2604.11297v12026
  27. LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety

    Junxiao Yang, Haoran Liu, Jinzhe Tu +9

    cs.LGcs.AIcs.CLarXiv:2604.12710v22026
  28. ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents

    Fei Tang, Zhiqiong Lu, Boxuan Zhang +4

    cs.LGcs.AIcs.CLarXiv:2604.11784v12026
  29. Towards Autonomous Mechanistic Reasoning in Virtual Cells

    Yunhui Jang, Lu Zhu, Jake Fawkes +3

    cs.LGcs.AIarXiv:2604.11661v32026
  30. UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy Optimization

    Zhengxi Lu, Fei Tang, Guangyi Liu +8

    cs.LGarXiv:2604.13822v12026
  31. VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization

    Andrei Atanov, Jesse Allardice, Roman Bachmann +6

    cs.CVcs.LGarXiv:2604.12887v12026
  32. Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems

    Jiacheng Liu, Xiaohan Zhao, Xinyi Shang +1

    cs.SEcs.AIcs.CLarXiv:2604.14228v22026
  33. Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding

    Eun Woo Im, Dhruv Madhwal, Vivek Gupta

    cs.LGarXiv:2604.13313v12026
  34. TIP: Token Importance in On-Policy Distillation

    Yuanda Xu, Hejian Sang, Zhengze Zhou +3

    cs.LGcs.AIarXiv:2604.14084v42026
  35. From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space

    Yuqiao Tan, Minzheng Wang, Bo Liu +5

    cs.LGcs.AIcs.CLarXiv:2604.14142v12026
  36. C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences

    Akira Kawabata, Saku Sugawara

    cs.CLcs.LGarXiv:2604.13618v12026
  37. Three-Phase Transformer

    Mohammad R. Abu Ayyash

    cs.CLcs.AIcs.LGarXiv:2604.14430v12026
  38. (1D) Ordered Tokens Enable Efficient Test-Time Search

    Zhitong Gao, Parham Rezaei, Ali Cy +7

    cs.CVcs.AIcs.LGarXiv:2604.15453v12026
  39. AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization

    Genghan Zhang, Shaowei Zhu, Anjiang Wei +6

    cs.LGcs.CLarXiv:2511.15915v22025
  40. Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

    Xiaohua Wang, Muzhao Tian, Yuqi Zeng +20

    cs.LGarXiv:2604.13602v12026
  41. Where does output diversity collapse in post-training?

    Constantinos Karouzos, Xingwei Tan, Nikolaos Aletras

    cs.CLcs.AIcs.LGarXiv:2604.16027v12026
  42. GFT: From Imitation to Reward Fine-Tuning with Unbiased Group Advantages and Dynamic Coefficient Rectification

    Wangjie Gan, Miao Pan, Linbo Xi +4

    cs.AIcs.LGarXiv:2604.14258v32026
  43. An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental Learning

    Quyen Tran, Hai Nguyen, Hoang Phan +6

    cs.LGcs.CVarXiv:2211.16780v42022
  44. LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning

    Bowen Ping, Zijun Chen, Tingfeng Hui +4

    cs.LGcs.CLarXiv:2604.14922v12026
  45. PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research

    Tingjia Miao, Wenkai Jin, Muhua Zhang +19

    cs.LGcs.AIphysics.data-anarXiv:2604.15411v12026
  46. Maximal Brain Damage Without Data or Optimization: Disrupting Neural Networks via Sign-Bit Flips

    Ido Galil, Moshe Kimhi, Ran El-Yaniv

    cs.LGcs.AIcs.CVarXiv:2502.07408v22025
  47. PerturbRx: Learning Treatment-Conditioned Latent Transitions for Patient Drug Response Prediction

    Yoshitaka Inoue, Minoh Jeong, Alfred Hero +2

    q-bio.QMcs.LGarXiv:2608.21349v12026
  48. TwinTrack: Post-hoc Multi-Rater Calibration for Medical Image Segmentation

    Tristan Kirscher, Alexandra Ertl, Klaus Maier-Hein +3

    cs.LGarXiv:2604.15950v22026
  49. Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints

    Xinge Liu, Terry Jingchen Zhang, Bernhard Schölkopf +2

    cs.LGcs.AIarXiv:2604.15664v22026
  50. Structured Scaling of AI Discovery Across Diverse Scientific Domains

    Haotian Ye, Haowei Lin, Jingyi Tang +30

    cs.LGcs.AIarXiv:2604.19341v22026
  51. Advanced Linear Algebra with Applications - Part I (Numerical linear algebra for PDEs, machine learning, and data assimilation)

    Victorita Dolean, Jemima Tabeart

    math.NAcs.LGarXiv:2608.21234v12026
  52. The Exceedance Design Effect: Effective Sample Size for Thresholds under Clustering

    Adam Noonan

    stat.MLcs.LGarXiv:2608.21262v12026
  53. EasyVideoR1: Easier RL for Video Understanding

    Chuanyu Qin, Chenxu Yang, Qingyi Si +6

    cs.CVcs.LGarXiv:2604.16893v12026
  54. Truthful Calibration Measures for Sequential Prediction

    Anagha Gokul, Jason Hartline, Lunjia Hu +2

    cs.DScs.GTcs.LGarXiv:2608.21348v12026
  55. Test-Time Adaptation for EEG Foundation Models: A Systematic Study under Real-World Distribution Shifts

    Gabriel Jason Lee, Jathurshan Pradeepkumar, Jimeng Sun

    cs.LGcs.AIeess.SParXiv:2604.16926v22026
  56. Agents Explore but Agents Ignore: LLMs Lack Environmental Curiosity

    Leon Engländer, Sophia Althammer, Ahmet Üstün +2

    cs.CLcs.LGarXiv:2604.17609v12026
  57. Just Repair: A Minimal Denoising Network for Time Series Anomaly Detection

    Kadir-Kaan Özer, René Ebeling, Markus Enzweiler

    cs.LGcs.AIarXiv:2604.17388v32026
  58. On the Transferability of Agricultural Weed Detection Under Cross-Field Distribution Shift

    Nikhilesh Prabhakar, Pranuthi Tenali, Wilfredo Abudeye Fernandez +5

    cs.CVcs.LGarXiv:2608.21254v12026
  59. When Can LLMs Learn to Reason with Weak Supervision?

    Salman Rahman, Jingyan Shen, Anna Mordvina +3

    cs.LGcs.AIarXiv:2604.18574v12026
  60. MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval

    Shaden Alshammari, Kevin Wen, Abrar Zainal +5

    cs.AIcs.DLcs.IRarXiv:2604.18584v22026