Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
17,761 to 17,820 of 20,193
Meta-learning In-Context Enables Training-Free Cross Subject Brain Decoding
Mu Nan, Muquan Yu, Weijian Mai +12
cs.LGq-bio.NCarXiv:2604.08537v12026Envisioning the Future, One Step at a Time
Stefan Andreas Baumann, Jannik Wiese, Tommaso Martorella +2
cs.CVcs.AIcs.LGarXiv:2604.09527v12026MixFlow: Mixed Source Distributions Improve Rectified Flows
Nazir Nayal, Christopher Wewer, Jan Eric Lenssen
cs.CVcs.LGarXiv:2604.09181v12026DMax: Aggressive Parallel Decoding for dLLMs
Zigeng Chen, Gongfan Fang, Xinyin Ma +2
cs.LGcs.AIarXiv:2604.08302v32026Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Video
Chanhyuk Choi, Taesoo Kim, Donggyu Lee +2
cs.CVcs.LGarXiv:2604.07786v22026Efficient RL Training for LLMs with Experience Replay
Charles Arnal, Vivien Cabannes, Taco Cohen +2
cs.LGarXiv:2604.08706v120263D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding
Makanjuola Ogunleye, Eman Abdelrahman, Ismini Lourentzou
cs.CVcs.AIcs.LGarXiv:2604.08645v12026EquiformerV3: Scaling Efficient, Expressive, and General SE(3)-Equivariant Graph Attention Transformers
Yi-Lun Liao, Alexander J. Hoffman, Sabrina C. Shen +3
cs.LGcs.AIphysics.comp-pharXiv:2604.09130v12026Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types
Hadas Orgad, Boyi Wei, Kaden Zheng +4
cs.CLcs.AIcs.LGarXiv:2604.09544v22026ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion
Lifeng Chen, Tianqi You, Hao Liu +8
cs.LGcs.AIeess.IVarXiv:2604.09450v22026Hierarchical SVG Tokenization: Learning Compact Visual Programs for Scalable Vector Graphics Modeling
Ximing Xing, Ziteng Xue, Zhenxi Li +8
cs.LGarXiv:2604.05072v22026Not All Denoising Steps Are Equal: Model Scheduling for Faster Masked Diffusion Language Models
Ivan Sedykh, Nikita Sorokin, Valentin Malykh
cs.LGcs.CLarXiv:2604.02340v22026A Temporally Augmented Graph Attention Network for Affordance Classification
Ami Chopra, Supriya Bordoloi, Shyamanta M. Hazarika
cs.LGcs.AIarXiv:2604.10149v12026Beyond Perception Errors: Semantic Fixation in Large Vision-Language Models
Md Tanvirul Alam
cs.CVcs.LGarXiv:2604.12119v12026ADD for Multi-Bit Image Watermarking
An Luo, Jie Ding
stat.MLcs.AIcs.LGarXiv:2604.11491v12026How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models
Gregory N. Frank
cs.CLcs.AIcs.LGarXiv:2604.04385v52026Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind
Hanqi Xiao, Vaidehi Patil, Zaid Khan +3
cs.CLcs.AIcs.LGarXiv:2604.11666v22026Diversity Without Fidelity: A Solver-Sampler Mismatch in Multi-Agent LLM Negotiation Simulation
Sandro Andric
cs.LGcs.AIcs.CYarXiv:2604.11840v32026Eliciting Medical Reasoning with Knowledge-enhanced Data Synthesis: A Semi-Supervised Reinforcement Learning Approach
Haolin Li, Shuyang Jiang, Ruipeng Zhang +3
cs.LGcs.CLarXiv:2604.11547v12026Continuous Adversarial Flow Models
Shanchuan Lin, Ceyuan Yang, Zhijie Lin +2
cs.LGcs.CVarXiv:2604.11521v12026Solving Physics Olympiad via Reinforcement Learning on Physics Simulators
Mihir Prabhudesai, Aryan Satpathy, Yangmin Li +6
cs.LGcs.AIcs.CVarXiv:2604.11805v12026IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs
Yuzhen Mao, Qitong Wang, Martin Ester +1
cs.LGcs.AIarXiv:2604.10539v12026Rethinking the Diffusion Model from a Langevin Perspective
Candi Zheng, Yuan Lan
cs.LGcs.AIcs.CVarXiv:2604.10465v12026PokeRL: Reinforcement Learning for Pokemon Red
Dheeraj Mudireddy, Sai Patibandla
cs.LGarXiv:2604.10812v12026Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration
Zhipeng Chen, Tao Qian, Wayne Xin Zhao +1
cs.LGcs.AIcs.CLarXiv:2604.11446v12026The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping
Yang Liu, Enxi Wang, Yufei Gao +6
cs.LGcs.AIcs.CLarXiv:2604.11297v12026LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
Junxiao Yang, Haoran Liu, Jinzhe Tu +9
cs.LGcs.AIcs.CLarXiv:2604.12710v22026ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
Fei Tang, Zhiqiong Lu, Boxuan Zhang +4
cs.LGcs.AIcs.CLarXiv:2604.11784v12026Towards Autonomous Mechanistic Reasoning in Virtual Cells
Yunhui Jang, Lu Zhu, Jake Fawkes +3
cs.LGcs.AIarXiv:2604.11661v32026UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy Optimization
Zhengxi Lu, Fei Tang, Guangyi Liu +8
cs.LGarXiv:2604.13822v12026VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
Andrei Atanov, Jesse Allardice, Roman Bachmann +6
cs.CVcs.LGarXiv:2604.12887v12026Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
Jiacheng Liu, Xiaohan Zhao, Xinyi Shang +1
cs.SEcs.AIcs.CLarXiv:2604.14228v22026Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding
Eun Woo Im, Dhruv Madhwal, Vivek Gupta
cs.LGarXiv:2604.13313v12026TIP: Token Importance in On-Policy Distillation
Yuanda Xu, Hejian Sang, Zhengze Zhou +3
cs.LGcs.AIarXiv:2604.14084v42026From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space
Yuqiao Tan, Minzheng Wang, Bo Liu +5
cs.LGcs.AIcs.CLarXiv:2604.14142v12026C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences
Akira Kawabata, Saku Sugawara
cs.CLcs.LGarXiv:2604.13618v12026Three-Phase Transformer
Mohammad R. Abu Ayyash
cs.CLcs.AIcs.LGarXiv:2604.14430v12026(1D) Ordered Tokens Enable Efficient Test-Time Search
Zhitong Gao, Parham Rezaei, Ali Cy +7
cs.CVcs.AIcs.LGarXiv:2604.15453v12026AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization
Genghan Zhang, Shaowei Zhu, Anjiang Wei +6
cs.LGcs.CLarXiv:2511.15915v22025Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Xiaohua Wang, Muzhao Tian, Yuqi Zeng +20
cs.LGarXiv:2604.13602v12026Where does output diversity collapse in post-training?
Constantinos Karouzos, Xingwei Tan, Nikolaos Aletras
cs.CLcs.AIcs.LGarXiv:2604.16027v12026GFT: From Imitation to Reward Fine-Tuning with Unbiased Group Advantages and Dynamic Coefficient Rectification
Wangjie Gan, Miao Pan, Linbo Xi +4
cs.AIcs.LGarXiv:2604.14258v32026An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental Learning
Quyen Tran, Hai Nguyen, Hoang Phan +6
cs.LGcs.CVarXiv:2211.16780v42022LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning
Bowen Ping, Zijun Chen, Tingfeng Hui +4
cs.LGcs.CLarXiv:2604.14922v12026PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research
Tingjia Miao, Wenkai Jin, Muhua Zhang +19
cs.LGcs.AIphysics.data-anarXiv:2604.15411v12026Maximal Brain Damage Without Data or Optimization: Disrupting Neural Networks via Sign-Bit Flips
Ido Galil, Moshe Kimhi, Ran El-Yaniv
cs.LGcs.AIcs.CVarXiv:2502.07408v22025PerturbRx: Learning Treatment-Conditioned Latent Transitions for Patient Drug Response Prediction
Yoshitaka Inoue, Minoh Jeong, Alfred Hero +2
q-bio.QMcs.LGarXiv:2608.21349v12026TwinTrack: Post-hoc Multi-Rater Calibration for Medical Image Segmentation
Tristan Kirscher, Alexandra Ertl, Klaus Maier-Hein +3
cs.LGarXiv:2604.15950v22026Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints
Xinge Liu, Terry Jingchen Zhang, Bernhard Schölkopf +2
cs.LGcs.AIarXiv:2604.15664v22026Structured Scaling of AI Discovery Across Diverse Scientific Domains
Haotian Ye, Haowei Lin, Jingyi Tang +30
cs.LGcs.AIarXiv:2604.19341v22026Advanced Linear Algebra with Applications - Part I (Numerical linear algebra for PDEs, machine learning, and data assimilation)
Victorita Dolean, Jemima Tabeart
math.NAcs.LGarXiv:2608.21234v12026The Exceedance Design Effect: Effective Sample Size for Thresholds under Clustering
Adam Noonan
stat.MLcs.LGarXiv:2608.21262v12026EasyVideoR1: Easier RL for Video Understanding
Chuanyu Qin, Chenxu Yang, Qingyi Si +6
cs.CVcs.LGarXiv:2604.16893v12026Truthful Calibration Measures for Sequential Prediction
Anagha Gokul, Jason Hartline, Lunjia Hu +2
cs.DScs.GTcs.LGarXiv:2608.21348v12026Test-Time Adaptation for EEG Foundation Models: A Systematic Study under Real-World Distribution Shifts
Gabriel Jason Lee, Jathurshan Pradeepkumar, Jimeng Sun
cs.LGcs.AIeess.SParXiv:2604.16926v22026Agents Explore but Agents Ignore: LLMs Lack Environmental Curiosity
Leon Engländer, Sophia Althammer, Ahmet Üstün +2
cs.CLcs.LGarXiv:2604.17609v12026Just Repair: A Minimal Denoising Network for Time Series Anomaly Detection
Kadir-Kaan Özer, René Ebeling, Markus Enzweiler
cs.LGcs.AIarXiv:2604.17388v32026On the Transferability of Agricultural Weed Detection Under Cross-Field Distribution Shift
Nikhilesh Prabhakar, Pranuthi Tenali, Wilfredo Abudeye Fernandez +5
cs.CVcs.LGarXiv:2608.21254v12026When Can LLMs Learn to Reason with Weak Supervision?
Salman Rahman, Jingyan Shen, Anna Mordvina +3
cs.LGcs.AIarXiv:2604.18574v12026MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
Shaden Alshammari, Kevin Wen, Abrar Zainal +5
cs.AIcs.DLcs.IRarXiv:2604.18584v22026