Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
19,681 to 19,740 of 20,192
Looped World Models
Hongyuan Adam Lu, Z. L. Victor Wei, Qun Zhang +28
cs.LGcs.AIcs.CLarXiv:2606.18208v12026SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior
Mingyue Cui, Linghui Shen, Xingyi Yang
cs.LGcs.AIarXiv:2606.18322v12026EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts
Minseo Kim, Minjae Lee, Seunghyuk Oh +7
cs.LGarXiv:2606.18967v12026BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation
Max Van Puyvelde, Ibrahim Gulluk, Wim Van Criekinge +1
cs.AIcs.CVcs.LGarXiv:2606.19651v22026GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents
Zhe Ren, Yibo Yang, Yimeng Chen +7
cs.LGcs.CLarXiv:2606.18829v12026Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning
Jiayu Yang, Chao Chen, Shengen Wu +6
cs.LGcs.CLarXiv:2606.13106v12026Q-RAG: Long Context Multi-step Retrieval via Value-based Embedder Training
Artyom Sorokin, Nazar Buzun, Alexander Anokhin +7
cs.LGcs.IRarXiv:2511.07328v22025FedADB: Class Anchor-Driven Dual-Branch Federated Learning for Mitigating Forgetting
Zhenyan Liu, Hua Zhang, Haoran Gao +6
cs.CRcs.LGarXiv:2608.15310v12026MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents
Lawrence Keunho Jang, Andrew Keunwoo Jang, Jing Yu Koh +1
cs.LGcs.CLarXiv:2606.16748v12026Why Fine-Tuning Encourages Hallucinations and How to Fix It
Guy Kaplan, Zorik Gekhman, Zhen Zhu +5
cs.CLcs.AIcs.LGarXiv:2604.15574v12026Diffusion Templates: A Unified Plugin Framework for Controllable Diffusion
Zhongjie Duan, Hong Zhang, Yingda Chen
cs.LGcs.AIcs.CVarXiv:2604.24351v12026Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning
Chengshuai Shi, Wenzhe Li, Xinran Liang +10
cs.LGcs.AIcs.CLarXiv:2605.00347v12026PianoCoRe: Combined and Refined Piano MIDI Dataset
Ilya Borovik
cs.SDcs.LGarXiv:2605.06627v12026CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing
Cheng Qian, Hyeonjeong Ha, Jiayu Liu +10
cs.AIcs.CLcs.LGarXiv:2605.02910v22026MARBLE: Multi-Aspect Reward Balance for Diffusion RL
Canyu Zhao, Hao Chen, Yunze Tong +3
cs.CVcs.LGarXiv:2605.06507v12026Continuous Quantum Feedback Control via Kraus-Parameterized Belief Reinforcement Learning
Priyanshi Singh, Krishna Bhatia
quant-phcs.LGarXiv:2608.15715v12026SCALE: State-Calibrated Latent Embeddings for JEPA Planning in the Right Geometry
Jiaming Hu, Yan Zheng, Tian Wang
cs.LGarXiv:2608.16287v12026How Many Samples Are Needed to Determine Causal Direction? Sharp Minimax Bounds for Bivariate LiNGAM
Jikai Jin
math.STcs.LGecon.EMarXiv:2608.15840v12026FeatCal: Feature Calibration for Post-Merging Models
Yanggan Gu, Shuo Cai, Zihao Wang +7
cs.LGcs.AIarXiv:2605.13030v12026Conformal Agent Error Attribution
Naihe Feng, Yi Sui, Shiyi Hou +2
cs.LGcs.MAarXiv:2605.06788v12026SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training
Shengkun Tang, Zekun Wang, Bo Zheng +7
cs.LGcs.AIcs.CLarXiv:2605.08738v22026TD3B: Transition-Directed Discrete Diffusion for Allosteric Binder Generation
Hanqun Cao, Aastha Pal, Sophia Tang +4
q-bio.BMcs.LGarXiv:2605.09810v12026Learning Visual Feature-Based World Models via Residual Latent Action
Xinyu Zhang, Zhengtong Xu, Yutian Tao +3
cs.CVcs.AIcs.LGarXiv:2605.07079v12026Normalizing Trajectory Models
Jiatao Gu, Tianrong Chen, Ying Shen +3
cs.CVcs.LGarXiv:2605.08078v22026F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking
Rohan Surana, Gagan Mundada, Junda Wu +9
cs.LGarXiv:2605.12995v12026Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining
Weimin Xiong, Shuhao Gu, Bowen Ye +5
cs.CLcs.AIcs.CVarXiv:2605.14747v12026Inferential Evaluation of Surrogate-Derived Models under Covariate Shift
Longtian Shi, Molei Liu, Doudou Zhou
stat.MLcs.LGstat.AParXiv:2608.15783v12026A Banach-Space Theory of Markovian Halpern Iteration for Non-Expansive Maps
Ege C. Kaya, Arda Fazla, M. Berk Sahin +1
cs.LGmath.OCarXiv:2608.15966v12026Decorrelation Is Not Complementarity: Skill, Not Lineage, Governs Trusted-Monitor Ensembles
Anik Jha
cs.CRcs.LGarXiv:2608.16190v12026SkillComposer: Learning Reusable Skills for Natural-Language Robot Programming
John Woods, Hasti Seifi
cs.ROcs.CLcs.LGarXiv:2608.14944v12026Invariant Pretraining for Robust Code Representations
Yifeng He, Yundi Xu, Christopher Castro Gaw Gonzalo +2
cs.LGcs.AIcs.SEarXiv:2608.15412v12026Data-Driven Reconstruction of Spatially Resolved Electron and Ion Energy Distributions from Macroscopic Plasma Quantities with Deep Neural Networks
Libin Varghese, Kaushik Prajapati, Bhaskar Chaudhury
physics.plasm-phcs.LGphysics.comp-pharXiv:2608.16519v12026Self-Routed Tensor Adapters for Parameter-Efficient Universal Visual Adaptation
Suraj Yadav
cs.CVcs.LGarXiv:2608.16384v12026Whose Gold? Annotator-Pool Disagreement Is Large at the Item Level, and Hidden by Small Leaderboards
Anik Jha
cs.CLcs.LGarXiv:2608.15980v12026Data-Efficient and Interpretable Classification of Circulating Tumor Cell Phenotypes in Microfluidic Devices via Deep Learning
Serena Su, Yifan Wang, Senwei Liang
cs.LGarXiv:2608.16870v12026AutoSR: Automatic Symbolic Regression by Searching Research States
Kejia Zhang, Youran Sun, Xinyu Ren +2
cs.SCcs.AIcs.LGarXiv:2608.16876v12026MiNO: Cotangent-bundle propagator learning for PDEs
Gnankan Landry Regis N'guessan, Bum Jun Kim
cs.LGcs.CEmath.NAarXiv:2608.15187v12026Self-Supervised Auxiliary Task Discovery for Stable Reinforcement Learning in Stock Trading
Arishi Orra, Himanshu Choudhary, Manoj Thakur
cs.LGq-fin.CPstat.MLarXiv:2608.15841v12026Learning Stock Trading Policies via Barycenter-Based Adversarial Inverse Reinforcement Learning
Arishi Orra, Himanshu Choudhary, Manoj Thakur
cs.LGstat.MLarXiv:2608.15770v12026SchurQuant: Groupwise Discrete Optimization for Layer-Wise LLM Quantization
Gunjun Lee, Sehwan Son, Younjoo Lee +2
cs.LGarXiv:2608.15567v12026Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation
Cedar Site Bai, Duanshun Li, Zhenyu Liao +6
cs.IRcs.AIcs.CLarXiv:2608.15949v12026FirstDiff: One-Step Diffusion-Based Anomaly Detection for Multivariate Time Series via Initial Noise Prediction
Ali Boudaghi, Alireza Nemati, Hadi Zare
cs.LGcs.AIstat.MLarXiv:2608.15727v12026When Is Shallow Enough? Adaptive Split Federated Learning with Client-Specific Sufficiency Estimation
Wenhao Yuan, Chenchen Lin, Wenhao Hu +4
cs.DCcs.AIcs.LGarXiv:2608.15639v12026TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning
Heming Zou, Qi Wang, Yun Qu +9
cs.LGcs.AIcs.CLarXiv:2606.11119v12026Skip a Layer or Loop It? Learning Program-of-Layers in LLMs
Ziyue Li, Yang Li, Tianyi Zhou
cs.LGarXiv:2606.06574v22026$S^3$: A Smooth Simulation Surrogate for Optimizing Discrete Abstractions of Dynamical Systems
Jordan Peper, James Mathias Gast, Vignesh Nanduri +3
eess.SYcs.LGarXiv:2608.15920v12026The Null Token Knows: Reducing Message-Free Hallucination in ASR and NMT
Kirill Borodin, Vasiliy Kudryavtsev, Ivan Viakhirev +1
cs.CLcs.LGcs.SDarXiv:2608.15940v22026Beat the Counter First: A Baseline for Temporal-Graph Anomaly Detectors
Omair Shafi Ahmed, Zohair Shafi
cs.LGarXiv:2608.15965v12026TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity
Armin Steinhauser
cs.LGcs.AIarXiv:2608.15767v12026When Should Models Change Their Minds? Contextual Belief Management in Large Language Models
Haoming Xu, Weihong Xu, Zongrui Li +6
cs.AIcs.CLcs.LGarXiv:2605.30219v12026Machine Learning Approaches to Decoding Topological Quantum Codes
Changwon Lee, Tak Hur, Jeongwoo Jae +1
quant-phcs.LGarXiv:2608.15760v12026PERO: Efficient Robust Post-Training Foundation Models for Encrypted Traffic Classification
Wumei Du, Jiarong Wen, Kaiyu Zhang +5
cs.LGstat.MLarXiv:2608.15504v12026Feasible and Novel Synthetic Population Generation with Tabular and Sequential Travel Attributes
Farbod Abbasi, Zachary Patterson, Bilal Farooq
cs.LGcs.AIarXiv:2608.15867v12026DeltaLog: Deferred Materialization of Recurrent States for Linear Attention Decoding
Junqing Lin, Jingwei Sun, Guangzhong Sun
cs.DCcs.LGarXiv:2608.15533v12026UniFed-VLM: Federated Instruction Tuning for Vision-Language Models with Multiple Heterogeneity
Pengyu Wang, Baochen Xiong, Xiaoshan Yang +4
cs.LGarXiv:2608.15516v12026Domain-Agnostic Neural Topic Modeling with Contextual Token-Level Semantic Graph Representation
Seung-Won Seo, Won Ik Cho, Yongmin Yoo
cs.CLcs.LGarXiv:2608.16269v12026ComboStoc: Combinatorial Stochasticity for Diffusion Generative Models
Rui Xu, Jiepeng Wang, Hao Pan +6
cs.LGcs.AIcs.CVarXiv:2405.13729v32024Co-Evolving Policy Distillation
Naibin Gu, Chenxu Yang, Qingyi Si +7
cs.LGarXiv:2604.27083v12026Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning
Xuekang Wang, Zhuoyuan Hao, Shuo Hou +3
cs.LGcs.AIcs.CLarXiv:2606.04923v22026Flash-WAM: Modality-Aware Distillation for World Action Models
Arman Akbari, Ci Zhang, Arash Akbari +6
cs.LGcs.CVcs.ROarXiv:2606.05254v12026