Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

19,681 to 19,740 of 20,192

  1. Looped World Models

    Hongyuan Adam Lu, Z. L. Victor Wei, Qun Zhang +28

    cs.LGcs.AIcs.CLarXiv:2606.18208v12026
  2. SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior

    Mingyue Cui, Linghui Shen, Xingyi Yang

    cs.LGcs.AIarXiv:2606.18322v12026
  3. EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts

    Minseo Kim, Minjae Lee, Seunghyuk Oh +7

    cs.LGarXiv:2606.18967v12026
  4. BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation

    Max Van Puyvelde, Ibrahim Gulluk, Wim Van Criekinge +1

    cs.AIcs.CVcs.LGarXiv:2606.19651v22026
  5. GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents

    Zhe Ren, Yibo Yang, Yimeng Chen +7

    cs.LGcs.CLarXiv:2606.18829v12026
  6. Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning

    Jiayu Yang, Chao Chen, Shengen Wu +6

    cs.LGcs.CLarXiv:2606.13106v12026
  7. Q-RAG: Long Context Multi-step Retrieval via Value-based Embedder Training

    Artyom Sorokin, Nazar Buzun, Alexander Anokhin +7

    cs.LGcs.IRarXiv:2511.07328v22025
  8. FedADB: Class Anchor-Driven Dual-Branch Federated Learning for Mitigating Forgetting

    Zhenyan Liu, Hua Zhang, Haoran Gao +6

    cs.CRcs.LGarXiv:2608.15310v12026
  9. MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents

    Lawrence Keunho Jang, Andrew Keunwoo Jang, Jing Yu Koh +1

    cs.LGcs.CLarXiv:2606.16748v12026
  10. Why Fine-Tuning Encourages Hallucinations and How to Fix It

    Guy Kaplan, Zorik Gekhman, Zhen Zhu +5

    cs.CLcs.AIcs.LGarXiv:2604.15574v12026
  11. Diffusion Templates: A Unified Plugin Framework for Controllable Diffusion

    Zhongjie Duan, Hong Zhang, Yingda Chen

    cs.LGcs.AIcs.CVarXiv:2604.24351v12026
  12. Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning

    Chengshuai Shi, Wenzhe Li, Xinran Liang +10

    cs.LGcs.AIcs.CLarXiv:2605.00347v12026
  13. PianoCoRe: Combined and Refined Piano MIDI Dataset

    Ilya Borovik

    cs.SDcs.LGarXiv:2605.06627v12026
  14. CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing

    Cheng Qian, Hyeonjeong Ha, Jiayu Liu +10

    cs.AIcs.CLcs.LGarXiv:2605.02910v22026
  15. MARBLE: Multi-Aspect Reward Balance for Diffusion RL

    Canyu Zhao, Hao Chen, Yunze Tong +3

    cs.CVcs.LGarXiv:2605.06507v12026
  16. Continuous Quantum Feedback Control via Kraus-Parameterized Belief Reinforcement Learning

    Priyanshi Singh, Krishna Bhatia

    quant-phcs.LGarXiv:2608.15715v12026
  17. SCALE: State-Calibrated Latent Embeddings for JEPA Planning in the Right Geometry

    Jiaming Hu, Yan Zheng, Tian Wang

    cs.LGarXiv:2608.16287v12026
  18. How Many Samples Are Needed to Determine Causal Direction? Sharp Minimax Bounds for Bivariate LiNGAM

    Jikai Jin

    math.STcs.LGecon.EMarXiv:2608.15840v12026
  19. FeatCal: Feature Calibration for Post-Merging Models

    Yanggan Gu, Shuo Cai, Zihao Wang +7

    cs.LGcs.AIarXiv:2605.13030v12026
  20. Conformal Agent Error Attribution

    Naihe Feng, Yi Sui, Shiyi Hou +2

    cs.LGcs.MAarXiv:2605.06788v12026
  21. SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training

    Shengkun Tang, Zekun Wang, Bo Zheng +7

    cs.LGcs.AIcs.CLarXiv:2605.08738v22026
  22. TD3B: Transition-Directed Discrete Diffusion for Allosteric Binder Generation

    Hanqun Cao, Aastha Pal, Sophia Tang +4

    q-bio.BMcs.LGarXiv:2605.09810v12026
  23. Learning Visual Feature-Based World Models via Residual Latent Action

    Xinyu Zhang, Zhengtong Xu, Yutian Tao +3

    cs.CVcs.AIcs.LGarXiv:2605.07079v12026
  24. Normalizing Trajectory Models

    Jiatao Gu, Tianrong Chen, Ying Shen +3

    cs.CVcs.LGarXiv:2605.08078v22026
  25. F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking

    Rohan Surana, Gagan Mundada, Junda Wu +9

    cs.LGarXiv:2605.12995v12026
  26. Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining

    Weimin Xiong, Shuhao Gu, Bowen Ye +5

    cs.CLcs.AIcs.CVarXiv:2605.14747v12026
  27. Inferential Evaluation of Surrogate-Derived Models under Covariate Shift

    Longtian Shi, Molei Liu, Doudou Zhou

    stat.MLcs.LGstat.AParXiv:2608.15783v12026
  28. A Banach-Space Theory of Markovian Halpern Iteration for Non-Expansive Maps

    Ege C. Kaya, Arda Fazla, M. Berk Sahin +1

    cs.LGmath.OCarXiv:2608.15966v12026
  29. Decorrelation Is Not Complementarity: Skill, Not Lineage, Governs Trusted-Monitor Ensembles

    Anik Jha

    cs.CRcs.LGarXiv:2608.16190v12026
  30. SkillComposer: Learning Reusable Skills for Natural-Language Robot Programming

    John Woods, Hasti Seifi

    cs.ROcs.CLcs.LGarXiv:2608.14944v12026
  31. Invariant Pretraining for Robust Code Representations

    Yifeng He, Yundi Xu, Christopher Castro Gaw Gonzalo +2

    cs.LGcs.AIcs.SEarXiv:2608.15412v12026
  32. Data-Driven Reconstruction of Spatially Resolved Electron and Ion Energy Distributions from Macroscopic Plasma Quantities with Deep Neural Networks

    Libin Varghese, Kaushik Prajapati, Bhaskar Chaudhury

    physics.plasm-phcs.LGphysics.comp-pharXiv:2608.16519v12026
  33. Self-Routed Tensor Adapters for Parameter-Efficient Universal Visual Adaptation

    Suraj Yadav

    cs.CVcs.LGarXiv:2608.16384v12026
  34. Whose Gold? Annotator-Pool Disagreement Is Large at the Item Level, and Hidden by Small Leaderboards

    Anik Jha

    cs.CLcs.LGarXiv:2608.15980v12026
  35. Data-Efficient and Interpretable Classification of Circulating Tumor Cell Phenotypes in Microfluidic Devices via Deep Learning

    Serena Su, Yifan Wang, Senwei Liang

    cs.LGarXiv:2608.16870v12026
  36. AutoSR: Automatic Symbolic Regression by Searching Research States

    Kejia Zhang, Youran Sun, Xinyu Ren +2

    cs.SCcs.AIcs.LGarXiv:2608.16876v12026
  37. MiNO: Cotangent-bundle propagator learning for PDEs

    Gnankan Landry Regis N'guessan, Bum Jun Kim

    cs.LGcs.CEmath.NAarXiv:2608.15187v12026
  38. Self-Supervised Auxiliary Task Discovery for Stable Reinforcement Learning in Stock Trading

    Arishi Orra, Himanshu Choudhary, Manoj Thakur

    cs.LGq-fin.CPstat.MLarXiv:2608.15841v12026
  39. Learning Stock Trading Policies via Barycenter-Based Adversarial Inverse Reinforcement Learning

    Arishi Orra, Himanshu Choudhary, Manoj Thakur

    cs.LGstat.MLarXiv:2608.15770v12026
  40. SchurQuant: Groupwise Discrete Optimization for Layer-Wise LLM Quantization

    Gunjun Lee, Sehwan Son, Younjoo Lee +2

    cs.LGarXiv:2608.15567v12026
  41. Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation

    Cedar Site Bai, Duanshun Li, Zhenyu Liao +6

    cs.IRcs.AIcs.CLarXiv:2608.15949v12026
  42. FirstDiff: One-Step Diffusion-Based Anomaly Detection for Multivariate Time Series via Initial Noise Prediction

    Ali Boudaghi, Alireza Nemati, Hadi Zare

    cs.LGcs.AIstat.MLarXiv:2608.15727v12026
  43. When Is Shallow Enough? Adaptive Split Federated Learning with Client-Specific Sufficiency Estimation

    Wenhao Yuan, Chenchen Lin, Wenhao Hu +4

    cs.DCcs.AIcs.LGarXiv:2608.15639v12026
  44. TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning

    Heming Zou, Qi Wang, Yun Qu +9

    cs.LGcs.AIcs.CLarXiv:2606.11119v12026
  45. Skip a Layer or Loop It? Learning Program-of-Layers in LLMs

    Ziyue Li, Yang Li, Tianyi Zhou

    cs.LGarXiv:2606.06574v22026
  46. $S^3$: A Smooth Simulation Surrogate for Optimizing Discrete Abstractions of Dynamical Systems

    Jordan Peper, James Mathias Gast, Vignesh Nanduri +3

    eess.SYcs.LGarXiv:2608.15920v12026
  47. The Null Token Knows: Reducing Message-Free Hallucination in ASR and NMT

    Kirill Borodin, Vasiliy Kudryavtsev, Ivan Viakhirev +1

    cs.CLcs.LGcs.SDarXiv:2608.15940v22026
  48. Beat the Counter First: A Baseline for Temporal-Graph Anomaly Detectors

    Omair Shafi Ahmed, Zohair Shafi

    cs.LGarXiv:2608.15965v12026
  49. TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity

    Armin Steinhauser

    cs.LGcs.AIarXiv:2608.15767v12026
  50. When Should Models Change Their Minds? Contextual Belief Management in Large Language Models

    Haoming Xu, Weihong Xu, Zongrui Li +6

    cs.AIcs.CLcs.LGarXiv:2605.30219v12026
  51. Machine Learning Approaches to Decoding Topological Quantum Codes

    Changwon Lee, Tak Hur, Jeongwoo Jae +1

    quant-phcs.LGarXiv:2608.15760v12026
  52. PERO: Efficient Robust Post-Training Foundation Models for Encrypted Traffic Classification

    Wumei Du, Jiarong Wen, Kaiyu Zhang +5

    cs.LGstat.MLarXiv:2608.15504v12026
  53. Feasible and Novel Synthetic Population Generation with Tabular and Sequential Travel Attributes

    Farbod Abbasi, Zachary Patterson, Bilal Farooq

    cs.LGcs.AIarXiv:2608.15867v12026
  54. DeltaLog: Deferred Materialization of Recurrent States for Linear Attention Decoding

    Junqing Lin, Jingwei Sun, Guangzhong Sun

    cs.DCcs.LGarXiv:2608.15533v12026
  55. UniFed-VLM: Federated Instruction Tuning for Vision-Language Models with Multiple Heterogeneity

    Pengyu Wang, Baochen Xiong, Xiaoshan Yang +4

    cs.LGarXiv:2608.15516v12026
  56. Domain-Agnostic Neural Topic Modeling with Contextual Token-Level Semantic Graph Representation

    Seung-Won Seo, Won Ik Cho, Yongmin Yoo

    cs.CLcs.LGarXiv:2608.16269v12026
  57. ComboStoc: Combinatorial Stochasticity for Diffusion Generative Models

    Rui Xu, Jiepeng Wang, Hao Pan +6

    cs.LGcs.AIcs.CVarXiv:2405.13729v32024
  58. Co-Evolving Policy Distillation

    Naibin Gu, Chenxu Yang, Qingyi Si +7

    cs.LGarXiv:2604.27083v12026
  59. Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning

    Xuekang Wang, Zhuoyuan Hao, Shuo Hou +3

    cs.LGcs.AIcs.CLarXiv:2606.04923v22026
  60. Flash-WAM: Modality-Aware Distillation for World Action Models

    Arman Akbari, Ci Zhang, Arash Akbari +6

    cs.LGcs.CVcs.ROarXiv:2606.05254v12026