Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

19,741 to 19,800 of 19,970

  1. MuScriptor: An Open Model for Multi-Instrument Music Transcription

    Simon Rouard, Michael Krause, Axel Roebel +2

    cs.SDcs.LGarXiv:2607.08168v22026
  2. On Locality and Length Generalization in Visual Reasoning

    Pulkit Madan, Sanjay Haresh, Reza Ebrahimi +3

    cs.CVcs.AIcs.LGarXiv:2607.09061v12026
  3. Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes

    Minh-Quan Le, Armand Comas, Alexandros Lattas +7

    cs.LGarXiv:2607.13188v12026
  4. Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel

    Sietse Schelpe

    cs.CLcs.AIcs.LGarXiv:2607.14431v12026
  5. Understanding Reasoning from Pretraining to Post-Training

    Jingyan Shen, Ang Li, Salman Rahman +4

    cs.LGcs.AIcs.CLarXiv:2607.16097v22026
  6. Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

    Weiwen Xu, Jia Liu, Hou Pong Chan +4

    cs.LGcs.AIcs.CLarXiv:2607.14614v12026
  7. When Does Muon Help Agentic Reinforcement Learning?

    Kai Ruan, Jinghao Lin, Zihe Huang +4

    cs.LGcs.AIarXiv:2607.16169v42026
  8. Generated Contents Enrichment

    Mahdi Naseri, Jiayan Qiu, Zhou Wang

    cs.CVcs.LGarXiv:2405.03650v42024
  9. PixCon: Clean-Positive Contrastive Learning for Foundation-Model Semi-Supervised Segmentation

    Ebenezer Tarubinga

    cs.CVcs.AIcs.LGarXiv:2607.03068v12026
  10. Exploring the Design Space of Reward Backpropagation for Flow Matching

    Ruoyu Wang, Boye Niu, Xiangxin Zhou +3

    cs.LGarXiv:2606.11075v12026
  11. RL-Index: Reinforcement Learning for Retrieval Index Reasoning

    Yongjia Lei, Nedim Lipka, Zhisheng Qi +7

    cs.IRcs.AIcs.LGarXiv:2606.16316v22026
  12. Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation

    Zhilin Huang, Hang Gao, Ziqiang Dong +6

    cs.LGarXiv:2606.18844v12026
  13. Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining

    Michael K. Chen, Xikun Zhang, Fan Bai +2

    cs.LGcs.AIcs.CLarXiv:2606.16246v22026
  14. Improving Text-to-Music Generation with Human Preference Rewards

    Yonghyun Kim, Junwon Lee, Haiwen Xia +2

    cs.SDcs.AIcs.LGarXiv:2606.21670v12026
  15. RoPE-Aware Bit Allocation for KV-Cache Quantization

    Fengfeng Liang, Yuechen Zhang, Jiaya Jia

    cs.LGcs.CLarXiv:2606.24033v12026
  16. Fast LeWorldModel

    Yuntian Gao, Xiangyu Xu

    cs.LGcs.CVcs.ROarXiv:2606.26217v12026
  17. Hallucination in World Models is Predictable and Preventable

    Nicklas Hansen, Xiaolong Wang

    cs.LGcs.CVcs.ROarXiv:2606.27326v12026
  18. Taste-aware music retrieval from audio embeddings

    Matteo Spanio, Antonio Rodà

    cs.SDcs.IRcs.LGarXiv:2607.03296v12026
  19. ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation

    Kaifeng Zhao, Mathis Petrovich, Haotian Zhang +3

    cs.GRcs.CVcs.LGarXiv:2607.08741v12026
  20. A Sovereign, Open-Source Foundation Model for German and English

    Soofi-Team, :, Benedikt Droste +30

    cs.CLcs.AIcs.LGarXiv:2607.09424v32026
  21. PAST-TIDE: Prototype-Anchored Statement Tuning with Topic-Invariant Normalization for Stance Detection

    Md. Shakhoyat Rahman Shujon, MD Jahid Hasan Jim, Md. Milon Islam +2

    cs.CLcs.LGarXiv:2607.04690v12026
  22. Simplified Sparse Attention via Gist Tokens

    Yuzhen Mao, Michael Y. Li, Emily B. Fox

    cs.LGarXiv:2604.20920v22026
  23. Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction

    Chenguang Wang, Ming Li, Xinyue Zeng +4

    cs.CLcs.AIcs.CYarXiv:2606.28186v32026
  24. SPEAR: A Simulator for Photorealistic Embodied AI Research

    Mike Roberts, Renhan Wang, Rushikesh Zawar +10

    cs.CVcs.AIcs.GRarXiv:2607.06701v12026
  25. Wan-Streamer v0.2: Higher Resolution, Same Latency

    Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +23

    cs.CVcs.AIcs.GRarXiv:2607.04443v32026
  26. SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History

    Zhiwei Li, Yong Hu

    cs.LGarXiv:2606.08671v32026
  27. LLM Program Optimization via Retrieval Augmented Search

    Sagnik Anupam, Alexander Shypula, Osbert Bastani

    cs.LGarXiv:2501.18916v22025
  28. When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling

    Yong Yi Bay, Kathleen A. Yearick

    cs.LGcs.AIcs.CLarXiv:2606.28661v12026
  29. Parameter-Efficient Quantum-Inspired Fast Weight Programmers for Traffic-Matrix Forecasting

    Kuo-Chung Peng, Jiun-Cheng Jiang, Chun-Hua Lin +3

    quant-phcs.AIcs.LGarXiv:2606.27821v12026
  30. SciForma: Structure-Faithful Generation of Scientific Diagrams

    Yuxuan Luo, Peng Zhang, Xinjie Zhang +3

    cs.CVcs.GRcs.LGarXiv:2607.18091v12026
  31. Hard Cases, Bad Labels: Testing Error Exposure and Error Location in Uncertainty Sampling Under Bounded Label Noise

    John Myron Uy

    cs.LGarXiv:2608.13601v12026
  32. No Universal Signal Predicts Sample-Level LLM Regression under Version Updates

    Jia Sheng, Yiwei Lu

    cs.AIcs.CLcs.LGarXiv:2608.13607v12026
  33. HI-MeshGraphNets: Efficient and Accurate Mesh-based Physics Learning with Hierarchical Multi-scale Graph Neural Networks

    SiHun Lee, Dong-Hyuk Park, Taesoo Bang +1

    cs.LGarXiv:2608.13827v12026
  34. Trajectory Dynamics in Self-Supervised Learning Latent Space for Audio Deepfake Detection

    Tomás Andrade Weber

    eess.AScs.LGcs.SDarXiv:2608.13817v12026
  35. Language-Specific Gaps in AI Safety Training Datasets

    Chialuka Prisca-Mary Onuoha, Bright Etornam Sunu, Rashidat Sikiru

    cs.CYcs.LGarXiv:2608.13695v12026
  36. The Integer Alibi: Localizing Cross-Kernel Divergence in INT8-Quantized LLM Inference

    Teng-Ruei Chen

    cs.LGarXiv:2608.13756v12026
  37. Does ISO-Grounded NFR Specification Improve LLM Code Generation? A Comparison of Rich and Structured Interventions against a Natural-Language Baseline

    Joào Pedro Monteiro Pereira, Vinicius Cardoso Garcia

    cs.SEcs.AIcs.LGarXiv:2608.13742v12026
  38. A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure

    Dekun Yang

    cs.AIcs.CLcs.LGarXiv:2608.13626v12026
  39. High-dimensional nonparametric changepoint detection via low-rank degree-two density projection

    Guoqing Zhang, Zhaixin Chen

    cs.LGstat.MLarXiv:2608.13922v12026
  40. How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF

    Venkata Naga Sai Vishnu Rohit Pulipaka, Anish Katta, Deva Rohit Reddy Peddireddy

    cs.LGarXiv:2607.19712v12026
  41. FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

    Krish Agarwal, Zhuoming Chen, Yanyuan Qin +3

    cs.LGarXiv:2607.18171v22026
  42. Differentiable Logic Gate Networks for Low-Latency EEG Classification on Edge Devices

    Shyamal Y. Dharia, Stephen D. Smith, Camilo E. Valderrama

    cs.LGcs.AIarXiv:2607.18149v12026
  43. Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift

    Zitong Huang, Gustavo Lucas Carvalho, Deqing Fu +1

    cs.CLcs.LGarXiv:2607.17524v12026
  44. DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification

    Yuhang Wang, Lingyao Li, Hao Zhou

    cs.LGarXiv:2607.23822v12026
  45. AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition

    Yuqi Li, Yi-Cheng Lin, Xianglong Wang +5

    cs.LGarXiv:2607.25289v12026
  46. Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization

    Alexey Gavrilov, Alan-Barsag Gazzaev, Sergey Muravyov

    cs.CLcs.LGarXiv:2607.24176v12026
  47. A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever

    Sietse Schelpe

    cs.CLcs.AIcs.IRarXiv:2607.23806v12026
  48. Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations

    Pere Martra, Eugenio Martínez Cámara, Alfonso Ureña López

    cs.CLcs.CYcs.LGarXiv:2607.28319v12026
  49. H$^2$SD: Hybrid Hindsight Self-Distillation

    Qiye Cai, Yichuan Ma, Peiji Li +6

    cs.LGcs.CLarXiv:2607.18955v42026
  50. WorldDiT: A Unified Diffusion Architecture for World and Action Modeling

    Sen Wang, R. Gnana Praveen, Bidhan Roy +1

    cs.LGcs.ROarXiv:2607.23909v22026
  51. Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training

    Nuemaan Malik

    cs.LGcs.AIarXiv:2607.19058v22026
  52. OPERA: Offline Policy-guided Expert Routing and Adaptation for Universal Biomedical Image Analysis

    Zihan Li, Feiyang Liu, Dandan Shan +2

    cs.CVcs.AIcs.LGarXiv:2607.25108v12026
  53. LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks

    Tianzhu Ye, Li Dong, Guanheng Chen +4

    cs.LGcs.CLarXiv:2607.18110v12026
  54. Distilled Reinforcement Learning for LLM Post-training

    Chen Wang, Zhaochun Li, Jionghao Bai +4

    cs.LGcs.AIarXiv:2607.17247v12026
  55. CriPO: Enhancing Rubric-based RL via Self-Distillation

    Mingxuan Xia, Yuhang Yang, Chao Ye +7

    cs.LGcs.AIarXiv:2607.18082v32026
  56. WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting

    Zhaokai Wang, Tianlin Gui, Jiayuan Rao +3

    cs.AIcs.CLcs.LGarXiv:2607.18084v12026
  57. DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations

    Jiazhen Jiang, Boxi Cao, Lingyong Yan +6

    cs.AIcs.CLcs.LGarXiv:2607.19865v12026
  58. Codifying the Judge: Scalable Evaluation via Program Distillation

    Tzu-Heng Huang, Shengqi Qiu, Frederic Sala

    cs.AIcs.LGarXiv:2607.22561v12026
  59. Towards Robust Reinforcement Learning for Small-Scale Language Model Agents

    Md Rezwanul Haque, Md. Milon Islam, Fakhri Karray

    cs.AIcs.CLcs.LGarXiv:2607.25091v12026
  60. DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

    Hao Liang, Qifeng Cai, Yibo Lin +11

    cs.LGcs.CLarXiv:2607.20465v12026