Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
19,741 to 19,800 of 19,970
MuScriptor: An Open Model for Multi-Instrument Music Transcription
Simon Rouard, Michael Krause, Axel Roebel +2
cs.SDcs.LGarXiv:2607.08168v22026On Locality and Length Generalization in Visual Reasoning
Pulkit Madan, Sanjay Haresh, Reza Ebrahimi +3
cs.CVcs.AIcs.LGarXiv:2607.09061v12026Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes
Minh-Quan Le, Armand Comas, Alexandros Lattas +7
cs.LGarXiv:2607.13188v12026Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel
Sietse Schelpe
cs.CLcs.AIcs.LGarXiv:2607.14431v12026Understanding Reasoning from Pretraining to Post-Training
Jingyan Shen, Ang Li, Salman Rahman +4
cs.LGcs.AIcs.CLarXiv:2607.16097v22026Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization
Weiwen Xu, Jia Liu, Hou Pong Chan +4
cs.LGcs.AIcs.CLarXiv:2607.14614v12026When Does Muon Help Agentic Reinforcement Learning?
Kai Ruan, Jinghao Lin, Zihe Huang +4
cs.LGcs.AIarXiv:2607.16169v42026Generated Contents Enrichment
Mahdi Naseri, Jiayan Qiu, Zhou Wang
cs.CVcs.LGarXiv:2405.03650v42024PixCon: Clean-Positive Contrastive Learning for Foundation-Model Semi-Supervised Segmentation
Ebenezer Tarubinga
cs.CVcs.AIcs.LGarXiv:2607.03068v12026Exploring the Design Space of Reward Backpropagation for Flow Matching
Ruoyu Wang, Boye Niu, Xiangxin Zhou +3
cs.LGarXiv:2606.11075v12026RL-Index: Reinforcement Learning for Retrieval Index Reasoning
Yongjia Lei, Nedim Lipka, Zhisheng Qi +7
cs.IRcs.AIcs.LGarXiv:2606.16316v22026Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation
Zhilin Huang, Hang Gao, Ziqiang Dong +6
cs.LGarXiv:2606.18844v12026Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining
Michael K. Chen, Xikun Zhang, Fan Bai +2
cs.LGcs.AIcs.CLarXiv:2606.16246v22026Improving Text-to-Music Generation with Human Preference Rewards
Yonghyun Kim, Junwon Lee, Haiwen Xia +2
cs.SDcs.AIcs.LGarXiv:2606.21670v12026RoPE-Aware Bit Allocation for KV-Cache Quantization
Fengfeng Liang, Yuechen Zhang, Jiaya Jia
cs.LGcs.CLarXiv:2606.24033v12026Fast LeWorldModel
Yuntian Gao, Xiangyu Xu
cs.LGcs.CVcs.ROarXiv:2606.26217v12026Hallucination in World Models is Predictable and Preventable
Nicklas Hansen, Xiaolong Wang
cs.LGcs.CVcs.ROarXiv:2606.27326v12026Taste-aware music retrieval from audio embeddings
Matteo Spanio, Antonio Rodà
cs.SDcs.IRcs.LGarXiv:2607.03296v12026ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation
Kaifeng Zhao, Mathis Petrovich, Haotian Zhang +3
cs.GRcs.CVcs.LGarXiv:2607.08741v12026A Sovereign, Open-Source Foundation Model for German and English
Soofi-Team, :, Benedikt Droste +30
cs.CLcs.AIcs.LGarXiv:2607.09424v32026PAST-TIDE: Prototype-Anchored Statement Tuning with Topic-Invariant Normalization for Stance Detection
Md. Shakhoyat Rahman Shujon, MD Jahid Hasan Jim, Md. Milon Islam +2
cs.CLcs.LGarXiv:2607.04690v12026Simplified Sparse Attention via Gist Tokens
Yuzhen Mao, Michael Y. Li, Emily B. Fox
cs.LGarXiv:2604.20920v22026Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction
Chenguang Wang, Ming Li, Xinyue Zeng +4
cs.CLcs.AIcs.CYarXiv:2606.28186v32026SPEAR: A Simulator for Photorealistic Embodied AI Research
Mike Roberts, Renhan Wang, Rushikesh Zawar +10
cs.CVcs.AIcs.GRarXiv:2607.06701v12026Wan-Streamer v0.2: Higher Resolution, Same Latency
Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +23
cs.CVcs.AIcs.GRarXiv:2607.04443v32026SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History
Zhiwei Li, Yong Hu
cs.LGarXiv:2606.08671v32026LLM Program Optimization via Retrieval Augmented Search
Sagnik Anupam, Alexander Shypula, Osbert Bastani
cs.LGarXiv:2501.18916v22025When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling
Yong Yi Bay, Kathleen A. Yearick
cs.LGcs.AIcs.CLarXiv:2606.28661v12026Parameter-Efficient Quantum-Inspired Fast Weight Programmers for Traffic-Matrix Forecasting
Kuo-Chung Peng, Jiun-Cheng Jiang, Chun-Hua Lin +3
quant-phcs.AIcs.LGarXiv:2606.27821v12026SciForma: Structure-Faithful Generation of Scientific Diagrams
Yuxuan Luo, Peng Zhang, Xinjie Zhang +3
cs.CVcs.GRcs.LGarXiv:2607.18091v12026Hard Cases, Bad Labels: Testing Error Exposure and Error Location in Uncertainty Sampling Under Bounded Label Noise
John Myron Uy
cs.LGarXiv:2608.13601v12026No Universal Signal Predicts Sample-Level LLM Regression under Version Updates
Jia Sheng, Yiwei Lu
cs.AIcs.CLcs.LGarXiv:2608.13607v12026HI-MeshGraphNets: Efficient and Accurate Mesh-based Physics Learning with Hierarchical Multi-scale Graph Neural Networks
SiHun Lee, Dong-Hyuk Park, Taesoo Bang +1
cs.LGarXiv:2608.13827v12026Trajectory Dynamics in Self-Supervised Learning Latent Space for Audio Deepfake Detection
Tomás Andrade Weber
eess.AScs.LGcs.SDarXiv:2608.13817v12026Language-Specific Gaps in AI Safety Training Datasets
Chialuka Prisca-Mary Onuoha, Bright Etornam Sunu, Rashidat Sikiru
cs.CYcs.LGarXiv:2608.13695v12026The Integer Alibi: Localizing Cross-Kernel Divergence in INT8-Quantized LLM Inference
Teng-Ruei Chen
cs.LGarXiv:2608.13756v12026Does ISO-Grounded NFR Specification Improve LLM Code Generation? A Comparison of Rich and Structured Interventions against a Natural-Language Baseline
Joào Pedro Monteiro Pereira, Vinicius Cardoso Garcia
cs.SEcs.AIcs.LGarXiv:2608.13742v12026A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure
Dekun Yang
cs.AIcs.CLcs.LGarXiv:2608.13626v12026High-dimensional nonparametric changepoint detection via low-rank degree-two density projection
Guoqing Zhang, Zhaixin Chen
cs.LGstat.MLarXiv:2608.13922v12026How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF
Venkata Naga Sai Vishnu Rohit Pulipaka, Anish Katta, Deva Rohit Reddy Peddireddy
cs.LGarXiv:2607.19712v12026FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications
Krish Agarwal, Zhuoming Chen, Yanyuan Qin +3
cs.LGarXiv:2607.18171v22026Differentiable Logic Gate Networks for Low-Latency EEG Classification on Edge Devices
Shyamal Y. Dharia, Stephen D. Smith, Camilo E. Valderrama
cs.LGcs.AIarXiv:2607.18149v12026Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift
Zitong Huang, Gustavo Lucas Carvalho, Deqing Fu +1
cs.CLcs.LGarXiv:2607.17524v12026DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification
Yuhang Wang, Lingyao Li, Hao Zhou
cs.LGarXiv:2607.23822v12026AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition
Yuqi Li, Yi-Cheng Lin, Xianglong Wang +5
cs.LGarXiv:2607.25289v12026Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization
Alexey Gavrilov, Alan-Barsag Gazzaev, Sergey Muravyov
cs.CLcs.LGarXiv:2607.24176v12026A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever
Sietse Schelpe
cs.CLcs.AIcs.IRarXiv:2607.23806v12026Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations
Pere Martra, Eugenio Martínez Cámara, Alfonso Ureña López
cs.CLcs.CYcs.LGarXiv:2607.28319v12026H$^2$SD: Hybrid Hindsight Self-Distillation
Qiye Cai, Yichuan Ma, Peiji Li +6
cs.LGcs.CLarXiv:2607.18955v42026WorldDiT: A Unified Diffusion Architecture for World and Action Modeling
Sen Wang, R. Gnana Praveen, Bidhan Roy +1
cs.LGcs.ROarXiv:2607.23909v22026Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training
Nuemaan Malik
cs.LGcs.AIarXiv:2607.19058v22026OPERA: Offline Policy-guided Expert Routing and Adaptation for Universal Biomedical Image Analysis
Zihan Li, Feiyang Liu, Dandan Shan +2
cs.CVcs.AIcs.LGarXiv:2607.25108v12026LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks
Tianzhu Ye, Li Dong, Guanheng Chen +4
cs.LGcs.CLarXiv:2607.18110v12026Distilled Reinforcement Learning for LLM Post-training
Chen Wang, Zhaochun Li, Jionghao Bai +4
cs.LGcs.AIarXiv:2607.17247v12026CriPO: Enhancing Rubric-based RL via Self-Distillation
Mingxuan Xia, Yuhang Yang, Chao Ye +7
cs.LGcs.AIarXiv:2607.18082v32026WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting
Zhaokai Wang, Tianlin Gui, Jiayuan Rao +3
cs.AIcs.CLcs.LGarXiv:2607.18084v12026DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
Jiazhen Jiang, Boxi Cao, Lingyong Yan +6
cs.AIcs.CLcs.LGarXiv:2607.19865v12026Codifying the Judge: Scalable Evaluation via Program Distillation
Tzu-Heng Huang, Shengqi Qiu, Frederic Sala
cs.AIcs.LGarXiv:2607.22561v12026Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
Md Rezwanul Haque, Md. Milon Islam, Fakhri Karray
cs.AIcs.CLcs.LGarXiv:2607.25091v12026DataPrep-Bench: Benchmarking LLMs as Training Data Preparators
Hao Liang, Qifeng Cai, Yibo Lin +11
cs.LGcs.CLarXiv:2607.20465v12026