Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
19,501 to 19,560 of 20,193
ExpRL: Exploratory RL for LLM Mid-Training
Violet Xiang, Amrith Setlur, Chase Blagden +2
cs.LGarXiv:2606.17024v12026Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes
Tongyan Fang, Siyuan Huang, Naiyu Fang +6
cs.ROcs.LGarXiv:2606.17043v12026ProCUA-SFT Technical Report
Jaehun Jung, Ximing Lu, Brandon Cui +11
cs.LGcs.CVarXiv:2606.17321v12026Rethinking Reverse KL as Adaptive Entropy Distillation
Shizhen Li, Zhiyu Shen, Yuyin Lu +4
cs.LGstat.MLarXiv:2608.14685v12026Does the Heart Show Your Pain? Tackling the X-ITE Pain Challenge with Self-Supervised ECG Representation Learning
Dominika Kunc, Przemysław Kazienko, Stanisław Saganowski
eess.SPcs.AIcs.LGarXiv:2608.14662v12026Randomly initialized autoencoders: fixed points and edge-of-chaos
Leonid Berlyand, Roman Sarapin, Yitzchak Shmalo +2
cs.LGmath.PRmath.STarXiv:2608.14638v12026Ring-based Spatial Transformer: Learning Non-linear Spatial Interactions between Building Distribution and Pedestrian Flow
Shun Nakayama, Takahiro Kanamori, Wanglin Yan
cs.LGcs.AIcs.CYarXiv:2608.14660v12026The Quantum Shortcut: Complex Phase-State Dynamics Reduce the Optimization Steps of Sequence Models
Ahmed Nebli, Hadi Saadatdoorabi, Christopher Keibel +1
cs.LGquant-pharXiv:2608.14691v12026A Low-Cost IoT Device for Environmental Monitoring and Embedded Solar Forecasting with On-Device Incremental Learning
Erick Michel Lara Pinal, Abhinav Das, Stephan Schlüter
eess.SPcs.LGarXiv:2608.14698v12026One Score, Two Decisions: Selective Prediction on the Rare-Disease Tail
Zhaoyang Jiang, Zhizhong Fu, Yunsoo Kim +5
cs.LGarXiv:2608.14683v12026Take it Personally: The Limits of General SSL Representations for Real-Life PPG Emotion Detection
Dominika Kunc, Przemysław Kazienko, Stanisław Saganowski
cs.LGcs.AIarXiv:2608.14675v12026Mitigating Rubric Interference in LLM Judges via On-Policy Self-Distillation
Dingyao Yu, Tong Zhang, Yutao Mou +3
cs.LGcs.AIarXiv:2608.14684v12026p-Spin Glass Network Efficient Single-Batch Continual Learning
Vladimer Khasia
cs.LGarXiv:2608.14774v12026Hardware-in-the-Loop Phase-Aware CNN for Real-Time 5G Channel Estimation
Javad Zolfaghari-Bengar, Rakibul Rony, Elisa Gomez-de-Lope +3
eess.SPcs.CVcs.ITarXiv:2608.14709v12026Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling
Yang Zhao, Peisong Niu, Tian Zhou +5
cs.LGcs.AIcs.CVarXiv:2608.14652v12026When Uncertainty Isn't Enough: An Empirical Study of Self-Correction in Code Generation
Pranav Rakasi, Maanas Lalwani, Arnav Srivastava +4
cs.AIcs.LGcs.SEarXiv:2608.14659v12026Training and Evaluating Ethical Reinforcement Learning Agents on Per-Episode Distributions
Prabhjyot Singh, Majid Ghasemi, Mark Crowley
cs.LGarXiv:2608.14642v12026pico-type: A 1.5M-Parameter Byte-Level Multi-Head Content Classifier
Gautam Kishore
cs.LGcs.AIcs.CLarXiv:2608.14658v12026Phase-Aware CNN for Real-Time 5G/6G Channel Estimation with Hardware-in-the-loop Validation
Javad Zolfaghari-Bengar, Rakibul Rony, Elisa Gomez-de-Lope +3
eess.SPcs.LGarXiv:2608.14676v12026Belayer: Efficient Fault Tolerance for LLM Agentic RL Training
Jiecheng Zhou, Qinghao Hu, Peng Sun +2
cs.DCcs.LGarXiv:2608.14635v22026Offline Ambient-Controlled Latent Diffusion: Architecture, Telemetry, and On-Device Evaluation
Lech Kalinowski, Artur Morys-Magiera, Piotr Miłkowski
eess.SPcs.AIcs.LGarXiv:2608.14677v12026Efficient Neural-Network-Based High-Resolution Radiative Transfer for CO___ Retrieval, and Application to Interferometric Sensing
Jordan Lontsi Tedongmo, Yann Ferrec, Laurence Croizé +3
cs.LGphysics.ao-pharXiv:2608.14645v12026In-Context Learning to Assess Built Environment Impacts on Perceived Neighborhood Walkability Among Mobility-impaired Older Adults
Houhao Liang, Kresimir Friganovic, Joanne Kua +5
cs.LGstat.AParXiv:2608.14663v12026DUET: Dual-Teacher On-Policy Distillation via Same-Weight Disagreement for Prohibition Compliance
Zihan Li, Feifei Li, Wenhui Que
cs.LGcs.CLarXiv:2608.14644v12026ARGUS: Attention-Guided Transformers for Scalable Person Identification Using Wi-Fi Telemetry
Nayan Sanjay Bhatia, Pranay Kocheta, Yuhan Li +1
cs.LGcs.AIcs.CVarXiv:2608.14670v12026Metaplasticity as adaptive gradient preconditioning for incremental learning
Isabelle Aguilar, Zayn Andre Zainal, Omid Kavehei
cs.LGarXiv:2608.14634v12026Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays
Bhaskar Gurram
cs.LGcs.AIcs.CLarXiv:2608.14639v12026Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation
Zhizhao Liu, Zhiliang Tian, Xi Wang +4
cs.LGcs.AIcs.CLarXiv:2608.17941v12026NeuRoute: Logit-Guided Neural Routing for Billion-Scale Vector Search with Sub-Hour Index Construction
Xingqiao Wang, Zi Wang, Xiaowei Xu
cs.DBcs.IRcs.LGarXiv:2608.15438v12026Unraveling the Size Determination Mechanism of Nanocrystal Synthesis via Interpretable Neural Networks
Kai Gu, Haizheng Zhong
cs.LGcond-mat.mtrl-scics.AIarXiv:2608.14734v12026Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning
Joanikij Chulev, Hendrik Baier
cs.LGcs.AIcs.HCarXiv:2608.14963v12026Convolution Smoothed Quantile Regression for XGBoost
Mandy Yao, Meredith Franklin
stat.MLcs.LGarXiv:2608.15290v12026Detecting Money Laundering in Rwandan Mobile Money: A Machine Learning Framework
Emmanuel Nahimana, Yaé Ulrich Gaba
cs.LGq-fin.RMarXiv:2608.15447v12026MAPLE: MoE Adaptive Plug-and-play Layer-wise Expert allocation
Lie Li, Wen Li, Junxiao Shen +1
cs.LGcs.AIarXiv:2608.15299v12026Iterative Refinement Diffusion for Super-Resolved Data Assimilation of Multiscale Physical Systems
Mrigank Dhingra, Ramchandran Muthukumar, Rebecca Willett +1
cs.LGphysics.flu-dynarXiv:2608.14744v12026A Parameter-Free Few-Shot Evaluation for Elephant Vocalisation Classification
Christiaan M. Geldenhuys, Thomas R. Niesler
eess.AScs.LGcs.SDarXiv:2608.14824v12026On Cross-Validation for Hyperparameter Optimization of Deep Learning Image Classifiers
Ljubomir Buturovic
cs.CVcs.LGarXiv:2608.14705v12026Beyond Boundary Noise: Aggregated Aleatoric Uncertainty Fails to Capture Presence Ambiguity in 3D Lung Nodule Segmentation
Simon Baur, Arne Schernich, Ekin Böke +2
cs.CVcs.LGarXiv:2608.14766v12026PWLR: Pairwise Witness Local Rejection for Boundary-Aware Out-of-Distribution Detection
Chengyao Jia, Ruixuan Wang
cs.CVcs.LGarXiv:2608.15802v12026Geometry of Forgetting: Representation Flux in Continual Learning
Maksim A. Kazanskii
cs.LGcs.CVarXiv:2608.15854v12026Cross-Entropy Risk Estimation for Language Models: Inconsistency Must Be Dense, and the Holdout Method Is No Exception
Hanti Lin
cs.LGstat.MLarXiv:2608.15798v12026Not All Attention Is Equal: A Quantitative Survey of the EEI Trade-off
Aditya Singh
cs.LGcs.AIarXiv:2608.15459v12026Solvable Sokoban Without a Solver via Diffusion
Sina Baghal
cs.AIcs.GTcs.LGarXiv:2608.15958v12026PandasCorpus: A Resource of Real-World Pandas Workflows and Usage Patterns
Syrym Abdikhan, Mazhar Hameed
cs.SEcs.LGarXiv:2608.14742v12026EMASAM: a Computationally Efficient Sharpness-Aware Minimization via EMA-Guided Perturbations
Tanapat Ratchatorn, Masayuki Tanaka
cs.LGcs.CVarXiv:2608.15105v12026SAGA: Structure-Attended Generative Action Embedding Model that encodes Multi-Surface User Action Sequences
Tsz Fung Pang, Po Jen Chen, Nimish Ronghe +2
cs.LGcs.IRarXiv:2608.15429v12026M-LINKX: Multiview Graph Learning for Brain Cognitive Disease Detection
An Phan, Yufei Jin, Xingquan Zhu
cs.LGarXiv:2608.14847v12026Distinguishing AI-Generated Music from Edited Audio as a Hard-Negative Robustness Task
Alexandru-Stefan Morosanu, Valerian Cecan, Stefan-Daniel Achirei +1
cs.SDcs.AIcs.LGarXiv:2608.14916v12026Adaptive Volumetric Mechanical Property Fields Invariant to Resolution
Rishit Dagli, Donglai Xiang, Vismay Modi +4
cs.CVcs.LGcs.ROarXiv:2606.18231v12026LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI
Lalit Yadav, Akshaj Gurugubelli
cs.AIcs.CLcs.LGarXiv:2606.18021v12026Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models
Nikita Kachaev, Andrey Moskalenko, Matvey Skripkin +10
cs.LGcs.ROarXiv:2606.19297v12026Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States
Denis Peskoff, Joe Barrow, Christopher Vu +1
cs.CLcs.CYcs.LGarXiv:2606.19334v12026How Post-Training Shapes Biological Reasoning Models
Lukas Fesser, Hanlin Zhang, Michelle M. Li +5
cs.LGq-bio.QMarXiv:2606.16517v22026Characterization of Thermal Systems from Noisy and Low-resolution Measurements Using Dynamic Mode Decomposition
M. E. P. Silva, L. S. Araujo, F. T. Colombo +2
physics.comp-phcs.LGeess.SParXiv:2608.14581v12026Evaluating the impact of adversarial traffic patterns on vanet communication using veins simulation
Henry Agyapong
cs.NIcs.LGarXiv:2608.14583v12026Sumi: Open Uniform Diffusion Language Model from Scratch
Mengyu Ye, Keito Kudo, Wataru Ikeda +3
cs.CLcs.LGarXiv:2606.19005v12026When, Where, and How: Adaptive Binning for Tabular Self-Supervised Learning
Daehwan Kim, Haejun Chung, Ikbeom Jang
cs.LGcs.AIarXiv:2606.19827v12026Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention
Vishesh Tripathi, Abhay Kumar
cs.LGarXiv:2606.20945v22026Causal Discovery in the Era of Agents
Yujia Zheng, Vishal Verma, Mantej Gill +3
cs.AIcs.LGcs.SEarXiv:2606.23608v12026VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct
Haoling Li, Kai Zheng, Jie Wu +4
cs.AIcs.CLcs.CVarXiv:2606.23543v12026