Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

19,501 to 19,560 of 20,193

  1. ExpRL: Exploratory RL for LLM Mid-Training

    Violet Xiang, Amrith Setlur, Chase Blagden +2

    cs.LGarXiv:2606.17024v12026
  2. Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes

    Tongyan Fang, Siyuan Huang, Naiyu Fang +6

    cs.ROcs.LGarXiv:2606.17043v12026
  3. ProCUA-SFT Technical Report

    Jaehun Jung, Ximing Lu, Brandon Cui +11

    cs.LGcs.CVarXiv:2606.17321v12026
  4. Rethinking Reverse KL as Adaptive Entropy Distillation

    Shizhen Li, Zhiyu Shen, Yuyin Lu +4

    cs.LGstat.MLarXiv:2608.14685v12026
  5. Does the Heart Show Your Pain? Tackling the X-ITE Pain Challenge with Self-Supervised ECG Representation Learning

    Dominika Kunc, Przemysław Kazienko, Stanisław Saganowski

    eess.SPcs.AIcs.LGarXiv:2608.14662v12026
  6. Randomly initialized autoencoders: fixed points and edge-of-chaos

    Leonid Berlyand, Roman Sarapin, Yitzchak Shmalo +2

    cs.LGmath.PRmath.STarXiv:2608.14638v12026
  7. Ring-based Spatial Transformer: Learning Non-linear Spatial Interactions between Building Distribution and Pedestrian Flow

    Shun Nakayama, Takahiro Kanamori, Wanglin Yan

    cs.LGcs.AIcs.CYarXiv:2608.14660v12026
  8. The Quantum Shortcut: Complex Phase-State Dynamics Reduce the Optimization Steps of Sequence Models

    Ahmed Nebli, Hadi Saadatdoorabi, Christopher Keibel +1

    cs.LGquant-pharXiv:2608.14691v12026
  9. A Low-Cost IoT Device for Environmental Monitoring and Embedded Solar Forecasting with On-Device Incremental Learning

    Erick Michel Lara Pinal, Abhinav Das, Stephan Schlüter

    eess.SPcs.LGarXiv:2608.14698v12026
  10. One Score, Two Decisions: Selective Prediction on the Rare-Disease Tail

    Zhaoyang Jiang, Zhizhong Fu, Yunsoo Kim +5

    cs.LGarXiv:2608.14683v12026
  11. Take it Personally: The Limits of General SSL Representations for Real-Life PPG Emotion Detection

    Dominika Kunc, Przemysław Kazienko, Stanisław Saganowski

    cs.LGcs.AIarXiv:2608.14675v12026
  12. Mitigating Rubric Interference in LLM Judges via On-Policy Self-Distillation

    Dingyao Yu, Tong Zhang, Yutao Mou +3

    cs.LGcs.AIarXiv:2608.14684v12026
  13. p-Spin Glass Network Efficient Single-Batch Continual Learning

    Vladimer Khasia

    cs.LGarXiv:2608.14774v12026
  14. Hardware-in-the-Loop Phase-Aware CNN for Real-Time 5G Channel Estimation

    Javad Zolfaghari-Bengar, Rakibul Rony, Elisa Gomez-de-Lope +3

    eess.SPcs.CVcs.ITarXiv:2608.14709v12026
  15. Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling

    Yang Zhao, Peisong Niu, Tian Zhou +5

    cs.LGcs.AIcs.CVarXiv:2608.14652v12026
  16. When Uncertainty Isn't Enough: An Empirical Study of Self-Correction in Code Generation

    Pranav Rakasi, Maanas Lalwani, Arnav Srivastava +4

    cs.AIcs.LGcs.SEarXiv:2608.14659v12026
  17. Training and Evaluating Ethical Reinforcement Learning Agents on Per-Episode Distributions

    Prabhjyot Singh, Majid Ghasemi, Mark Crowley

    cs.LGarXiv:2608.14642v12026
  18. pico-type: A 1.5M-Parameter Byte-Level Multi-Head Content Classifier

    Gautam Kishore

    cs.LGcs.AIcs.CLarXiv:2608.14658v12026
  19. Phase-Aware CNN for Real-Time 5G/6G Channel Estimation with Hardware-in-the-loop Validation

    Javad Zolfaghari-Bengar, Rakibul Rony, Elisa Gomez-de-Lope +3

    eess.SPcs.LGarXiv:2608.14676v12026
  20. Belayer: Efficient Fault Tolerance for LLM Agentic RL Training

    Jiecheng Zhou, Qinghao Hu, Peng Sun +2

    cs.DCcs.LGarXiv:2608.14635v22026
  21. Offline Ambient-Controlled Latent Diffusion: Architecture, Telemetry, and On-Device Evaluation

    Lech Kalinowski, Artur Morys-Magiera, Piotr Miłkowski

    eess.SPcs.AIcs.LGarXiv:2608.14677v12026
  22. Efficient Neural-Network-Based High-Resolution Radiative Transfer for CO___ Retrieval, and Application to Interferometric Sensing

    Jordan Lontsi Tedongmo, Yann Ferrec, Laurence Croizé +3

    cs.LGphysics.ao-pharXiv:2608.14645v12026
  23. In-Context Learning to Assess Built Environment Impacts on Perceived Neighborhood Walkability Among Mobility-impaired Older Adults

    Houhao Liang, Kresimir Friganovic, Joanne Kua +5

    cs.LGstat.AParXiv:2608.14663v12026
  24. DUET: Dual-Teacher On-Policy Distillation via Same-Weight Disagreement for Prohibition Compliance

    Zihan Li, Feifei Li, Wenhui Que

    cs.LGcs.CLarXiv:2608.14644v12026
  25. ARGUS: Attention-Guided Transformers for Scalable Person Identification Using Wi-Fi Telemetry

    Nayan Sanjay Bhatia, Pranay Kocheta, Yuhan Li +1

    cs.LGcs.AIcs.CVarXiv:2608.14670v12026
  26. Metaplasticity as adaptive gradient preconditioning for incremental learning

    Isabelle Aguilar, Zayn Andre Zainal, Omid Kavehei

    cs.LGarXiv:2608.14634v12026
  27. Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays

    Bhaskar Gurram

    cs.LGcs.AIcs.CLarXiv:2608.14639v12026
  28. Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation

    Zhizhao Liu, Zhiliang Tian, Xi Wang +4

    cs.LGcs.AIcs.CLarXiv:2608.17941v12026
  29. NeuRoute: Logit-Guided Neural Routing for Billion-Scale Vector Search with Sub-Hour Index Construction

    Xingqiao Wang, Zi Wang, Xiaowei Xu

    cs.DBcs.IRcs.LGarXiv:2608.15438v12026
  30. Unraveling the Size Determination Mechanism of Nanocrystal Synthesis via Interpretable Neural Networks

    Kai Gu, Haizheng Zhong

    cs.LGcond-mat.mtrl-scics.AIarXiv:2608.14734v12026
  31. Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning

    Joanikij Chulev, Hendrik Baier

    cs.LGcs.AIcs.HCarXiv:2608.14963v12026
  32. Convolution Smoothed Quantile Regression for XGBoost

    Mandy Yao, Meredith Franklin

    stat.MLcs.LGarXiv:2608.15290v12026
  33. Detecting Money Laundering in Rwandan Mobile Money: A Machine Learning Framework

    Emmanuel Nahimana, Yaé Ulrich Gaba

    cs.LGq-fin.RMarXiv:2608.15447v12026
  34. MAPLE: MoE Adaptive Plug-and-play Layer-wise Expert allocation

    Lie Li, Wen Li, Junxiao Shen +1

    cs.LGcs.AIarXiv:2608.15299v12026
  35. Iterative Refinement Diffusion for Super-Resolved Data Assimilation of Multiscale Physical Systems

    Mrigank Dhingra, Ramchandran Muthukumar, Rebecca Willett +1

    cs.LGphysics.flu-dynarXiv:2608.14744v12026
  36. A Parameter-Free Few-Shot Evaluation for Elephant Vocalisation Classification

    Christiaan M. Geldenhuys, Thomas R. Niesler

    eess.AScs.LGcs.SDarXiv:2608.14824v12026
  37. On Cross-Validation for Hyperparameter Optimization of Deep Learning Image Classifiers

    Ljubomir Buturovic

    cs.CVcs.LGarXiv:2608.14705v12026
  38. Beyond Boundary Noise: Aggregated Aleatoric Uncertainty Fails to Capture Presence Ambiguity in 3D Lung Nodule Segmentation

    Simon Baur, Arne Schernich, Ekin Böke +2

    cs.CVcs.LGarXiv:2608.14766v12026
  39. PWLR: Pairwise Witness Local Rejection for Boundary-Aware Out-of-Distribution Detection

    Chengyao Jia, Ruixuan Wang

    cs.CVcs.LGarXiv:2608.15802v12026
  40. Geometry of Forgetting: Representation Flux in Continual Learning

    Maksim A. Kazanskii

    cs.LGcs.CVarXiv:2608.15854v12026
  41. Cross-Entropy Risk Estimation for Language Models: Inconsistency Must Be Dense, and the Holdout Method Is No Exception

    Hanti Lin

    cs.LGstat.MLarXiv:2608.15798v12026
  42. Not All Attention Is Equal: A Quantitative Survey of the EEI Trade-off

    Aditya Singh

    cs.LGcs.AIarXiv:2608.15459v12026
  43. Solvable Sokoban Without a Solver via Diffusion

    Sina Baghal

    cs.AIcs.GTcs.LGarXiv:2608.15958v12026
  44. PandasCorpus: A Resource of Real-World Pandas Workflows and Usage Patterns

    Syrym Abdikhan, Mazhar Hameed

    cs.SEcs.LGarXiv:2608.14742v12026
  45. EMASAM: a Computationally Efficient Sharpness-Aware Minimization via EMA-Guided Perturbations

    Tanapat Ratchatorn, Masayuki Tanaka

    cs.LGcs.CVarXiv:2608.15105v12026
  46. SAGA: Structure-Attended Generative Action Embedding Model that encodes Multi-Surface User Action Sequences

    Tsz Fung Pang, Po Jen Chen, Nimish Ronghe +2

    cs.LGcs.IRarXiv:2608.15429v12026
  47. M-LINKX: Multiview Graph Learning for Brain Cognitive Disease Detection

    An Phan, Yufei Jin, Xingquan Zhu

    cs.LGarXiv:2608.14847v12026
  48. Distinguishing AI-Generated Music from Edited Audio as a Hard-Negative Robustness Task

    Alexandru-Stefan Morosanu, Valerian Cecan, Stefan-Daniel Achirei +1

    cs.SDcs.AIcs.LGarXiv:2608.14916v12026
  49. Adaptive Volumetric Mechanical Property Fields Invariant to Resolution

    Rishit Dagli, Donglai Xiang, Vismay Modi +4

    cs.CVcs.LGcs.ROarXiv:2606.18231v12026
  50. LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI

    Lalit Yadav, Akshaj Gurugubelli

    cs.AIcs.CLcs.LGarXiv:2606.18021v12026
  51. Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models

    Nikita Kachaev, Andrey Moskalenko, Matvey Skripkin +10

    cs.LGcs.ROarXiv:2606.19297v12026
  52. Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States

    Denis Peskoff, Joe Barrow, Christopher Vu +1

    cs.CLcs.CYcs.LGarXiv:2606.19334v12026
  53. How Post-Training Shapes Biological Reasoning Models

    Lukas Fesser, Hanlin Zhang, Michelle M. Li +5

    cs.LGq-bio.QMarXiv:2606.16517v22026
  54. Characterization of Thermal Systems from Noisy and Low-resolution Measurements Using Dynamic Mode Decomposition

    M. E. P. Silva, L. S. Araujo, F. T. Colombo +2

    physics.comp-phcs.LGeess.SParXiv:2608.14581v12026
  55. Evaluating the impact of adversarial traffic patterns on vanet communication using veins simulation

    Henry Agyapong

    cs.NIcs.LGarXiv:2608.14583v12026
  56. Sumi: Open Uniform Diffusion Language Model from Scratch

    Mengyu Ye, Keito Kudo, Wataru Ikeda +3

    cs.CLcs.LGarXiv:2606.19005v12026
  57. When, Where, and How: Adaptive Binning for Tabular Self-Supervised Learning

    Daehwan Kim, Haejun Chung, Ikbeom Jang

    cs.LGcs.AIarXiv:2606.19827v12026
  58. Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention

    Vishesh Tripathi, Abhay Kumar

    cs.LGarXiv:2606.20945v22026
  59. Causal Discovery in the Era of Agents

    Yujia Zheng, Vishal Verma, Mantej Gill +3

    cs.AIcs.LGcs.SEarXiv:2606.23608v12026
  60. VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct

    Haoling Li, Kai Zheng, Jie Wu +4

    cs.AIcs.CLcs.CVarXiv:2606.23543v12026