Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

16,621 to 16,680 of 20,199

  1. Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping

    Jesse Dodge, Gabriel Ilharco, Roy Schwartz +3

    cs.CLcs.LGarXiv:2002.06305v12020
  2. PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning

    Shaoxuan Li, Zhixuan Zhao, Hanze Deng +9

    cs.CVcs.AIcs.CLarXiv:2603.26653v12026
  3. DScribe: Library of Descriptors for Machine Learning in Materials Science

    Lauri Himanen, Marc O. J. Jäger, Eiaki V. Morooka +5

    cond-mat.mtrl-scics.LGarXiv:1904.08875v12019
  4. MemRerank: Preference Memory for Personalized Product Reranking

    Zhiyuan Peng, Xuyang Wu, Huaixiao Tou +2

    cs.CLcs.AIcs.LGarXiv:2603.29247v32026
  5. Think Anywhere in Code Generation

    Xue Jiang, Tianyu Zhang, Ge Li +8

    cs.SEcs.LGarXiv:2603.29957v32026
  6. WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

    Haipeng Luo, Qingfeng Sun, Can Xu +8

    cs.CLcs.AIcs.LGarXiv:2308.09583v32023
  7. Visual Memory Injection Attacks for Multi-Turn Conversations

    Christian Schlarmann, Matthias Hein

    cs.CVcs.LGarXiv:2602.15927v12026
  8. Global Filter Networks for Image Classification

    Yongming Rao, Wenliang Zhao, Zheng Zhu +2

    cs.CVcs.AIcs.LGarXiv:2107.00645v22021
  9. ProBel: Propaganda Detection with Techniques, Spans, and Explanations

    Mohamed Bayan Kmainasi, Ali Ezzat Shahroor, Elisa Sartori +2

    cs.CLcs.AIcs.LGarXiv:2608.22388v12026
  10. Evaluating Very Long-Term Conversational Memory of LLM Agents

    Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov +3

    cs.CLcs.AIcs.LGarXiv:2402.17753v12024
  11. GET: Generative Embedding Translation for Medical Image Segmentation

    Md Maklachur Rahman, Md Hasan Al Banna, Saraf Anjum +2

    eess.IVcs.CVcs.LGarXiv:2608.22619v12026
  12. Arbitrage-Aware Multi-Step Forecasting of Implied Volatility Surfaces: Modelling Surface Trajectories Using Latent Diffusion

    Dominik Manuel Buchegger, Lukas Gonon

    q-fin.MFcs.LGarXiv:2608.22478v12026
  13. ROCKET: Rapid Optimization via Calibration-guided Knapsack Enhanced Truncation for Efficient Model Compression

    Ammar Ali, Baher Mohammad, Denis Makhov +3

    cs.LGcs.AIcs.CLarXiv:2602.11008v12026
  14. A Tour of Reinforcement Learning: The View from Continuous Control

    Benjamin Recht

    math.OCcs.LGstat.MLarXiv:1806.09460v22018
  15. Towards Fast Computation of Certified Robustness for ReLU Networks

    Tsui-Wei Weng, Huan Zhang, Hongge Chen +5

    stat.MLcs.CRcs.CVarXiv:1804.09699v42018
  16. The Internal State of an LLM Knows When It's Lying

    Amos Azaria, Tom Mitchell

    cs.CLcs.AIcs.LGarXiv:2304.13734v22023
  17. Safe RLHF: Safe Reinforcement Learning from Human Feedback

    Josef Dai, Xuehai Pan, Ruiyang Sun +5

    cs.AIcs.LGarXiv:2310.12773v12023
  18. Sci-Reasoning: A Dataset Decoding AI Innovation Patterns

    Jiachen Liu, Maestro Harmon, Zechen Zhang

    cs.AIcs.LGarXiv:2601.04577v12026
  19. Learning to Propagate Labels: Transductive Propagation Network for Few-shot Learning

    Yanbin Liu, Juho Lee, Minseop Park +4

    cs.LGcs.CVcs.NEarXiv:1805.10002v52018
  20. VERDICT: Agreement Beats Pixel-Space Verification in Real-Document OCSR

    Yani Guan, Dengpan Dong, Shuang Luo +6

    cs.CVcs.IRcs.LGarXiv:2608.22183v12026
  21. Discriminative Embeddings of Latent Variable Models for Structured Data

    Hanjun Dai, Bo Dai, Le Song

    cs.LGarXiv:1603.05629v52016
  22. MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

    Shengding Hu, Yuge Tu, Xu Han +22

    cs.CLcs.LGarXiv:2404.06395v32024
  23. GutenOCR: A Grounded Vision-Language Front-End for Documents

    Hunter Heidenreich, Ben Elliott, Olivia Dinica +1

    cs.CVcs.AIcs.CLarXiv:2601.14490v22026
  24. Improving the Adversarial Robustness and Interpretability of Deep Neural Networks by Regularizing their Input Gradients

    Andrew Slavin Ross, Finale Doshi-Velez

    cs.LGcs.CRcs.CVarXiv:1711.09404v12017
  25. NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval

    Zhuchenyang Liu, Yao Zhang, Yu Xiao

    cs.IRcs.CVcs.LGarXiv:2603.12824v22026
  26. CodeT5+: Open Code Large Language Models for Code Understanding and Generation

    Yue Wang, Hung Le, Akhilesh Deepak Gotmare +3

    cs.CLcs.LGcs.PLarXiv:2305.07922v22023
  27. Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAML

    Aniruddh Raghu, Maithra Raghu, Samy Bengio +1

    cs.LGstat.MLarXiv:1909.09157v22019
  28. Less is more: sampling chemical space with active learning

    Justin S. Smith, Ben Nebgen, Nicholas Lubbers +2

    physics.comp-phcs.LGphysics.chem-pharXiv:1801.09319v22018
  29. JudgeRLVR: Judge First, Generate Second for Efficient Reasoning

    Jiangshan Duo, Hanyu Li, Hailin Zhang +3

    cs.CLcs.AIcs.LGarXiv:2601.08468v12026
  30. Joint Causal Structure and Cluster Discovery Using Variational Inference

    Avni Rajpal, Anubhav Kumar, Rishabh Karnad +2

    cs.LGcs.AIstat.MLarXiv:2608.22212v12026
  31. LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action

    Dhruv Shah, Blazej Osinski, Brian Ichter +1

    cs.ROcs.AIcs.CLarXiv:2207.04429v22022
  32. Lost in the Prompt Order: Revealing the Limitations of Causal Attention in Language Models

    Hyunjong Ok, Jaeho Lee

    cs.CLcs.AIcs.LGarXiv:2601.14152v22026
  33. On-policy Distillation with Verifiable Reward

    Wenze Lin, Jiale Zhao, Xitai Jiang +5

    cs.LGcs.AIarXiv:2608.24696v12026
  34. A Comprehensive Overhaul of Feature Distillation

    Byeongho Heo, Jeesoo Kim, Sangdoo Yun +3

    cs.CVcs.LGarXiv:1904.01866v22019
  35. Safe Flow Q-Learning: Offline Safe Reinforcement Learning with Reachability-Based Flow Policies

    Mumuksh Tayal, Manan Tayal, Ravi Prakash

    cs.LGcs.AIarXiv:2603.15136v22026
  36. A Unified Game-Theoretic Approach to Multiagent Reinforcement Learning

    Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys +5

    cs.AIcs.GTcs.LGarXiv:1711.00832v22017
  37. Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning

    Minwu Kim, Safal Shrestha, Anubhav Shrestha +1

    cs.LGcs.AIcs.CLarXiv:2601.20829v22026
  38. Deep Learning COVID-19 Features on CXR using Limited Training Data Sets

    Yujin Oh, Sangjoon Park, Jong Chul Ye

    eess.IVcs.CVcs.LGarXiv:2004.05758v22020
  39. ECO: Quantized Training without Full-Precision Master Weights

    Mahdi Nikdan, Amir Zandieh, Dan Alistarh +1

    cs.CLcs.AIcs.LGarXiv:2601.22101v12026
  40. Lightweight Multi-scale Hierarchical Anomaly Detection and Localization for Geospatial Big Data Applications at the Edge

    Thomas Benton Townsend, Joshua Bean, Benjamin K Tkach +2

    eess.SPcs.LGeess.SYarXiv:2608.22648v12026
  41. DRPG (Decompose, Retrieve, Plan, Generate): An Agentic Framework for Academic Rebuttal

    Peixuan Han, Yingjie Yu, Jingjun Xu +1

    cs.LGarXiv:2601.18081v22026
  42. Visual Transformers: Token-based Image Representation and Processing for Computer Vision

    Bichen Wu, Chenfeng Xu, Xiaoliang Dai +7

    cs.CVcs.LGeess.IVarXiv:2006.03677v42020
  43. Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech

    Vadim Popov, Ivan Vovk, Vladimir Gogoryan +2

    cs.LGcs.CLstat.MLarXiv:2105.06337v22021
  44. RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

    Hanze Dong, Wei Xiong, Deepanshu Goyal +7

    cs.LGcs.AIcs.CLarXiv:2304.06767v42023
  45. Automatic Prompt Optimization with "Gradient Descent" and Beam Search

    Reid Pryzant, Dan Iter, Jerry Li +3

    cs.CLcs.AIcs.LGarXiv:2305.03495v22023
  46. Where Cognition Lives: Dissecting Emergent from Computed Function in a Minimal Complete Cognitive Architecture

    Francisco M. Arrabal-Campos, Francisco G. Montoya, Alfredo Alcayde +1

    cs.AIcs.CLcs.LGarXiv:2608.22347v12026
  47. Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning

    Felipe Petroski Such, Vashisht Madhavan, Edoardo Conti +3

    cs.NEcs.LGarXiv:1712.06567v32017
  48. Multi-Modal Fusion Transformer for End-to-End Autonomous Driving

    Aditya Prakash, Kashyap Chitta, Andreas Geiger

    cs.CVcs.AIcs.LGarXiv:2104.09224v12021
  49. Variational Information Distillation for Knowledge Transfer

    Sungsoo Ahn, Shell Xu Hu, Andreas Damianou +2

    cs.CVcs.AIcs.LGarXiv:1904.05835v12019
  50. Adversarial Attacks and Defenses in Images, Graphs and Text: A Review

    Han Xu, Yao Ma, Haochen Liu +4

    cs.LGcs.CRstat.MLarXiv:1909.08072v22019
  51. Big Self-Supervised Models Advance Medical Image Classification

    Shekoofeh Azizi, Basil Mustafa, Fiona Ryan +9

    eess.IVcs.CVcs.LGarXiv:2101.05224v22021
  52. RAMP: Reinforcement Adaptive Mixed Precision Quantization for Efficient On Device LLM Inference

    Arpit Singh Gautam, Saurabh Jha

    cs.LGcs.AIarXiv:2603.17891v12026
  53. Fairness in Machine Learning: A Survey

    Simon Caton, Christian Haas

    cs.LGstat.MLarXiv:2010.04053v12020
  54. Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules

    Florian Rottach, Sebastian Schieferdecker, William Rudman +2

    cs.LGcs.AIarXiv:2608.22642v12026
  55. Tracing the Unlabeled Storm: Cross-Variable Transfer in a Lagrangian Atmospheric JEPA Framework

    K M Anirudh, S Sandeep, Hariprasad Kodamana

    cs.LGphysics.geo-pharXiv:2608.22358v12026
  56. The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models

    Taebong Kim, Youngsik Hong, Minsik Kim +3

    cs.LGcs.AIarXiv:2608.22876v12026
  57. TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages

    Jonathan H. Clark, Eunsol Choi, Michael Collins +4

    cs.CLcs.LGarXiv:2003.05002v12020
    Summaries:한국어
  58. Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch

    Le Yu, Bowen Yu, Haiyang Yu +2

    cs.CLcs.LGarXiv:2311.03099v32023
  59. SimpleGPT: Improving GPT via A Simple Normalization Strategy

    Marco Chen, Xianbiao Qi, Yelin He +2

    cs.LGcs.CLcs.CVarXiv:2602.01212v12026
  60. On the (Statistical) Detection of Adversarial Examples

    Kathrin Grosse, Praveen Manoharan, Nicolas Papernot +2

    cs.CRcs.LGstat.MLarXiv:1702.06280v22017