Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
16,621 to 16,680 of 20,199
Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz +3
cs.CLcs.LGarXiv:2002.06305v12020PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning
Shaoxuan Li, Zhixuan Zhao, Hanze Deng +9
cs.CVcs.AIcs.CLarXiv:2603.26653v12026DScribe: Library of Descriptors for Machine Learning in Materials Science
Lauri Himanen, Marc O. J. Jäger, Eiaki V. Morooka +5
cond-mat.mtrl-scics.LGarXiv:1904.08875v12019MemRerank: Preference Memory for Personalized Product Reranking
Zhiyuan Peng, Xuyang Wu, Huaixiao Tou +2
cs.CLcs.AIcs.LGarXiv:2603.29247v32026Think Anywhere in Code Generation
Xue Jiang, Tianyu Zhang, Ge Li +8
cs.SEcs.LGarXiv:2603.29957v32026WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
Haipeng Luo, Qingfeng Sun, Can Xu +8
cs.CLcs.AIcs.LGarXiv:2308.09583v32023Visual Memory Injection Attacks for Multi-Turn Conversations
Christian Schlarmann, Matthias Hein
cs.CVcs.LGarXiv:2602.15927v12026Global Filter Networks for Image Classification
Yongming Rao, Wenliang Zhao, Zheng Zhu +2
cs.CVcs.AIcs.LGarXiv:2107.00645v22021ProBel: Propaganda Detection with Techniques, Spans, and Explanations
Mohamed Bayan Kmainasi, Ali Ezzat Shahroor, Elisa Sartori +2
cs.CLcs.AIcs.LGarXiv:2608.22388v12026Evaluating Very Long-Term Conversational Memory of LLM Agents
Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov +3
cs.CLcs.AIcs.LGarXiv:2402.17753v12024GET: Generative Embedding Translation for Medical Image Segmentation
Md Maklachur Rahman, Md Hasan Al Banna, Saraf Anjum +2
eess.IVcs.CVcs.LGarXiv:2608.22619v12026Arbitrage-Aware Multi-Step Forecasting of Implied Volatility Surfaces: Modelling Surface Trajectories Using Latent Diffusion
Dominik Manuel Buchegger, Lukas Gonon
q-fin.MFcs.LGarXiv:2608.22478v12026ROCKET: Rapid Optimization via Calibration-guided Knapsack Enhanced Truncation for Efficient Model Compression
Ammar Ali, Baher Mohammad, Denis Makhov +3
cs.LGcs.AIcs.CLarXiv:2602.11008v12026A Tour of Reinforcement Learning: The View from Continuous Control
Benjamin Recht
math.OCcs.LGstat.MLarXiv:1806.09460v22018Towards Fast Computation of Certified Robustness for ReLU Networks
Tsui-Wei Weng, Huan Zhang, Hongge Chen +5
stat.MLcs.CRcs.CVarXiv:1804.09699v42018The Internal State of an LLM Knows When It's Lying
Amos Azaria, Tom Mitchell
cs.CLcs.AIcs.LGarXiv:2304.13734v22023Safe RLHF: Safe Reinforcement Learning from Human Feedback
Josef Dai, Xuehai Pan, Ruiyang Sun +5
cs.AIcs.LGarXiv:2310.12773v12023Sci-Reasoning: A Dataset Decoding AI Innovation Patterns
Jiachen Liu, Maestro Harmon, Zechen Zhang
cs.AIcs.LGarXiv:2601.04577v12026Learning to Propagate Labels: Transductive Propagation Network for Few-shot Learning
Yanbin Liu, Juho Lee, Minseop Park +4
cs.LGcs.CVcs.NEarXiv:1805.10002v52018VERDICT: Agreement Beats Pixel-Space Verification in Real-Document OCSR
Yani Guan, Dengpan Dong, Shuang Luo +6
cs.CVcs.IRcs.LGarXiv:2608.22183v12026Discriminative Embeddings of Latent Variable Models for Structured Data
Hanjun Dai, Bo Dai, Le Song
cs.LGarXiv:1603.05629v52016MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Shengding Hu, Yuge Tu, Xu Han +22
cs.CLcs.LGarXiv:2404.06395v32024GutenOCR: A Grounded Vision-Language Front-End for Documents
Hunter Heidenreich, Ben Elliott, Olivia Dinica +1
cs.CVcs.AIcs.CLarXiv:2601.14490v22026Improving the Adversarial Robustness and Interpretability of Deep Neural Networks by Regularizing their Input Gradients
Andrew Slavin Ross, Finale Doshi-Velez
cs.LGcs.CRcs.CVarXiv:1711.09404v12017NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval
Zhuchenyang Liu, Yao Zhang, Yu Xiao
cs.IRcs.CVcs.LGarXiv:2603.12824v22026CodeT5+: Open Code Large Language Models for Code Understanding and Generation
Yue Wang, Hung Le, Akhilesh Deepak Gotmare +3
cs.CLcs.LGcs.PLarXiv:2305.07922v22023Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAML
Aniruddh Raghu, Maithra Raghu, Samy Bengio +1
cs.LGstat.MLarXiv:1909.09157v22019Less is more: sampling chemical space with active learning
Justin S. Smith, Ben Nebgen, Nicholas Lubbers +2
physics.comp-phcs.LGphysics.chem-pharXiv:1801.09319v22018JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
Jiangshan Duo, Hanyu Li, Hailin Zhang +3
cs.CLcs.AIcs.LGarXiv:2601.08468v12026Joint Causal Structure and Cluster Discovery Using Variational Inference
Avni Rajpal, Anubhav Kumar, Rishabh Karnad +2
cs.LGcs.AIstat.MLarXiv:2608.22212v12026LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action
Dhruv Shah, Blazej Osinski, Brian Ichter +1
cs.ROcs.AIcs.CLarXiv:2207.04429v22022Lost in the Prompt Order: Revealing the Limitations of Causal Attention in Language Models
Hyunjong Ok, Jaeho Lee
cs.CLcs.AIcs.LGarXiv:2601.14152v22026On-policy Distillation with Verifiable Reward
Wenze Lin, Jiale Zhao, Xitai Jiang +5
cs.LGcs.AIarXiv:2608.24696v12026A Comprehensive Overhaul of Feature Distillation
Byeongho Heo, Jeesoo Kim, Sangdoo Yun +3
cs.CVcs.LGarXiv:1904.01866v22019Safe Flow Q-Learning: Offline Safe Reinforcement Learning with Reachability-Based Flow Policies
Mumuksh Tayal, Manan Tayal, Ravi Prakash
cs.LGcs.AIarXiv:2603.15136v22026A Unified Game-Theoretic Approach to Multiagent Reinforcement Learning
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys +5
cs.AIcs.GTcs.LGarXiv:1711.00832v22017Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning
Minwu Kim, Safal Shrestha, Anubhav Shrestha +1
cs.LGcs.AIcs.CLarXiv:2601.20829v22026Deep Learning COVID-19 Features on CXR using Limited Training Data Sets
Yujin Oh, Sangjoon Park, Jong Chul Ye
eess.IVcs.CVcs.LGarXiv:2004.05758v22020ECO: Quantized Training without Full-Precision Master Weights
Mahdi Nikdan, Amir Zandieh, Dan Alistarh +1
cs.CLcs.AIcs.LGarXiv:2601.22101v12026Lightweight Multi-scale Hierarchical Anomaly Detection and Localization for Geospatial Big Data Applications at the Edge
Thomas Benton Townsend, Joshua Bean, Benjamin K Tkach +2
eess.SPcs.LGeess.SYarXiv:2608.22648v12026DRPG (Decompose, Retrieve, Plan, Generate): An Agentic Framework for Academic Rebuttal
Peixuan Han, Yingjie Yu, Jingjun Xu +1
cs.LGarXiv:2601.18081v22026Visual Transformers: Token-based Image Representation and Processing for Computer Vision
Bichen Wu, Chenfeng Xu, Xiaoliang Dai +7
cs.CVcs.LGeess.IVarXiv:2006.03677v42020Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech
Vadim Popov, Ivan Vovk, Vladimir Gogoryan +2
cs.LGcs.CLstat.MLarXiv:2105.06337v22021RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
Hanze Dong, Wei Xiong, Deepanshu Goyal +7
cs.LGcs.AIcs.CLarXiv:2304.06767v42023Automatic Prompt Optimization with "Gradient Descent" and Beam Search
Reid Pryzant, Dan Iter, Jerry Li +3
cs.CLcs.AIcs.LGarXiv:2305.03495v22023Where Cognition Lives: Dissecting Emergent from Computed Function in a Minimal Complete Cognitive Architecture
Francisco M. Arrabal-Campos, Francisco G. Montoya, Alfredo Alcayde +1
cs.AIcs.CLcs.LGarXiv:2608.22347v12026Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning
Felipe Petroski Such, Vashisht Madhavan, Edoardo Conti +3
cs.NEcs.LGarXiv:1712.06567v32017Multi-Modal Fusion Transformer for End-to-End Autonomous Driving
Aditya Prakash, Kashyap Chitta, Andreas Geiger
cs.CVcs.AIcs.LGarXiv:2104.09224v12021Variational Information Distillation for Knowledge Transfer
Sungsoo Ahn, Shell Xu Hu, Andreas Damianou +2
cs.CVcs.AIcs.LGarXiv:1904.05835v12019Adversarial Attacks and Defenses in Images, Graphs and Text: A Review
Han Xu, Yao Ma, Haochen Liu +4
cs.LGcs.CRstat.MLarXiv:1909.08072v22019Big Self-Supervised Models Advance Medical Image Classification
Shekoofeh Azizi, Basil Mustafa, Fiona Ryan +9
eess.IVcs.CVcs.LGarXiv:2101.05224v22021RAMP: Reinforcement Adaptive Mixed Precision Quantization for Efficient On Device LLM Inference
Arpit Singh Gautam, Saurabh Jha
cs.LGcs.AIarXiv:2603.17891v12026Fairness in Machine Learning: A Survey
Simon Caton, Christian Haas
cs.LGstat.MLarXiv:2010.04053v12020Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules
Florian Rottach, Sebastian Schieferdecker, William Rudman +2
cs.LGcs.AIarXiv:2608.22642v12026Tracing the Unlabeled Storm: Cross-Variable Transfer in a Lagrangian Atmospheric JEPA Framework
K M Anirudh, S Sandeep, Hariprasad Kodamana
cs.LGphysics.geo-pharXiv:2608.22358v12026The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models
Taebong Kim, Youngsik Hong, Minsik Kim +3
cs.LGcs.AIarXiv:2608.22876v12026TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages
Jonathan H. Clark, Eunsol Choi, Michael Collins +4
cs.CLcs.LGarXiv:2003.05002v12020Summaries:한국어Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch
Le Yu, Bowen Yu, Haiyang Yu +2
cs.CLcs.LGarXiv:2311.03099v32023SimpleGPT: Improving GPT via A Simple Normalization Strategy
Marco Chen, Xianbiao Qi, Yelin He +2
cs.LGcs.CLcs.CVarXiv:2602.01212v12026On the (Statistical) Detection of Adversarial Examples
Kathrin Grosse, Praveen Manoharan, Nicolas Papernot +2
cs.CRcs.LGstat.MLarXiv:1702.06280v22017