Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
8,281 to 8,340 of 20,219
The Lessons of Developing Process Reward Models in Mathematical Reasoning
Zhenru Zhang, Chujie Zheng, Yangzhen Wu +6
cs.CLcs.AIcs.LGarXiv:2501.07301v22025The Right Tool for the Job: Matching Model and Instance Complexities
Roy Schwartz, Gabriel Stanovsky, Swabha Swayamdipta +2
cs.CLcs.LGarXiv:2004.07453v22020A Unified Approach to Error Bounds for Structured Convex Optimization Problems
Zirui Zhou, Anthony Man-Cho So
math.OCcs.LGmath.NAarXiv:1512.03518v12015Deep Probabilistic Programming
Dustin Tran, Matthew D. Hoffman, Rif A. Saurous +3
stat.MLcs.AIcs.LGarXiv:1701.03757v22017Diffusion Transformers with Representation Autoencoders
Boyang Zheng, Nanye Ma, Shengbang Tong +1
cs.CVcs.LGarXiv:2510.11690v12025Differentiable plasticity: training plastic neural networks with backpropagation
Thomas Miconi, Jeff Clune, Kenneth O. Stanley
cs.NEcs.LGstat.MLarXiv:1804.02464v32018Contrastive Learning for Label-Efficient Semantic Segmentation
Xiangyun Zhao, Raviteja Vemulapalli, Philip Mansfield +4
cs.CVcs.AIcs.LGarXiv:2012.06985v42020Few-Shot Class-Incremental Learning by Sampling Multi-Phase Tasks
Da-Wei Zhou, Han-Jia Ye, Liang Ma +3
cs.CVcs.LGarXiv:2203.17030v22022Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Kanishk Gandhi, Ayush Chakravarthy, Anikait Singh +2
cs.CLcs.LGarXiv:2503.01307v22025Overparameterized Nonlinear Learning: Gradient Descent Takes the Shortest Path?
Samet Oymak, Mahdi Soltanolkotabi
cs.LGmath.OCstat.MLarXiv:1812.10004v12018Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
Qizheng Zhang, Changran Hu, Shubhangi Upasani +10
cs.LGcs.AIcs.CLarXiv:2510.04618v32025Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models
Jingfeng Yao, Bin Yang, Xinggang Wang
cs.CVcs.LGarXiv:2501.01423v32025Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
Marianne Arriola, Aaron Gokaslan, Justin T. Chiu +5
cs.LGcs.AIarXiv:2503.09573v32025Who Should I Trust: AI or Myself? Leveraging Human and AI Correctness Likelihood to Promote Appropriate Trust in AI-Assisted Decision-Making
Shuai Ma, Ying Lei, Xinru Wang +4
cs.HCcs.AIcs.LGarXiv:2301.05809v12023Process Reinforcement through Implicit Rewards
Ganqu Cui, Lifan Yuan, Zefan Wang +22
cs.LGcs.AIcs.CLarXiv:2502.01456v22025TrajMind: Chaining Role-Specialized LoRAs for Fast-and-Slow Collective Trajectory Anomaly Diagnosis
Jiahao Wu, Zhenqun Yang, Chen Jason Zhang +1
cs.LGarXiv:2609.02540v12026Summaries:한국어From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution
Yuzhang Luo, Chenpeng Wang, Jianhui Chen +1
cs.CLcs.AIcs.LGarXiv:2609.02771v12026Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment
Chenyu Zhou, Qiliang Jiang, Shuning Wu +1
cs.LGcs.AIarXiv:2609.02417v12026Node Feature Extraction by Self-Supervised Multi-scale Neighborhood Prediction
Eli Chien, Wei-Cheng Chang, Cho-Jui Hsieh +4
cs.LGarXiv:2111.00064v32021Machine-Learning-Based Diagnostics of EEG Pathology
Lukas Alexander Wilhelm Gemein, Robin Tibor Schirrmeister, Patryk Chrabąszcz +5
eess.IVcs.LGeess.SParXiv:2002.05115v12020Molecule Attention Transformer
Łukasz Maziarka, Tomasz Danel, Sławomir Mucha +3
cs.LGphysics.comp-phstat.MLarXiv:2002.08264v12020DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search
Huajian Xin, Z. Z. Ren, Junxiao Song +14
cs.CLcs.AIcs.LGarXiv:2408.08152v12024Inferring deterministic causal relations
Povilas Daniusis, Dominik Janzing, Joris Mooij +4
cs.LGstat.MLarXiv:1203.3475v12012ELEVATER: A Benchmark and Toolkit for Evaluating Language-Augmented Visual Models
Chunyuan Li, Haotian Liu, Liunian Harold Li +8
cs.CVcs.CLcs.LGarXiv:2204.08790v62022MAZE: Data-Free Model Stealing Attack Using Zeroth-Order Gradient Estimation
Sanjay Kariyappa, Atul Prakash, Moinuddin Qureshi
stat.MLcs.LGarXiv:2005.03161v22020Trace Lasso: a trace norm regularization for correlated designs
Edouard Grave, Guillaume Obozinski, Francis Bach
cs.LGstat.MLarXiv:1109.1990v12011RAB: Provable Robustness Against Backdoor Attacks
Maurice Weber, Xiaojun Xu, Bojan Karlaš +2
cs.LGstat.MLarXiv:2003.08904v82020Summaries:한국어C3oT: Generating Shorter Chain-of-Thought without Compromising Effectiveness
Yu Kang, Xianghui Sun, Liangyu Chen +1
cs.CLcs.LGarXiv:2412.11664v12024Distributed Matrix Completion and Robust Factorization
Lester Mackey, Ameet Talwalkar, Michael I. Jordan
cs.LGcs.DSmath.NAarXiv:1107.0789v72011Dyna-Style Planning with Linear Function Approximation and Prioritized Sweeping
Richard S. Sutton, Csaba Szepesvari, Alborz Geramifard +1
cs.AIcs.LGeess.SYarXiv:1206.3285v12012What Is Worth Representing? Representational Empowerment for Continual Model Construction
Fei Dai, Hanqi Zhou, Alison Gopnik +1
cs.LGcs.AIarXiv:2609.02322v12026Meta-DETR: Image-Level Few-Shot Detection with Inter-Class Correlation Exploitation
Gongjie Zhang, Zhipeng Luo, Kaiwen Cui +2
cs.CVcs.AIcs.LGarXiv:2208.00219v12022Learning from Distributions via Support Measure Machines
Krikamol Muandet, Kenji Fukumizu, Francesco Dinuzzo +1
stat.MLcs.LGarXiv:1202.6504v22012Kernel Reboot: Breaking the Boundaries of Neural Tangent Kernels for Neural Fields
Amir Mallak, Alaa Maalouf, Lior Wolf +2
cs.LGcs.CVarXiv:2609.03117v12026Self-Supervised Generation of Spatial Audio for 360 Video
Pedro Morgado, Nuno Vasconcelos, Timothy Langlois +1
cs.SDcs.CVcs.LGarXiv:1809.02587v12018Can We Gain More from Orthogonality Regularizations in Training Deep CNNs?
Nitin Bansal, Xiaohan Chen, Zhangyang Wang
cs.LGcs.CVstat.MLarXiv:1810.09102v12018FlowSeq: Non-Autoregressive Conditional Sequence Generation with Generative Flow
Xuezhe Ma, Chunting Zhou, Xian Li +2
cs.CLcs.LGarXiv:1909.02480v32019Targeted Adversarial Examples for Black Box Audio Systems
Rohan Taori, Amog Kamsetty, Brenton Chu +1
cs.LGcs.CRcs.SDarXiv:1805.07820v22018FastSHAP: Real-Time Shapley Value Estimation
Neil Jethani, Mukund Sudarshan, Ian Covert +2
stat.MLcs.CVcs.LGarXiv:2107.07436v32021Learn from Whoever Is Right: Answer-Verified Multi-Teacher Distillation for Multi-Domain LLMs
Xixiang He, Xingming Li, Baiqi Wu +4
cs.LGcs.AIarXiv:2609.02548v12026$π^{*}_{0.6}$: a VLA That Learns From Experience
Physical Intelligence, Ali Amin, Raichelle Aniceto +53
cs.LGcs.ROarXiv:2511.14759v22025Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
Jingyang Yuan, Huazuo Gao, Damai Dai +12
cs.CLcs.AIcs.LGarXiv:2502.11089v22025LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model
Dilxat Muhtar, Zhenshi Li, Feng Gu +2
cs.CVcs.AIcs.LGarXiv:2402.02544v42024Multi-class Classification without Multi-class Labels
Yen-Chang Hsu, Zhaoyang Lv, Joel Schlosser +2
cs.LGcs.AIcs.CVarXiv:1901.00544v12019Recursive Value Learning for Long-Horizon Offline Goal-Conditioned RL
Hyeonseong Jeon, Youngwoon Lee
cs.LGcs.ROarXiv:2609.02237v12026AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
AgiBot-World-Contributors, Qingwen Bu, Jisong Cai +49
cs.ROcs.CVcs.LGarXiv:2503.06669v42025Spectral Initialization and Scheduled Graph Smoothness for Uncertain Knowledge Graph Completion
Md Abrar Jahin, Taufikur Rahman Fuad, Jay Pujara +1
cs.LGcs.AIarXiv:2609.02519v12026Interactive and Visual Prompt Engineering for Ad-hoc Task Adaptation with Large Language Models
Hendrik Strobelt, Albert Webson, Victor Sanh +4
cs.CLcs.HCcs.LGarXiv:2208.07852v12022Automated Concatenation of Embeddings for Structured Prediction
Xinyu Wang, Yong Jiang, Nguyen Bach +4
cs.CLcs.AIcs.LGarXiv:2010.05006v42020LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style Updates
Dmitrii Andriianov, Andrey Veprikov, Aleksandr Beznosikov
cs.LGarXiv:2609.02734v12026Graph Meta Learning via Local Subgraphs
Kexin Huang, Marinka Zitnik
cs.LGstat.MLarXiv:2006.07889v42020Reasoning Models Don't Always Say What They Think
Yanda Chen, Joe Benton, Ansh Radhakrishnan +12
cs.CLcs.AIcs.LGarXiv:2505.05410v12025Causal Bandits: Learning Good Interventions via Causal Inference
Finnian Lattimore, Tor Lattimore, Mark D. Reid
stat.MLcs.LGarXiv:1606.03203v12016Best-Arm Identification in Linear Bandits
Marta Soare, Alessandro Lazaric, Rémi Munos
cs.LGarXiv:1409.6110v22014Kimi K2: Open Agentic Intelligence
Kimi Team, Yifan Bai, Yiping Bao +197
cs.LGcs.AIcs.CLarXiv:2507.20534v22025Leveraging Big Data Analytics in Healthcare Enhancement: Trends, Challenges and Opportunities
Arshia Rehman, Saeeda Naz, Imran Razzak
stat.OTcs.LGstat.MLarXiv:2004.09010v12020Getting aligned on representational alignment
Ilia Sucholutsky, Lukas Muttenthaler, Adrian Weller +30
q-bio.NCcs.AIcs.LGarXiv:2310.13018v32023Brain-like associative learning using a nanoscale non-volatile phase change synaptic device array
Sukru Burc Eryilmaz, Duygu Kuzum, Rakesh Jeyasingh +4
cs.NEcond-mat.mtrl-scics.LGarXiv:1406.4951v42014Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Jingcheng Hu, Yinmin Zhang, Qi Han +3
cs.LGcs.CLarXiv:2503.24290v22025Scalable Kronecker-Fisher Approximation: Efficient Hessian Analysis for Billion-Parameter Language Models Compression
Viacheslav Yusupov, Daria Cherniuk, Evgeny Frolov
cs.LGcs.AIcs.CLarXiv:2609.02451v12026