Machine Learning

Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

8,281 to 8,340 of 20,219

  1. The Lessons of Developing Process Reward Models in Mathematical Reasoning

    Zhenru Zhang, Chujie Zheng, Yangzhen Wu +6

    cs.CLcs.AIcs.LGarXiv:2501.07301v22025
  2. The Right Tool for the Job: Matching Model and Instance Complexities

    Roy Schwartz, Gabriel Stanovsky, Swabha Swayamdipta +2

    cs.CLcs.LGarXiv:2004.07453v22020
  3. A Unified Approach to Error Bounds for Structured Convex Optimization Problems

    Zirui Zhou, Anthony Man-Cho So

    math.OCcs.LGmath.NAarXiv:1512.03518v12015
  4. Deep Probabilistic Programming

    Dustin Tran, Matthew D. Hoffman, Rif A. Saurous +3

    stat.MLcs.AIcs.LGarXiv:1701.03757v22017
  5. Diffusion Transformers with Representation Autoencoders

    Boyang Zheng, Nanye Ma, Shengbang Tong +1

    cs.CVcs.LGarXiv:2510.11690v12025
  6. Differentiable plasticity: training plastic neural networks with backpropagation

    Thomas Miconi, Jeff Clune, Kenneth O. Stanley

    cs.NEcs.LGstat.MLarXiv:1804.02464v32018
  7. Contrastive Learning for Label-Efficient Semantic Segmentation

    Xiangyun Zhao, Raviteja Vemulapalli, Philip Mansfield +4

    cs.CVcs.AIcs.LGarXiv:2012.06985v42020
  8. Few-Shot Class-Incremental Learning by Sampling Multi-Phase Tasks

    Da-Wei Zhou, Han-Jia Ye, Liang Ma +3

    cs.CVcs.LGarXiv:2203.17030v22022
  9. Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

    Kanishk Gandhi, Ayush Chakravarthy, Anikait Singh +2

    cs.CLcs.LGarXiv:2503.01307v22025
  10. Overparameterized Nonlinear Learning: Gradient Descent Takes the Shortest Path?

    Samet Oymak, Mahdi Soltanolkotabi

    cs.LGmath.OCstat.MLarXiv:1812.10004v12018
  11. Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models

    Qizheng Zhang, Changran Hu, Shubhangi Upasani +10

    cs.LGcs.AIcs.CLarXiv:2510.04618v32025
  12. Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models

    Jingfeng Yao, Bin Yang, Xinggang Wang

    cs.CVcs.LGarXiv:2501.01423v32025
  13. Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models

    Marianne Arriola, Aaron Gokaslan, Justin T. Chiu +5

    cs.LGcs.AIarXiv:2503.09573v32025
  14. Who Should I Trust: AI or Myself? Leveraging Human and AI Correctness Likelihood to Promote Appropriate Trust in AI-Assisted Decision-Making

    Shuai Ma, Ying Lei, Xinru Wang +4

    cs.HCcs.AIcs.LGarXiv:2301.05809v12023
  15. Process Reinforcement through Implicit Rewards

    Ganqu Cui, Lifan Yuan, Zefan Wang +22

    cs.LGcs.AIcs.CLarXiv:2502.01456v22025
  16. TrajMind: Chaining Role-Specialized LoRAs for Fast-and-Slow Collective Trajectory Anomaly Diagnosis

    Jiahao Wu, Zhenqun Yang, Chen Jason Zhang +1

    cs.LGarXiv:2609.02540v12026
    Summaries:한국어
  17. From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution

    Yuzhang Luo, Chenpeng Wang, Jianhui Chen +1

    cs.CLcs.AIcs.LGarXiv:2609.02771v12026
  18. Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment

    Chenyu Zhou, Qiliang Jiang, Shuning Wu +1

    cs.LGcs.AIarXiv:2609.02417v12026
  19. Node Feature Extraction by Self-Supervised Multi-scale Neighborhood Prediction

    Eli Chien, Wei-Cheng Chang, Cho-Jui Hsieh +4

    cs.LGarXiv:2111.00064v32021
  20. Machine-Learning-Based Diagnostics of EEG Pathology

    Lukas Alexander Wilhelm Gemein, Robin Tibor Schirrmeister, Patryk Chrabąszcz +5

    eess.IVcs.LGeess.SParXiv:2002.05115v12020
  21. Molecule Attention Transformer

    Łukasz Maziarka, Tomasz Danel, Sławomir Mucha +3

    cs.LGphysics.comp-phstat.MLarXiv:2002.08264v12020
  22. DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search

    Huajian Xin, Z. Z. Ren, Junxiao Song +14

    cs.CLcs.AIcs.LGarXiv:2408.08152v12024
  23. Inferring deterministic causal relations

    Povilas Daniusis, Dominik Janzing, Joris Mooij +4

    cs.LGstat.MLarXiv:1203.3475v12012
  24. ELEVATER: A Benchmark and Toolkit for Evaluating Language-Augmented Visual Models

    Chunyuan Li, Haotian Liu, Liunian Harold Li +8

    cs.CVcs.CLcs.LGarXiv:2204.08790v62022
  25. MAZE: Data-Free Model Stealing Attack Using Zeroth-Order Gradient Estimation

    Sanjay Kariyappa, Atul Prakash, Moinuddin Qureshi

    stat.MLcs.LGarXiv:2005.03161v22020
  26. Trace Lasso: a trace norm regularization for correlated designs

    Edouard Grave, Guillaume Obozinski, Francis Bach

    cs.LGstat.MLarXiv:1109.1990v12011
  27. RAB: Provable Robustness Against Backdoor Attacks

    Maurice Weber, Xiaojun Xu, Bojan Karlaš +2

    cs.LGstat.MLarXiv:2003.08904v82020
    Summaries:한국어
  28. C3oT: Generating Shorter Chain-of-Thought without Compromising Effectiveness

    Yu Kang, Xianghui Sun, Liangyu Chen +1

    cs.CLcs.LGarXiv:2412.11664v12024
  29. Distributed Matrix Completion and Robust Factorization

    Lester Mackey, Ameet Talwalkar, Michael I. Jordan

    cs.LGcs.DSmath.NAarXiv:1107.0789v72011
  30. Dyna-Style Planning with Linear Function Approximation and Prioritized Sweeping

    Richard S. Sutton, Csaba Szepesvari, Alborz Geramifard +1

    cs.AIcs.LGeess.SYarXiv:1206.3285v12012
  31. What Is Worth Representing? Representational Empowerment for Continual Model Construction

    Fei Dai, Hanqi Zhou, Alison Gopnik +1

    cs.LGcs.AIarXiv:2609.02322v12026
  32. Meta-DETR: Image-Level Few-Shot Detection with Inter-Class Correlation Exploitation

    Gongjie Zhang, Zhipeng Luo, Kaiwen Cui +2

    cs.CVcs.AIcs.LGarXiv:2208.00219v12022
  33. Learning from Distributions via Support Measure Machines

    Krikamol Muandet, Kenji Fukumizu, Francesco Dinuzzo +1

    stat.MLcs.LGarXiv:1202.6504v22012
  34. Kernel Reboot: Breaking the Boundaries of Neural Tangent Kernels for Neural Fields

    Amir Mallak, Alaa Maalouf, Lior Wolf +2

    cs.LGcs.CVarXiv:2609.03117v12026
  35. Self-Supervised Generation of Spatial Audio for 360 Video

    Pedro Morgado, Nuno Vasconcelos, Timothy Langlois +1

    cs.SDcs.CVcs.LGarXiv:1809.02587v12018
  36. Can We Gain More from Orthogonality Regularizations in Training Deep CNNs?

    Nitin Bansal, Xiaohan Chen, Zhangyang Wang

    cs.LGcs.CVstat.MLarXiv:1810.09102v12018
  37. FlowSeq: Non-Autoregressive Conditional Sequence Generation with Generative Flow

    Xuezhe Ma, Chunting Zhou, Xian Li +2

    cs.CLcs.LGarXiv:1909.02480v32019
  38. Targeted Adversarial Examples for Black Box Audio Systems

    Rohan Taori, Amog Kamsetty, Brenton Chu +1

    cs.LGcs.CRcs.SDarXiv:1805.07820v22018
  39. FastSHAP: Real-Time Shapley Value Estimation

    Neil Jethani, Mukund Sudarshan, Ian Covert +2

    stat.MLcs.CVcs.LGarXiv:2107.07436v32021
  40. Learn from Whoever Is Right: Answer-Verified Multi-Teacher Distillation for Multi-Domain LLMs

    Xixiang He, Xingming Li, Baiqi Wu +4

    cs.LGcs.AIarXiv:2609.02548v12026
  41. $π^{*}_{0.6}$: a VLA That Learns From Experience

    Physical Intelligence, Ali Amin, Raichelle Aniceto +53

    cs.LGcs.ROarXiv:2511.14759v22025
  42. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

    Jingyang Yuan, Huazuo Gao, Damai Dai +12

    cs.CLcs.AIcs.LGarXiv:2502.11089v22025
  43. LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model

    Dilxat Muhtar, Zhenshi Li, Feng Gu +2

    cs.CVcs.AIcs.LGarXiv:2402.02544v42024
  44. Multi-class Classification without Multi-class Labels

    Yen-Chang Hsu, Zhaoyang Lv, Joel Schlosser +2

    cs.LGcs.AIcs.CVarXiv:1901.00544v12019
  45. Recursive Value Learning for Long-Horizon Offline Goal-Conditioned RL

    Hyeonseong Jeon, Youngwoon Lee

    cs.LGcs.ROarXiv:2609.02237v12026
  46. AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

    AgiBot-World-Contributors, Qingwen Bu, Jisong Cai +49

    cs.ROcs.CVcs.LGarXiv:2503.06669v42025
  47. Spectral Initialization and Scheduled Graph Smoothness for Uncertain Knowledge Graph Completion

    Md Abrar Jahin, Taufikur Rahman Fuad, Jay Pujara +1

    cs.LGcs.AIarXiv:2609.02519v12026
  48. Interactive and Visual Prompt Engineering for Ad-hoc Task Adaptation with Large Language Models

    Hendrik Strobelt, Albert Webson, Victor Sanh +4

    cs.CLcs.HCcs.LGarXiv:2208.07852v12022
  49. Automated Concatenation of Embeddings for Structured Prediction

    Xinyu Wang, Yong Jiang, Nguyen Bach +4

    cs.CLcs.AIcs.LGarXiv:2010.05006v42020
  50. LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style Updates

    Dmitrii Andriianov, Andrey Veprikov, Aleksandr Beznosikov

    cs.LGarXiv:2609.02734v12026
  51. Graph Meta Learning via Local Subgraphs

    Kexin Huang, Marinka Zitnik

    cs.LGstat.MLarXiv:2006.07889v42020
  52. Reasoning Models Don't Always Say What They Think

    Yanda Chen, Joe Benton, Ansh Radhakrishnan +12

    cs.CLcs.AIcs.LGarXiv:2505.05410v12025
  53. Causal Bandits: Learning Good Interventions via Causal Inference

    Finnian Lattimore, Tor Lattimore, Mark D. Reid

    stat.MLcs.LGarXiv:1606.03203v12016
  54. Best-Arm Identification in Linear Bandits

    Marta Soare, Alessandro Lazaric, Rémi Munos

    cs.LGarXiv:1409.6110v22014
  55. Kimi K2: Open Agentic Intelligence

    Kimi Team, Yifan Bai, Yiping Bao +197

    cs.LGcs.AIcs.CLarXiv:2507.20534v22025
  56. Leveraging Big Data Analytics in Healthcare Enhancement: Trends, Challenges and Opportunities

    Arshia Rehman, Saeeda Naz, Imran Razzak

    stat.OTcs.LGstat.MLarXiv:2004.09010v12020
  57. Getting aligned on representational alignment

    Ilia Sucholutsky, Lukas Muttenthaler, Adrian Weller +30

    q-bio.NCcs.AIcs.LGarXiv:2310.13018v32023
  58. Brain-like associative learning using a nanoscale non-volatile phase change synaptic device array

    Sukru Burc Eryilmaz, Duygu Kuzum, Rakesh Jeyasingh +4

    cs.NEcond-mat.mtrl-scics.LGarXiv:1406.4951v42014
  59. Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

    Jingcheng Hu, Yinmin Zhang, Qi Han +3

    cs.LGcs.CLarXiv:2503.24290v22025
  60. Scalable Kronecker-Fisher Approximation: Efficient Hessian Analysis for Billion-Parameter Language Models Compression

    Viacheslav Yusupov, Daria Cherniuk, Evgeny Frolov

    cs.LGcs.AIcs.CLarXiv:2609.02451v12026