Machine Learning
Papers filed under cs.LG on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
121 to 180 of 19,911
Training Generative Adversarial Networks with Limited Data
Tero Karras, Miika Aittala, Janne Hellsten +3
cs.CVcs.LGcs.NEarXiv:2006.06676v22020Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
Sreyan Ghosh, Zhifeng Kong, Sonal Kumar +6
cs.SDcs.CLcs.LGarXiv:2503.03983v12025Deep Learning is Not So Mysterious or Different
Andrew Gordon Wilson
cs.LGstat.MLarXiv:2503.02113v22025Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus +9
cs.CLcs.LGarXiv:2005.11401v42020Summaries:한국어Dataset Distillation with Neural Characteristic Function: A Minmax Perspective
Shaobo Wang, Yicun Yang, Zhiyuan Liu +4
cs.CVcs.AIcs.LGarXiv:2502.20653v12025Summaries:한국어Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment
Harrish Thasarathan, Julian Forsyth, Thomas Fel +2
cs.CVcs.LGarXiv:2502.03714v22025Maximum Density Divergence for Domain Adaptation
Li Jingjing, Chen Erpeng, Ding Zhengming +3
cs.CVcs.LGcs.MMarXiv:2004.12615v12020Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models
Junlin Fang, Do Nguyen-Thanh, Xiaogang Xu +2
cs.AIcs.LGarXiv:2607.22098v12026Don't Judge an Object by Its Context: Learning to Overcome Contextual Bias
Krishna Kumar Singh, Dhruv Mahajan, Kristen Grauman +3
cs.CVcs.LGarXiv:2001.03152v22020Frontier Models are Capable of In-context Scheming
Alexander Meinke, Bronson Schoen, Jérémy Scheurer +3
cs.AIcs.LGarXiv:2412.04984v22024Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training
Zijian Zhang, Rizhen Hu, Athanasios Glentis +4
cs.LGcs.CLarXiv:2607.01232v22026Omni-Scale CNNs: a simple and effective kernel size configuration for time series classification
Wensi Tang, Guodong Long, Lu Liu +3
cs.LGstat.MLarXiv:2002.10061v32020Beyond IID: How General Are Tabular Foundation Models, Really?
Lennart Purucker, Andrej Tschalzev, Nick Erickson +7
cs.LGcs.AIarXiv:2606.30410v12026Summaries:한국어Explaining Explanations: Axiomatic Feature Interactions for Deep Networks
Joseph D. Janizek, Pascal Sturmfels, Su-In Lee
cs.LGstat.MLarXiv:2002.04138v32020PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation
Shaowei Liu, Zhongzheng Ren, Saurabh Gupta +1
cs.CVcs.AIcs.LGarXiv:2409.18964v12024LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
Yijia Xiao, Edward Sun, Tianyu Liu +1
cs.AIcs.CLcs.CVarXiv:2407.04973v12024Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Xin Lai, Zhuotao Tian, Yukang Chen +3
cs.LGcs.AIcs.CLarXiv:2406.18629v12024Variational Mixture-of-Experts Autoencoders for Multi-Modal Deep Generative Models
Yuge Shi, N. Siddharth, Brooks Paige +1
stat.MLcs.LGarXiv:1911.03393v12019The Heidelberg spiking datasets for the systematic evaluation of spiking neural networks
Benjamin Cramer, Yannik Stradmann, Johannes Schemmel +1
cs.NEcs.LGq-bio.NCarXiv:1910.07407v32019CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression
Morayo Danielle Adeyemi, Ryan A. Rossi, Franck Dernoncourt
cs.CLcs.AIcs.LGarXiv:2606.24083v12026LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
Mostafa Elhoushi, Akshat Shrivastava, Diana Liskovich +10
cs.CLcs.AIcs.LGarXiv:2404.16710v42024DeepGCNs: Making GCNs Go as Deep as CNNs
Guohao Li, Matthias Müller, Guocheng Qian +4
cs.CVcs.LGeess.IVarXiv:1910.06849v32019InceptionTime: Finding AlexNet for Time Series Classification
Hassan Ismail Fawaz, Benjamin Lucas, Germain Forestier +7
cs.LGstat.MLarXiv:1909.04939v32019ControlNet++: Improving Conditional Controls with Efficient Consistency Feedback
Ming Li, Taojiannan Yang, Huafeng Kuang +4
cs.CVcs.AIcs.LGarXiv:2404.07987v42024Anomaly Detection in Video Sequence with Appearance-Motion Correspondence
Trong Nguyen Nguyen, Jean Meunier
cs.CVcs.LGcs.NEarXiv:1908.06351v12019Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking
Jiwan Chung, JiHyuk Byun, Vibhav Vineet +1
cs.AIcs.LGarXiv:2606.15673v22026Local Differential Privacy for Deep Learning
M. A. P. Chamikara, P. Bertok, I. Khalil +3
cs.LGcs.CRarXiv:1908.02997v32019Optuna: A Next-generation Hyperparameter Optimization Framework
Takuya Akiba, Shotaro Sano, Toshihiko Yanase +2
cs.LGstat.MLarXiv:1907.10902v12019Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
Zhichen Dong, Zhanhui Zhou, Chao Yang +2
cs.CLcs.AIcs.CYarXiv:2402.09283v32024Hierarchically Structured Meta-learning
Huaxiu Yao, Ying Wei, Junzhou Huang +1
cs.LGstat.MLarXiv:1905.05301v22019SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
Dongyang Liu, Renrui Zhang, Longtian Qiu +16
cs.CVcs.AIcs.CLarXiv:2402.05935v32024InstructIR: High-Quality Image Restoration Following Human Instructions
Marcos V. Conde, Gregor Geigle, Radu Timofte
cs.CVcs.LGeess.IVarXiv:2401.16468v52024A Principle of Targeted Intervention for Multi-Agent Reinforcement Learning
Anjie Liu, Jianhong Wang, Samuel Kaski +2
cs.AIcs.LGcs.MAarXiv:2510.17697v42025Fair Transfer of Multiple Style Attributes in Text
Karan Dabas, Nishtha Madan, Vijay Arya +3
cs.CLcs.AIcs.LGarXiv:2001.06693v12020Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration
Soham De, Anirbit Mukherjee, Enayat Ullah
cs.LGmath.OCstat.MLarXiv:1807.06766v32018Training Language Models with Language Feedback
Jérémy Scheurer, Jon Ander Campos, Jun Shern Chan +3
cs.CLcs.AIcs.LGarXiv:2204.14146v42022PhiNets: Brain-inspired Non-contrastive Learning Based on Temporal Prediction Hypothesis
Satoki Ishikawa, Makoto Yamada, Han Bao +1
cs.LGarXiv:2405.14650v22024Brain-Inspired Stochastic Joint Embedding Representation Learning
Makoto Yamada, Kian Ming A. Chai, Ayoub Rhim +3
cs.CVcs.AIcs.LGarXiv:2505.11129v22025SkyJEPA: Learning Long-Horizon World Models for Zero-Shot Sim-to-Real Control of Quadrotors
Pratyaksh Rao, Wancong Zhang, Randall Balestriero +2
cs.ROcs.LGarXiv:2606.23444v22026Siamese Masked Autoencoders
Agrim Gupta, Jiajun Wu, Jia Deng +1
cs.CVcs.LGarXiv:2305.14344v12023Visual Representation Learning with Stochastic Frame Prediction
Huiwon Jang, Dongyoung Kim, Junsu Kim +3
cs.CVcs.AIcs.LGarXiv:2406.07398v22024Pushing Stochastic Gradient towards Second-Order Methods -- Backpropagation Learning with Transformations in Nonlinearities
Tommi Vatanen, Tapani Raiko, Harri Valpola +1
cs.LGcs.CVstat.MLarXiv:1301.3476v32013TransferTraj: A Vehicle Trajectory Learning Model for Region and Task Transferability
Tonglong Wei, Yan Lin, Zeyu Zhou +6
cs.LGarXiv:2505.12672v12025Stable Learning Using Spiking Neural Networks Equipped With Affine Encoders and Decoders
A. Martina Neuman, Dominik Dold, Philipp Christian Petersen
cs.NEcs.LGmath.FAarXiv:2404.04549v32024Towards Automatic Concept-based Explanations
Amirata Ghorbani, James Wexler, James Zou +1
stat.MLcs.CVcs.LGarXiv:1902.03129v32019Revisiting minimum description length complexity in overparameterized models
Raaz Dwivedi, Chandan Singh, Bin Yu +1
cs.LGcs.ITmath.STarXiv:2006.10189v42020DoMINO: A Decomposable Multi-scale Iterative Neural Operator for Modeling Large Scale Engineering Simulations
Rishikesh Ranade, Mohammad Amin Nabian, Kaustubh Tangsali +4
cs.LGphysics.comp-pharXiv:2501.13350v12025TripNet: Learning Large-scale High-fidelity 3D Car Aerodynamics with Triplane Networks
Qian Chen, Mohamed Elrefaie, Angela Dai +1
physics.flu-dyncs.LGarXiv:2503.17400v22025A Tropical Approach to Neural Networks with Piecewise Linear Activations
Vasileios Charisopoulos, Petros Maragos
stat.MLcs.LGarXiv:1805.08749v22018AgentInstruct: Toward Generative Teaching with Agentic Flows
Arindam Mitra, Luciano Del Corro, Guoqing Zheng +11
cs.AIcs.CLcs.LGarXiv:2407.03502v12024BoardgameQA: A Dataset for Natural Language Reasoning with Contradictory Information
Mehran Kazemi, Quan Yuan, Deepti Bhatia +4
cs.CLcs.AIcs.LGarXiv:2306.07934v12023Soft Tokens, Hard Truths
Natasha Butt, Ariel Kwiatkowski, Ismail Labiad +2
cs.CLcs.AIcs.LGarXiv:2509.19170v22025Causal Language Modeling Can Elicit Search and Reasoning Capabilities on Logic Puzzles
Kulin Shah, Nishanth Dikkala, Xin Wang +1
cs.LGcs.CLarXiv:2409.10502v12024Monitoring Latent World States in Language Models with Propositional Probes
Jiahai Feng, Stuart Russell, Jacob Steinhardt
cs.CLcs.LGarXiv:2406.19501v22024Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning
Jifan Zhang, Lalit Jain, Yang Guo +9
cs.LGcs.AIcs.CLarXiv:2406.10522v22024Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
Bilal Chughtai, Alan Cooney, Neel Nanda
cs.LGcs.CLarXiv:2402.07321v12024Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
Shahar Katz, Yonatan Belinkov, Mor Geva +1
cs.CLcs.AIcs.LGarXiv:2402.12865v12024Measuring Semantic Similarity by Latent Relational Analysis
Peter D. Turney
cs.LGcs.CLcs.IRarXiv:cs/0508053v12005SyGuS-Comp 2016: Results and Analysis
Rajeev Alur, Dana Fisman, Rishabh Singh +1
cs.SEcs.LGcs.LOarXiv:1611.07627v12016TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
Zhangchen Xu, Yuetai Li, Fengqing Jiang +4
cs.LGcs.AIcs.CLarXiv:2505.14625v22025