Every paper with a summary
Every arXiv paper Paperlayer has summarized so far. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
21,181 to 21,240 of 61,335
Reinforcement Learning with Action Chunking
Qiyang Li, Zhiyuan Zhou, Sergey Levine
cs.LGcs.AIcs.ROarXiv:2507.07969v42025VIBE: Visual Instruction Based Editor
Grigorii Alekseenko, Aleksandr Gordeev, Irina Tolstykh +7
cs.CVcs.AIcs.LGarXiv:2601.02242v12026SampoNLP: A Self-Referential Toolkit for Morphological Analysis of Subword Tokenizers
Iaroslav Chelombitko, Ekaterina Chelombitko, Aleksey Komissarov
cs.CLcs.IRcs.LGarXiv:2601.04469v12026GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving
Lloyd Russell, Anthony Hu, Lorenzo Bertoni +4
cs.CVcs.AIcs.ROarXiv:2503.20523v12025OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction
Lujie Yang, Xiaoyu Huang, Zhen Wu +6
cs.ROcs.AIcs.LGarXiv:2509.26633v32025MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
Kunlun Zhu, Hongyi Du, Zhaochen Hong +8
cs.MAcs.AIcs.CLarXiv:2503.01935v12025Learning Smooth and Expressive Interatomic Potentials for Physical Property Prediction
Xiang Fu, Brandon M. Wood, Luis Barroso-Luque +4
physics.comp-phcs.LGarXiv:2502.12147v22025Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation
Ruotong Wang, Zihao Zhu, Siwei Lyu +2
cs.CVcs.AIarXiv:2609.00206v12026Learning a visuomotor controller for real world robotic grasping using simulated depth images
Ulrich Viereck, Andreas ten Pas, Kate Saenko +1
cs.ROcs.AIarXiv:1706.04652v32017Learning to Act from Actionless Videos through Dense Correspondences
Po-Chen Ko, Jiayuan Mao, Yilun Du +2
cs.ROcs.CVcs.LGarXiv:2310.08576v12023CoverM: Read alignment statistics for metagenomics
Samuel T. N. Aroney, Rhys J. P. Newell, Jakob N. Nissen +3
q-bio.GNarXiv:2501.11217v12025Poseidon: Efficient Foundation Models for PDEs
Maximilian Herde, Bogdan Raonić, Tobias Rohner +4
cs.LGarXiv:2405.19101v22024PaliGemma: A versatile 3B VLM for transfer
Lucas Beyer, Andreas Steiner, André Susano Pinto +32
cs.CVcs.AIcs.CLarXiv:2407.07726v22024DataComp-LM: In search of the next generation of training sets for language models
Jeffrey Li, Alex Fang, Georgios Smyrnis +56
cs.LGcs.CLarXiv:2406.11794v42024Simplified and Generalized Masked Diffusion for Discrete Data
Jiaxin Shi, Kehang Han, Zhe Wang +2
cs.LGstat.MLarXiv:2406.04329v42024YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-time Object Detection
Yuming Chen, Xinbin Yuan, Jiabao Wang +4
cs.CVarXiv:2308.05480v22023Joint Beam Training and Positioning For Intelligent Reflecting Surfaces Assisted Millimeter Wave Communications
Wei Wang, Wei Zhang
cs.ITeess.SParXiv:2009.03536v22020ReFT: Representation Finetuning for Language Models
Zhengxuan Wu, Aryaman Arora, Zheng Wang +4
cs.CLcs.AIcs.LGarXiv:2404.03592v32024Diffusion Model-Based Image Editing: A Survey
Yi Huang, Jiancheng Huang, Yifan Liu +7
cs.CVarXiv:2402.17525v42024Otter: A Multi-Modal Model with In-Context Instruction Tuning
Bo Li, Yuanhan Zhang, Liangyu Chen +5
cs.CVcs.CLarXiv:2305.03726v22023Integrated Sensing and Communication Signals Toward 5G-A and 6G: A Survey
Zhiqing Wei, Hanyang Qu, Yuan Wang +6
cs.ITeess.SParXiv:2301.03857v32023ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
Xin Men, Mingyu Xu, Qingyu Zhang +5
cs.CLarXiv:2403.03853v32024HCF-Net: Hierarchical Context Fusion Network for Infrared Small Object Detection
Shibiao Xu, ShuChen Zheng, Wenhao Xu +6
cs.CVarXiv:2403.10778v12024StylizedNeRF: Consistent 3D Scene Stylization as Stylized NeRF via 2D-3D Mutual Learning
Yi-Hua Huang, Yue He, Yu-Jie Yuan +2
cs.GRcs.CVarXiv:2205.12183v22022V-STaR: Training Verifiers for Self-Taught Reasoners
Arian Hosseini, Xingdi Yuan, Nikolay Malkin +3
cs.LGcs.AIcs.CLarXiv:2402.06457v22024Unexpected Improvements to Expected Improvement for Bayesian Optimization
Sebastian Ament, Samuel Daulton, David Eriksson +2
cs.LGmath.NAstat.MLarXiv:2310.20708v32023Summaries:한국어Localization Distillation for Dense Object Detection
Zhaohui Zheng, Rongguang Ye, Ping Wang +4
cs.CVarXiv:2102.12252v42021Advancements in Generative AI: A Comprehensive Review of GANs, GPT, Autoencoders, Diffusion Model, and Transformers
Staphord Bengesi, Hoda El-Sayed, Md Kamruzzaman Sarker +3
cs.LGcs.AIarXiv:2311.10242v22023FastVGGT: Training-Free Acceleration of Visual Geometry Transformer
You Shen, Zhipeng Zhang, Yansong Qu +4
cs.CVarXiv:2509.02560v22025RIS-Aided Cell-Free Massive MIMO Systems for 6G: Fundamentals, System Design, and Applications
Enyu Shi, Jiayi Zhang, Hongyang Du +5
cs.ITeess.SParXiv:2310.00263v32023Convolutional Recurrent Neural Networks for Small-Footprint Keyword Spotting
Sercan O. Arik, Markus Kliegl, Rewon Child +5
cs.CLcs.AIcs.LGarXiv:1703.05390v32017Benchmarking Distributed Stream Data Processing Systems
Jeyhun Karimov, Tilmann Rabl, Asterios Katsifodimos +3
cs.DBarXiv:1802.08496v22018Being-H0.7: A Latent World-Action Model from Egocentric Videos
Hao Luo, Wanpeng Zhang, Yicheng Feng +6
cs.ROcs.CVcs.LGarXiv:2605.00078v12026OmniControl: Control Any Joint at Any Time for Human Motion Generation
Yiming Xie, Varun Jampani, Lei Zhong +2
cs.CVcs.GRarXiv:2310.08580v22023iTransformer: Inverted Transformers Are Effective for Time Series Forecasting
Yong Liu, Tengge Hu, Haoran Zhang +4
cs.LGarXiv:2310.06625v42023Ferret: Refer and Ground Anything Anywhere at Any Granularity
Haoxuan You, Haotian Zhang, Zhe Gan +6
cs.CVcs.CLarXiv:2310.07704v12023Collaborative Regression of Expressive Bodies using Moderation
Yao Feng, Vasileios Choutas, Timo Bolkart +2
cs.CVarXiv:2105.05301v22021Large Language Models as Optimizers
Chengrun Yang, Xuezhi Wang, Yifeng Lu +4
cs.LGcs.AIcs.CLarXiv:2309.03409v32023LIMR: Less is More for RL Scaling
Xuefeng Li, Haoyang Zou, Pengfei Liu
cs.LGcs.AIcs.CLarXiv:2502.11886v12025Jointly Localizing and Describing Events for Dense Video Captioning
Yehao Li, Ting Yao, Yingwei Pan +2
cs.CVarXiv:1804.08274v12018Negative Momentum for Improved Game Dynamics
Gauthier Gidel, Reyhane Askari Hemmat, Mohammad Pezeshki +4
cs.LGstat.MLarXiv:1807.04740v52018Convolutional Bypasses Are Better Vision Transformer Adapters
Shibo Jie, Zhi-Hong Deng
cs.CVarXiv:2207.07039v32022A Comprehensive Survey of Deep Transfer Learning for Anomaly Detection in Industrial Time Series: Methods, Applications, and Directions
Peng Yan, Ahmed Abdulkadir, Paul-Philipp Luley +4
cs.LGcs.AIarXiv:2307.05638v22023Language is All a Graph Needs
Ruosong Ye, Caiqi Zhang, Runhui Wang +2
cs.CLcs.AIcs.IRarXiv:2308.07134v52023Feature Decomposition and Reconstruction Learning for Effective Facial Expression Recognition
Delian Ruan, Yan Yan, Shenqi Lai +3
cs.CVarXiv:2104.05160v22021Is Self-Repair a Silver Bullet for Code Generation?
Theo X. Olausson, Jeevana Priya Inala, Chenglong Wang +2
cs.CLcs.AIcs.PLarXiv:2306.09896v52023Portfolio Risk Bounds without Cross-Asset Return Covariances: Distributional Fields from Language-Model Representations
Marcus Gawronsky, Chun-Sung Huang
q-fin.STarXiv:2608.29692v12026Unleashing the Power of Edge-Cloud Generative AI in Mobile Networks: A Survey of AIGC Services
Minrui Xu, Hongyang Du, Dusit Niyato +9
cs.NIarXiv:2303.16129v42023DragDiffusion: Harnessing Diffusion Models for Interactive Point-based Image Editing
Yujun Shi, Chuhui Xue, Jun Hao Liew +5
cs.CVcs.LGarXiv:2306.14435v62023ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization
Xixi Wu, Kuan Li, Yida Zhao +13
cs.CLarXiv:2509.13313v32025Unified Vision-Language-Action Model
Yuqi Wang, Xinghang Li, Wenxuan Wang +5
cs.CVcs.ROarXiv:2506.19850v12025Can Language Models Solve Graph Problems in Natural Language?
Heng Wang, Shangbin Feng, Tianxing He +3
cs.CLcs.AIarXiv:2305.10037v32023mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Qinghao Ye, Haiyang Xu, Guohai Xu +15
cs.CLcs.CVcs.LGarXiv:2304.14178v32023Interactive and Explainable Region-guided Radiology Report Generation
Tim Tanida, Philip Müller, Georgios Kaissis +1
cs.CVcs.CLcs.LGarXiv:2304.08295v12023GraphPrompt: Unifying Pre-Training and Downstream Tasks for Graph Neural Networks
Zemin Liu, Xingtong Yu, Yuan Fang +1
cs.LGcs.CLarXiv:2302.08043v32023Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning
Thomas Carta, Clément Romac, Thomas Wolf +3
cs.LGarXiv:2302.02662v52023Experimentally realized in situ backpropagation for deep learning in nanophotonic neural networks
Sunil Pai, Zhanghao Sun, Tyler W. Hughes +11
cs.ETcs.LGphysics.opticsarXiv:2205.08501v12022Navigating to Objects in the Real World
Theophile Gervet, Soumith Chintala, Dhruv Batra +2
cs.ROcs.CVcs.LGarXiv:2212.00922v12022Open-vocabulary Queryable Scene Representations for Real World Planning
Boyuan Chen, Fei Xia, Brian Ichter +5
cs.ROcs.AIcs.CVarXiv:2209.09874v22022Data Augmentation techniques in time series domain: A survey and taxonomy
Guillermo Iglesias, Edgar Talavera, Ángel González-Prieto +2
cs.LGcs.AIarXiv:2206.13508v42022