Every paper with a summary

Every arXiv paper Paperlayer has summarized so far. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

21,181 to 21,240 of 61,335

  1. Reinforcement Learning with Action Chunking

    Qiyang Li, Zhiyuan Zhou, Sergey Levine

    cs.LGcs.AIcs.ROarXiv:2507.07969v42025
  2. VIBE: Visual Instruction Based Editor

    Grigorii Alekseenko, Aleksandr Gordeev, Irina Tolstykh +7

    cs.CVcs.AIcs.LGarXiv:2601.02242v12026
  3. SampoNLP: A Self-Referential Toolkit for Morphological Analysis of Subword Tokenizers

    Iaroslav Chelombitko, Ekaterina Chelombitko, Aleksey Komissarov

    cs.CLcs.IRcs.LGarXiv:2601.04469v12026
  4. GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving

    Lloyd Russell, Anthony Hu, Lorenzo Bertoni +4

    cs.CVcs.AIcs.ROarXiv:2503.20523v12025
  5. OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction

    Lujie Yang, Xiaoyu Huang, Zhen Wu +6

    cs.ROcs.AIcs.LGarXiv:2509.26633v32025
  6. MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

    Kunlun Zhu, Hongyi Du, Zhaochen Hong +8

    cs.MAcs.AIcs.CLarXiv:2503.01935v12025
  7. Learning Smooth and Expressive Interatomic Potentials for Physical Property Prediction

    Xiang Fu, Brandon M. Wood, Luis Barroso-Luque +4

    physics.comp-phcs.LGarXiv:2502.12147v22025
  8. Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation

    Ruotong Wang, Zihao Zhu, Siwei Lyu +2

    cs.CVcs.AIarXiv:2609.00206v12026
  9. Learning a visuomotor controller for real world robotic grasping using simulated depth images

    Ulrich Viereck, Andreas ten Pas, Kate Saenko +1

    cs.ROcs.AIarXiv:1706.04652v32017
  10. Learning to Act from Actionless Videos through Dense Correspondences

    Po-Chen Ko, Jiayuan Mao, Yilun Du +2

    cs.ROcs.CVcs.LGarXiv:2310.08576v12023
  11. CoverM: Read alignment statistics for metagenomics

    Samuel T. N. Aroney, Rhys J. P. Newell, Jakob N. Nissen +3

    q-bio.GNarXiv:2501.11217v12025
  12. Poseidon: Efficient Foundation Models for PDEs

    Maximilian Herde, Bogdan Raonić, Tobias Rohner +4

    cs.LGarXiv:2405.19101v22024
  13. PaliGemma: A versatile 3B VLM for transfer

    Lucas Beyer, Andreas Steiner, André Susano Pinto +32

    cs.CVcs.AIcs.CLarXiv:2407.07726v22024
  14. DataComp-LM: In search of the next generation of training sets for language models

    Jeffrey Li, Alex Fang, Georgios Smyrnis +56

    cs.LGcs.CLarXiv:2406.11794v42024
  15. Simplified and Generalized Masked Diffusion for Discrete Data

    Jiaxin Shi, Kehang Han, Zhe Wang +2

    cs.LGstat.MLarXiv:2406.04329v42024
  16. YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-time Object Detection

    Yuming Chen, Xinbin Yuan, Jiabao Wang +4

    cs.CVarXiv:2308.05480v22023
  17. Joint Beam Training and Positioning For Intelligent Reflecting Surfaces Assisted Millimeter Wave Communications

    Wei Wang, Wei Zhang

    cs.ITeess.SParXiv:2009.03536v22020
  18. ReFT: Representation Finetuning for Language Models

    Zhengxuan Wu, Aryaman Arora, Zheng Wang +4

    cs.CLcs.AIcs.LGarXiv:2404.03592v32024
  19. Diffusion Model-Based Image Editing: A Survey

    Yi Huang, Jiancheng Huang, Yifan Liu +7

    cs.CVarXiv:2402.17525v42024
  20. Otter: A Multi-Modal Model with In-Context Instruction Tuning

    Bo Li, Yuanhan Zhang, Liangyu Chen +5

    cs.CVcs.CLarXiv:2305.03726v22023
  21. Integrated Sensing and Communication Signals Toward 5G-A and 6G: A Survey

    Zhiqing Wei, Hanyang Qu, Yuan Wang +6

    cs.ITeess.SParXiv:2301.03857v32023
  22. ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

    Xin Men, Mingyu Xu, Qingyu Zhang +5

    cs.CLarXiv:2403.03853v32024
  23. HCF-Net: Hierarchical Context Fusion Network for Infrared Small Object Detection

    Shibiao Xu, ShuChen Zheng, Wenhao Xu +6

    cs.CVarXiv:2403.10778v12024
  24. StylizedNeRF: Consistent 3D Scene Stylization as Stylized NeRF via 2D-3D Mutual Learning

    Yi-Hua Huang, Yue He, Yu-Jie Yuan +2

    cs.GRcs.CVarXiv:2205.12183v22022
  25. V-STaR: Training Verifiers for Self-Taught Reasoners

    Arian Hosseini, Xingdi Yuan, Nikolay Malkin +3

    cs.LGcs.AIcs.CLarXiv:2402.06457v22024
  26. Unexpected Improvements to Expected Improvement for Bayesian Optimization

    Sebastian Ament, Samuel Daulton, David Eriksson +2

    cs.LGmath.NAstat.MLarXiv:2310.20708v32023
    Summaries:한국어
  27. Localization Distillation for Dense Object Detection

    Zhaohui Zheng, Rongguang Ye, Ping Wang +4

    cs.CVarXiv:2102.12252v42021
  28. Advancements in Generative AI: A Comprehensive Review of GANs, GPT, Autoencoders, Diffusion Model, and Transformers

    Staphord Bengesi, Hoda El-Sayed, Md Kamruzzaman Sarker +3

    cs.LGcs.AIarXiv:2311.10242v22023
  29. FastVGGT: Training-Free Acceleration of Visual Geometry Transformer

    You Shen, Zhipeng Zhang, Yansong Qu +4

    cs.CVarXiv:2509.02560v22025
  30. RIS-Aided Cell-Free Massive MIMO Systems for 6G: Fundamentals, System Design, and Applications

    Enyu Shi, Jiayi Zhang, Hongyang Du +5

    cs.ITeess.SParXiv:2310.00263v32023
  31. Convolutional Recurrent Neural Networks for Small-Footprint Keyword Spotting

    Sercan O. Arik, Markus Kliegl, Rewon Child +5

    cs.CLcs.AIcs.LGarXiv:1703.05390v32017
  32. Benchmarking Distributed Stream Data Processing Systems

    Jeyhun Karimov, Tilmann Rabl, Asterios Katsifodimos +3

    cs.DBarXiv:1802.08496v22018
  33. Being-H0.7: A Latent World-Action Model from Egocentric Videos

    Hao Luo, Wanpeng Zhang, Yicheng Feng +6

    cs.ROcs.CVcs.LGarXiv:2605.00078v12026
  34. OmniControl: Control Any Joint at Any Time for Human Motion Generation

    Yiming Xie, Varun Jampani, Lei Zhong +2

    cs.CVcs.GRarXiv:2310.08580v22023
  35. iTransformer: Inverted Transformers Are Effective for Time Series Forecasting

    Yong Liu, Tengge Hu, Haoran Zhang +4

    cs.LGarXiv:2310.06625v42023
  36. Ferret: Refer and Ground Anything Anywhere at Any Granularity

    Haoxuan You, Haotian Zhang, Zhe Gan +6

    cs.CVcs.CLarXiv:2310.07704v12023
  37. Collaborative Regression of Expressive Bodies using Moderation

    Yao Feng, Vasileios Choutas, Timo Bolkart +2

    cs.CVarXiv:2105.05301v22021
  38. Large Language Models as Optimizers

    Chengrun Yang, Xuezhi Wang, Yifeng Lu +4

    cs.LGcs.AIcs.CLarXiv:2309.03409v32023
  39. LIMR: Less is More for RL Scaling

    Xuefeng Li, Haoyang Zou, Pengfei Liu

    cs.LGcs.AIcs.CLarXiv:2502.11886v12025
  40. Jointly Localizing and Describing Events for Dense Video Captioning

    Yehao Li, Ting Yao, Yingwei Pan +2

    cs.CVarXiv:1804.08274v12018
  41. Negative Momentum for Improved Game Dynamics

    Gauthier Gidel, Reyhane Askari Hemmat, Mohammad Pezeshki +4

    cs.LGstat.MLarXiv:1807.04740v52018
  42. Convolutional Bypasses Are Better Vision Transformer Adapters

    Shibo Jie, Zhi-Hong Deng

    cs.CVarXiv:2207.07039v32022
  43. A Comprehensive Survey of Deep Transfer Learning for Anomaly Detection in Industrial Time Series: Methods, Applications, and Directions

    Peng Yan, Ahmed Abdulkadir, Paul-Philipp Luley +4

    cs.LGcs.AIarXiv:2307.05638v22023
  44. Language is All a Graph Needs

    Ruosong Ye, Caiqi Zhang, Runhui Wang +2

    cs.CLcs.AIcs.IRarXiv:2308.07134v52023
  45. Feature Decomposition and Reconstruction Learning for Effective Facial Expression Recognition

    Delian Ruan, Yan Yan, Shenqi Lai +3

    cs.CVarXiv:2104.05160v22021
  46. Is Self-Repair a Silver Bullet for Code Generation?

    Theo X. Olausson, Jeevana Priya Inala, Chenglong Wang +2

    cs.CLcs.AIcs.PLarXiv:2306.09896v52023
  47. Portfolio Risk Bounds without Cross-Asset Return Covariances: Distributional Fields from Language-Model Representations

    Marcus Gawronsky, Chun-Sung Huang

    q-fin.STarXiv:2608.29692v12026
  48. Unleashing the Power of Edge-Cloud Generative AI in Mobile Networks: A Survey of AIGC Services

    Minrui Xu, Hongyang Du, Dusit Niyato +9

    cs.NIarXiv:2303.16129v42023
  49. DragDiffusion: Harnessing Diffusion Models for Interactive Point-based Image Editing

    Yujun Shi, Chuhui Xue, Jun Hao Liew +5

    cs.CVcs.LGarXiv:2306.14435v62023
  50. ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization

    Xixi Wu, Kuan Li, Yida Zhao +13

    cs.CLarXiv:2509.13313v32025
  51. Unified Vision-Language-Action Model

    Yuqi Wang, Xinghang Li, Wenxuan Wang +5

    cs.CVcs.ROarXiv:2506.19850v12025
  52. Can Language Models Solve Graph Problems in Natural Language?

    Heng Wang, Shangbin Feng, Tianxing He +3

    cs.CLcs.AIarXiv:2305.10037v32023
  53. mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

    Qinghao Ye, Haiyang Xu, Guohai Xu +15

    cs.CLcs.CVcs.LGarXiv:2304.14178v32023
  54. Interactive and Explainable Region-guided Radiology Report Generation

    Tim Tanida, Philip Müller, Georgios Kaissis +1

    cs.CVcs.CLcs.LGarXiv:2304.08295v12023
  55. GraphPrompt: Unifying Pre-Training and Downstream Tasks for Graph Neural Networks

    Zemin Liu, Xingtong Yu, Yuan Fang +1

    cs.LGcs.CLarXiv:2302.08043v32023
  56. Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning

    Thomas Carta, Clément Romac, Thomas Wolf +3

    cs.LGarXiv:2302.02662v52023
  57. Experimentally realized in situ backpropagation for deep learning in nanophotonic neural networks

    Sunil Pai, Zhanghao Sun, Tyler W. Hughes +11

    cs.ETcs.LGphysics.opticsarXiv:2205.08501v12022
  58. Navigating to Objects in the Real World

    Theophile Gervet, Soumith Chintala, Dhruv Batra +2

    cs.ROcs.CVcs.LGarXiv:2212.00922v12022
  59. Open-vocabulary Queryable Scene Representations for Real World Planning

    Boyuan Chen, Fei Xia, Brian Ichter +5

    cs.ROcs.AIcs.CVarXiv:2209.09874v22022
  60. Data Augmentation techniques in time series domain: A survey and taxonomy

    Guillermo Iglesias, Edgar Talavera, Ángel González-Prieto +2

    cs.LGcs.AIarXiv:2206.13508v42022