Every paper with a summary

Every arXiv paper Paperlayer has summarized so far. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

54,421 to 54,480 of 61,111

  1. SAM 3D Body: Robust Full-Body Human Mesh Recovery

    Xitong Yang, Devansh Kukreja, Don Pinkus +11

    cs.CVarXiv:2602.15989v12026
  2. MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents

    Haozhen Zhang, Quanyu Long, Jianzhu Bao +4

    cs.CLcs.AIcs.LGarXiv:2602.02474v22026
  3. Qwen3-Coder-Next Technical Report

    Ruisheng Cao, Mouxiang Chen, Jiawei Chen +17

    cs.CLarXiv:2603.00729v12026
  4. SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?

    Tingxu Han, Yi Zhang, Wei Song +4

    cs.SEcs.AIarXiv:2603.15401v12026
  5. V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning

    Lorenzo Mur-Labadia, Matthew Muckley, Amir Bar +6

    cs.CVarXiv:2603.14482v32026
  6. Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

    Moo Jin Kim, Yihuai Gao, Tsung-Yi Lin +8

    cs.AIcs.ROarXiv:2601.16163v12026
  7. Advancing Open-source World Models

    Robbyant Team, Zelin Gao, Qiuyu Wang +21

    cs.CVarXiv:2601.20540v12026
  8. DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos

    Shenyuan Gao, William Liang, Kaiyuan Zheng +27

    cs.ROcs.AIcs.CVarXiv:2602.06949v12026
  9. World Action Models are Zero-shot Policies

    Seonghyeon Ye, Yunhao Ge, Kaiyuan Zheng +33

    cs.ROcs.CVcs.LGarXiv:2602.15922v12026
  10. Qwen3-ASR Technical Report

    Xian Shi, Xiong Wang, Zhifang Guo +10

    cs.CLcs.SDeess.ASarXiv:2601.21337v22026
    Summaries:한국어
  11. Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

    Jeonghye Kim, Xufang Luo, Minbeom Kim +5

    cs.CLcs.LGarXiv:2603.24472v42026
  12. GLM-5: from Vibe Coding to Agentic Engineering

    GLM-5-Team, :, Aohan Zeng +184

    cs.LGcs.CLarXiv:2602.15763v22026
  13. Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

    Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82

    cs.SEcs.AIarXiv:2601.11868v12026
  14. Human-Centric Intelligence in the Era of Foundation Models: A Survey

    Yang Chen, Tianqi Wang, Xiaorui Jiang +13

    cs.CVarXiv:2608.18184v12026
    Summaries:한국어
  15. EgoSim: Egocentric World Simulator for Embodied Interaction Generation

    Jinkun Hao, Mingda Jia, Ruiyan Wang +7

    cs.CVcs.AIarXiv:2604.01001v22026
  16. RecGen3D: Reconstruction-Guided 3D Generation in a Shared Canonical Space

    Zhisheng Huang, Jiahao Chen, Cheng Lin +10

    cs.CVarXiv:2604.01479v32026
  17. PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding

    Nan Wang, Zhiwei Jin, Chen Chen +1

    cs.CVcs.AIcs.CLarXiv:2604.00886v12026
  18. Signals: Trajectory Sampling and Triage for Agentic Interactions

    Shuguang Chen, Adil Hafeez, Salman Paracha

    cs.AIcs.CLarXiv:2604.00356v12026
  19. Investigating Autonomous Agent Contributions in the Wild: Activity Patterns and Code Change over Time

    Razvan Mihai Popescu, David Gros, Andrei Botocan +3

    cs.SEcs.AIcs.LGarXiv:2604.00917v12026
  20. Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants

    Deepak Nathani, Cheng Zhang, Chang Huan +7

    cs.AIcs.LGcs.MAarXiv:2604.00842v12026
  21. Benchmarking and Mechanistic Analysis of Vision-Language Models for Cross-Depiction Assembly Instruction Alignment

    Zhuchenyang Liu, Yao Zhang, Yu Xiao

    cs.CVcs.CLarXiv:2604.00913v22026
  22. FlowSlider: Training-Free Continuous Image Editing via Fidelity-Steering Decomposition

    Taichi Endo, Guoqing Hao, Kazuhiko Sumi

    cs.CVarXiv:2604.02088v12026
  23. Do Phone-Use Agents Respect Your Privacy?

    Zhengyang Tang, Ke Ji, Xidong Wang +19

    cs.CRcs.AIcs.CLarXiv:2604.00986v22026
  24. Revision or Re-Solving? Decomposing Second-Pass Gains in Multi-LLM Pipelines

    Jingjie Ning, Xueqi Li, Chengyu Yu

    cs.SEcs.AIcs.CLarXiv:2604.01029v22026
  25. Woosh: A Sound Effects Foundation Model

    Gaëtan Hadjeres, Marc Ferras, Khaled Koutini +7

    cs.SDcs.AIcs.LGarXiv:2604.01929v32026
  26. DynaVid: Learning to Generate Highly Dynamic Videos using Synthetic Motion Data

    Wonjoon Jin, Jiyun Won, Janghyeok Han +4

    cs.CVarXiv:2604.01666v12026
  27. VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification

    Jiahao Meng, Tan Yue, Qi Xu +7

    cs.CVcs.MMarXiv:2604.01569v12026
  28. VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors

    Haz Sameen Shahgir, Xiaofu Chen, Yu Fu +4

    cs.CVcs.CLarXiv:2604.02486v32026
  29. Generative World Renderer

    Zheng-Hui Huang, Zhixiang Wang, Jiaming Tan +6

    cs.CVarXiv:2604.02329v12026
  30. Apriel-1.5-OpenReasoner: RL Post-Training for General-Purpose and Efficient Reasoning

    Rafael Pardinas, Ehsan Kamalloo, David Vazquez +1

    cs.LGarXiv:2604.02007v22026
  31. Executing as You Generate: Hiding Execution Latency in LLM Code Interpreters

    Zhensu Sun, Zhihao Lin, Zhi Chen +4

    cs.PLcs.AIcs.SEarXiv:2604.00491v22026
  32. CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery

    Ao Qu, Han Zheng, Zijian Zhou +14

    cs.AIarXiv:2604.01658v22026
  33. LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation

    Patrick Amadeus Irawan, Erland Hilman Fuadi, Shanu Kumar +2

    cs.CVcs.CLarXiv:2604.00829v32026
  34. Forecasting Supply Chain Disruptions with Foresight Learning

    Benjamin Turtel, Paul Wilczewski, Kris Skotheim

    cs.LGarXiv:2604.01298v12026
  35. Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies

    Zhanzhi Lou, Hui Chen, Yibo Li +2

    cs.LGcs.AIarXiv:2604.00830v32026
  36. Do World Action Models Generalize Better than VLAs? A Robustness Study

    Zhanguang Zhang, Zhiyuan Li, Behnam Rahmati +11

    cs.ROarXiv:2603.22078v52026
  37. AgentSocialBench: Evaluating Privacy Risks in Human-Centered Agentic Social Networks

    Prince Zizhuang Wang, Shuli Jiang

    cs.AIcs.SIarXiv:2604.01487v22026
  38. Test-Time Scaling Makes Overtraining Compute-Optimal

    Nicholas Roberts, Sungjun Cho, Zhiqi Gao +7

    cs.LGcs.CLstat.MLarXiv:2604.01411v12026
  39. UniMixer: A Unified Architecture for Scaling Laws in Recommendation Systems

    Mingming Ha, Guanchen Wang, Linxun Chen +9

    cs.IRcs.AIarXiv:2604.00590v22026
  40. A Survey of On-Policy Distillation for Large Language Models

    Mingyang Song, Mao Zheng

    cs.LGcs.CLarXiv:2604.00626v42026
  41. All Roads Lead to Rome: Incentivizing Divergent Thinking in Vision-Language Models

    Xinyu Tian, Shu Zou, Zhaoyuan Yang +3

    cs.CVarXiv:2604.00479v12026
  42. AutoMIA: Improved Baselines for Membership Inference Attack via Agentic Self-Exploration

    Ruhao Liu, Weiqi Huang, Qi Li +1

    cs.CRcs.CVarXiv:2604.01014v12026
  43. AgentWatcher: A Rule-based Prompt Injection Monitor

    Yanting Wang, Wei Zou, Runpeng Geng +1

    cs.CRarXiv:2604.01194v22026
  44. HippoCamp: Benchmarking Contextual Agents on Personal Computers

    Zhe Yang, Shulin Tian, Kairui Hu +9

    cs.AIcs.CVarXiv:2604.01221v12026
    Summaries:한국어
  45. Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers

    Atsuyuki Miyai, Mashiro Toyooka, Zaiying Zhao +3

    cs.CLcs.AIcs.LGarXiv:2604.01128v12026
  46. Universal YOCO for Efficient Depth Scaling

    Yutao Sun, Li Dong, Tianzhu Ye +3

    cs.CLarXiv:2604.01220v12026
  47. Reasoning Shift: How Context Silently Shortens LLM Reasoning

    Gleb Rodionov, Roman Garipov, George Yakushev

    cs.LGarXiv:2604.01161v22026
  48. Brainstacks: Cross-Domain Cognitive Capabilities via Frozen MoE-LoRA Stacks for Continual LLM Learning

    Mohammad R. Abu Ayyash

    cs.CLcs.AIarXiv:2604.01152v12026
  49. ActionParty: Multi-Subject Action Binding in Generative Video Games

    Alexander Pondaven, Ziyi Wu, Igor Gilitschenski +4

    cs.CVcs.AIcs.LGarXiv:2604.02330v22026
  50. T5Gemma-TTS Technical Report

    Chihiro Arata, Kiyoshi Kurihara

    eess.ASarXiv:2604.01760v12026
  51. Tex3D: Objects as Attack Surfaces via Adversarial 3D Textures for Vision-Language-Action Models

    Jiawei Chen, Simin Huang, Jiawei Du +5

    cs.CVcs.AIarXiv:2604.01618v12026
  52. Omni123: Exploring 3D Native Foundation Models with Limited 3D Data by Unifying Text to 2D and 3D Generation

    Chongjie Ye, Cheng Cao, Chuanyu Pan +4

    cs.CVcs.AIarXiv:2604.02289v12026
  53. UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving

    Yongkang Li, Lijun Zhou, Sixu Yan +11

    cs.CVcs.ROarXiv:2604.02190v12026
  54. Omni-SimpleMem: Autoresearch-Guided Discovery of Lifelong Multimodal Agent Memory

    Jiaqi Liu, Zipeng Ling, Shi Qiu +9

    cs.AIarXiv:2604.01007v22026
  55. GPA: Learning GUI Process Automation from Demonstrations

    Zirui Zhao, Jun Hao Liew, Yan Yang +5

    cs.CVcs.AIcs.SEarXiv:2604.01676v22026
  56. LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model

    Jiachun Jin, Zetong Zhou, Xiao Yang +4

    cs.CVcs.LGarXiv:2604.02097v12026
  57. NearID: Identity Representation Learning via Near-identity Distractors

    Aleksandar Cvejic, Rameen Abdal, Abdelrahman Eldesokey +2

    cs.CVarXiv:2604.01973v22026
  58. Therefore I am. I Think

    Esakkivel Esakkiraja, Sai Rajeswar, Denis Akhiyarov +1

    cs.AIarXiv:2604.01202v32026
  59. Steerable Visual Representations

    Jona Ruthardt, Manu Gaur, Deva Ramanan +2

    cs.CVcs.AIarXiv:2604.02327v22026
  60. The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook

    Xinlei Yu, Zhangquan Chen, Yongbo He +36

    cs.AIarXiv:2604.02029v22026