Every paper with a summary

Every arXiv paper Paperlayer has summarized so far. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

54,961 to 55,020 of 61,393

  1. SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks

    Tianyi Wang, Yixia Li, Long Li +6

    cs.AIarXiv:2604.08865v12026
  2. Zero-shot World Models Are Developmentally Efficient Learners

    Khai Loong Aw, Klemen Kotar, Wanhee Lee +6

    cs.AIcs.CVarXiv:2604.10333v12026
  3. Counting to Four is still a Chore for VLMs

    Duy Le Dinh Anh, Patrick Amadeus Irawan, Tuan Van Vo

    cs.CVarXiv:2604.10039v12026
  4. Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation

    Gordon Chen, Ziqi Huang, Ziwei Liu

    cs.CVarXiv:2604.10030v12026
  5. EditCrafter: Tuning-free High-Resolution Image Editing via Pretrained Diffusion Model

    Kunho Kim, Sumin Seo, Yongjun Cho +1

    cs.CVarXiv:2604.10268v12026
  6. SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?

    Udari Madhushani Sehwag, Elaine Lau, Haniyeh Ehsani Oskouie +14

    cs.AIarXiv:2604.10718v12026
  7. BMdataset: A Musicologically Curated LilyPond Dataset

    Matteo Spanio, Ilay Guler, Antonio Rodà

    cs.SDcs.CLcs.IRarXiv:2604.10628v22026
  8. DiningBench: A Hierarchical Multi-view Benchmark for Perception and Reasoning in the Dietary Domain

    Song Jin, Juntian Zhang, Xun Zhang +6

    cs.CVarXiv:2604.10425v12026
  9. IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs

    Yuzhen Mao, Qitong Wang, Martin Ester +1

    cs.LGcs.AIarXiv:2604.10539v12026
  10. Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs

    Yu Li, Xiaoran Shang, Qizhi Pei +11

    cs.AIarXiv:2604.10480v12026
  11. PersonalAI: A Systematic Comparison of Knowledge Graph Storage and Retrieval Approaches for Personalized LLM agents

    Mikhail Menschikov, Dmitry Evseev, Victoria Dochkina +5

    cs.CLcs.IRarXiv:2506.17001v62025
  12. The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents

    Xuwei Ding, Skylar Zhai, Linxin Song +6

    cs.CRcs.AIarXiv:2604.10577v22026
  13. Rethinking the Diffusion Model from a Langevin Perspective

    Candi Zheng, Yuan Lan

    cs.LGcs.AIcs.CVarXiv:2604.10465v12026
  14. PokeRL: Reinforcement Learning for Pokemon Red

    Dheeraj Mudireddy, Sai Patibandla

    cs.LGarXiv:2604.10812v12026
  15. Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration

    Zhipeng Chen, Tao Qian, Wayne Xin Zhao +1

    cs.LGcs.AIcs.CLarXiv:2604.11446v12026
  16. Time is Not a Label: Continuous Phase Rotation for Temporal Knowledge Graphs and Agentic Memory

    Weixian Waylon Li, Jiaxin Zhang, Xianan Jim Yang +2

    cs.CLcs.AIarXiv:2604.11544v12026
  17. SWE-AGILE: A Software Agent Framework for Efficiently Managing Dynamic Reasoning Context

    Shuquan Lian, Juncheng Liu, Yazhe Chen +2

    cs.AIcs.CLarXiv:2604.11716v12026
  18. Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization

    Zhixin Lin, Jungang Li, Dongliang Xu +5

    cs.AIcs.CRarXiv:2604.11259v12026
  19. From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models

    Chenchen Zhang

    cs.CLarXiv:2604.09459v32026
  20. General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks

    Junlin Liu, Shengnan An, Shuang Zhou +10

    cs.CLcs.AIarXiv:2604.11778v12026
  21. Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks

    Yoonsang Lee, Howard Yen, Xi Ye +1

    cs.CLarXiv:2604.11753v32026
  22. Introspective Diffusion Language Models

    Yifan Yu, Yuqing Jian, Junxiong Wang +12

    cs.AIarXiv:2604.11035v12026
  23. CocoaBench: Evaluating Unified Digital Agents in the Wild

    CocoaBench Team, Shibo Hao, Zhining Zhang +29

    cs.CLcs.AIarXiv:2604.11201v22026
  24. Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models

    Songlin Yang, Xianghao Kong, Anyi Rao

    cs.CVcs.AIarXiv:2604.10949v12026
  25. OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation

    Donghao Zhou, Guisheng Liu, Hao Yang +9

    cs.CVarXiv:2604.11804v22026
  26. CodeTracer: Towards Traceable Agent States

    Han Li, Yifan Yao, Letian Zhu +13

    cs.SEcs.AIarXiv:2604.11641v32026
  27. The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping

    Yang Liu, Enxi Wang, Yufei Gao +6

    cs.LGcs.AIcs.CLarXiv:2604.11297v12026
  28. Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding

    Shivam Sharma, Sankalp Nagaonkar, Ashish Choithani +1

    cs.CVarXiv:2604.11177v12026
  29. You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass

    Yinuo Yang, Zixian Ma, Manasi Ganti +2

    cs.CVcs.AIarXiv:2604.10966v22026
  30. LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety

    Junxiao Yang, Haoran Liu, Jinzhe Tu +9

    cs.LGcs.AIcs.CLarXiv:2604.12710v22026
  31. ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents

    Fei Tang, Zhiqiong Lu, Boxuan Zhang +4

    cs.LGcs.AIcs.CLarXiv:2604.11784v12026
  32. Grid2Matrix: Revealing Digital Agnosia in Vision-Language Models

    Yunkai Zhang, Linda Li, Yingxin Cui +5

    cs.CVcs.AIarXiv:2604.09687v22026
  33. Motif-Video 2B: Technical Report

    Junghwan Lim, Wai Ting Cheung, Minsu Ha +25

    cs.CVcs.AIarXiv:2604.16503v22026
  34. OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video

    Junfu Pu, Yuxin Chen, Teng Wang +1

    cs.CVcs.MMarXiv:2604.11102v12026
  35. Sema Code: Decoupling AI Coding Agents into Programmable, Embeddable Infrastructure

    Huacan Wang, Jie Zhou, Ningyan Zhu +8

    cs.SEarXiv:2604.11045v12026
  36. Self-Adversarial One Step Generation via Condition Shifting

    Deyuan Liu, Peng Sun, Yansen Han +3

    cs.CVarXiv:2604.12322v12026
  37. GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts

    Amir Hossein Kargaran, Nafiseh Nikeghbal, Jana Diesner +2

    cs.CLcs.CVarXiv:2604.12978v12026
  38. Habitat-GS: A High-Fidelity Navigation Simulator with Dynamic Gaussian Splatting

    Ziyuan Xia, Jingyi Xu, Chong Cui +9

    cs.ROcs.CVarXiv:2604.12626v12026
  39. Accelerating Speculative Decoding with Block Diffusion Draft Trees

    Liran Ringel, Yaniv Romano

    cs.CLarXiv:2604.12989v12026
  40. Generative Refinement Networks for Visual Synthesis

    Jian Han, Jinlai Liu, Jiahuan Wang +2

    cs.CVarXiv:2604.13030v22026
  41. Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective

    Weijie Wang, Qihang Cao, Sensen Gao +10

    cs.CVcs.AIcs.GRarXiv:2604.14025v12026
  42. Mobile GUI Agents under Real-world Threats: Are We There Yet?

    Guohong Liu, Jialei Ye, Jiacheng Liu +5

    cs.CRcs.AIarXiv:2507.04227v22025
  43. Lyra 2.0: Explorable Generative 3D Worlds

    Tianchang Shen, Sherwin Bahmani, Kai He +12

    cs.CVarXiv:2604.13036v12026
  44. ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack

    Yein Park, Jungwoo Park, Jaewoo Kang

    cs.AIarXiv:2509.25843v22025
  45. Towards Autonomous Mechanistic Reasoning in Virtual Cells

    Yunhui Jang, Lu Zhu, Jake Fawkes +3

    cs.LGcs.AIarXiv:2604.11661v32026
  46. UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy Optimization

    Zhengxi Lu, Fei Tang, Guangyi Liu +8

    cs.LGarXiv:2604.13822v12026
  47. RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies

    Jenai Xuning Yang, Rishit Dagli, Alex Zook +5

    cs.ROcs.AIarXiv:2604.09860v42026
  48. SpatialEvo: Self-Evolving Spatial Intelligence via Deterministic Geometric Environments

    Dinging Li, Yingxiu Zhao, Xinrui Cheng +16

    cs.CVcs.CLarXiv:2604.14144v12026
  49. AgentSPEX: An Agent SPecification and EXecution Language

    Pengcheng Wang, Jerry Huang, Jiarui Yao +7

    cs.CLarXiv:2604.13346v12026
  50. Free Geometry: Refining 3D Reconstruction from Longer Versions of Itself

    Yuhang Dai, Xingyi Yang

    cs.CVarXiv:2604.14048v12026
  51. ROSE: Retrieval-Oriented Segmentation Enhancement

    Song Tang, Guangquan Jie, Henghui Ding +1

    cs.CVarXiv:2604.14147v12026
  52. TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration

    Zerun Ma, Guoqiang Wang, Xinchen Xie +7

    cs.AIcs.CLarXiv:2604.14116v22026
  53. OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation

    Xiaomeng Hu, Yinger Zhang, Fei Huang +7

    cs.CLarXiv:2604.10866v22026
  54. MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments

    Han Wang, David Wan, Hyunji Lee +6

    cs.CLcs.AIcs.CVarXiv:2604.13418v12026
  55. Cross-Tokenizer LLM Distillation through a Byte-Level Interface

    Avyav Kumar Singh, Yen-Chen Wu, Alexandru Cioba +2

    cs.CLarXiv:2604.07466v22026
  56. Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction

    Efstathios Karypidis, Spyros Gidaris, Nikos Komodakis

    cs.CVarXiv:2604.11707v12026
  57. Self-Evolving LLM Memory Extraction Across Heterogeneous Tasks

    Yuqing Yang, Tengxiao Liu, Wang Bill Zhu +3

    cs.CLarXiv:2604.11610v12026
  58. Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

    Yinghui He, Simran Kaur, Adithya Bhaskar +7

    cs.CLarXiv:2604.12002v22026
  59. Domain-Specific Latent Representations Improve the Fidelity of Diffusion-Based Medical Image Super-Resolution

    Sebastian Cajas, Ashaba Judith, Rahul Gorijavolu +8

    cs.CVcs.AIarXiv:2604.12152v12026
  60. Learning Versatile Humanoid Manipulation with Touch Dreaming

    Yaru Niu, Zhenlong Fang, Binghong Chen +8

    cs.ROarXiv:2604.13015v32026