Every paper with a summary

Every arXiv paper Paperlayer has summarized so far. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

54,901 to 54,960 of 61,280

  1. MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments

    Han Wang, David Wan, Hyunji Lee +6

    cs.CLcs.AIcs.CVarXiv:2604.13418v12026
  2. Cross-Tokenizer LLM Distillation through a Byte-Level Interface

    Avyav Kumar Singh, Yen-Chen Wu, Alexandru Cioba +2

    cs.CLarXiv:2604.07466v22026
  3. Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction

    Efstathios Karypidis, Spyros Gidaris, Nikos Komodakis

    cs.CVarXiv:2604.11707v12026
  4. Self-Evolving LLM Memory Extraction Across Heterogeneous Tasks

    Yuqing Yang, Tengxiao Liu, Wang Bill Zhu +3

    cs.CLarXiv:2604.11610v12026
  5. Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

    Yinghui He, Simran Kaur, Adithya Bhaskar +7

    cs.CLarXiv:2604.12002v22026
  6. Domain-Specific Latent Representations Improve the Fidelity of Diffusion-Based Medical Image Super-Resolution

    Sebastian Cajas, Ashaba Judith, Rahul Gorijavolu +8

    cs.CVcs.AIarXiv:2604.12152v12026
  7. Learning Versatile Humanoid Manipulation with Touch Dreaming

    Yaru Niu, Zhenlong Fang, Binghong Chen +8

    cs.ROarXiv:2604.13015v32026
  8. VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization

    Andrei Atanov, Jesse Allardice, Roman Bachmann +6

    cs.CVcs.LGarXiv:2604.12887v12026
  9. Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness

    Tomer Ashuach, Shai Gretz, Yoav Katz +2

    cs.CLarXiv:2604.12373v52026
  10. Towards Long-horizon Agentic Multimodal Search

    Yifan Du, Zikang Liu, Jinbiao Peng +5

    cs.CVcs.AIarXiv:2604.12890v22026
  11. Toward Autonomous Long-Horizon Engineering for ML Research

    Guoxin Chen, Jie Chen, Lei Chen +7

    cs.CLarXiv:2604.13018v22026
  12. Exploration and Exploitation Errors Are Measurable for Language Model Agents

    Jaden Park, Jungtaek Kim, Jongwon Jeong +3

    cs.AIarXiv:2604.13151v12026
  13. Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems

    Jiacheng Liu, Xiaohan Zhao, Xinyi Shang +1

    cs.SEcs.AIcs.CLarXiv:2604.14228v22026
  14. InfiniteScienceGym: An Unbounded, Procedurally-Generated Benchmark for Scientific Analysis

    Oliver Bentham, Vivek Srikumar

    cs.CLcs.AIarXiv:2604.13201v22026
  15. Boosting Visual Instruction Tuning with Self-Supervised Guidance

    Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc +2

    cs.CVarXiv:2604.12966v12026
  16. Forge-UGC: FX optimization and register-graph engine for universal graph compiler

    Satyam Kumar, Saurabh Jha

    cs.ARcs.AIcs.DCarXiv:2604.16498v12026
  17. Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding

    Eun Woo Im, Dhruv Madhwal, Vivek Gupta

    cs.LGarXiv:2604.13313v12026
  18. UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding

    Fei Tang, Bofan Chen, Zhengxi Lu +8

    cs.CVcs.AIcs.CLarXiv:2604.14113v12026
  19. Memory Transfer Learning: How Memories are Transferred Across Domains in Coding Agents

    Kangsan Kim, Minki Kang, Taeil Kim +3

    cs.AIcs.CLarXiv:2604.14004v12026
  20. TIP: Token Importance in On-Policy Distillation

    Yuanda Xu, Hejian Sang, Zhengze Zhou +3

    cs.LGcs.AIarXiv:2604.14084v42026
  21. From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space

    Yuqiao Tan, Minzheng Wang, Bo Liu +5

    cs.LGcs.AIcs.CLarXiv:2604.14142v12026
  22. Geometric Context Transformer for Streaming 3D Reconstruction

    Lin-Zhuo Chen, Jian Gao, Yihang Chen +8

    cs.CVarXiv:2604.14141v22026
  23. C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences

    Akira Kawabata, Saku Sugawara

    cs.CLcs.LGarXiv:2604.13618v12026
  24. Three-Phase Transformer

    Mohammad R. Abu Ayyash

    cs.CLcs.AIcs.LGarXiv:2604.14430v12026
  25. (1D) Ordered Tokens Enable Efficient Test-Time Search

    Zhitong Gao, Parham Rezaei, Ali Cy +7

    cs.CVcs.AIcs.LGarXiv:2604.15453v12026
  26. HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds

    Team HY-World, Chenjie Cao, Xuhui Zuo +42

    cs.CVarXiv:2604.14268v12026
  27. RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography

    Mélanie Roschewitz, Kenneth Styppa, Yitian Tao +10

    cs.AIarXiv:2604.15231v22026
  28. Target-Oriented Pretraining Data Selection via Neuron-Activated Graph

    Zijun Wang, Haoqin Tu, Weidong Zhou +7

    cs.CLarXiv:2604.15706v12026
  29. HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System

    Tianshuo Yang, Guanyu Chen, Yutian Chen +8

    cs.CVcs.AIcs.ROarXiv:2604.14125v22026
  30. AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization

    Genghan Zhang, Shaowei Zhu, Anjiang Wei +6

    cs.LGcs.CLarXiv:2511.15915v22025
  31. Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

    Xiaohua Wang, Muzhao Tian, Yuqi Zeng +20

    cs.LGarXiv:2604.13602v12026
  32. WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training

    Yifu Chen, Shengpeng Ji, Qian Chen +9

    cs.AIarXiv:2604.14932v12026
  33. OneHOI: Unifying Human-Object Interaction Generation and Editing

    Jiun Tian Hoe, Weipeng Hu, Xudong Jiang +2

    cs.CVcs.MMarXiv:2604.14062v12026
  34. MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation

    Yan Li, Zezi Zeng, Yifan Yang +12

    cs.CVcs.AIcs.CLarXiv:2604.15309v12026
  35. GlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene Tokens

    Roni Itkin, Noam Issachar, Yehonatan Keypur +3

    cs.CVarXiv:2604.15284v22026
  36. TRACER: Trace-Based Adaptive Cost-Efficient Routing for LLM Classification

    Adam Rida

    cs.AIarXiv:2604.14531v12026
  37. OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis

    Kanzhi Cheng, Zehao Li, Zheng Ma +11

    cs.AIcs.CLcs.CVarXiv:2604.15093v12026
  38. Where does output diversity collapse in post-training?

    Constantinos Karouzos, Xingwei Tan, Nikolaos Aletras

    cs.CLcs.AIcs.LGarXiv:2604.16027v12026
  39. Don't Make Models Guess Security and Safety: Symbolic Guardrails for Domain-Specific AI Agents

    Yining Hong, Yining She, Eunsuk Kang +2

    cs.SEcs.AIcs.CRarXiv:2604.15579v22026
  40. Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs

    Sai Srinivas Kancheti, Aditya Sanjiv Kanade, Vineeth N. Balasubramanian +1

    cs.CVcs.AIarXiv:2604.16060v12026
  41. Qwen3.5-Omni Technical Report

    Qwen Team

    cs.CLeess.ASarXiv:2604.15804v22026
  42. GFT: From Imitation to Reward Fine-Tuning with Unbiased Group Advantages and Dynamic Coefficient Rectification

    Wangjie Gan, Miao Pan, Linbo Xi +4

    cs.AIcs.LGarXiv:2604.14258v32026
  43. Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes

    Victoria Yue Chen, Emery Pierson, Léopold Maillard +1

    cs.CVarXiv:2604.14914v22026
  44. An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental Learning

    Quyen Tran, Hai Nguyen, Hoang Phan +6

    cs.LGcs.CVarXiv:2211.16780v42022
  45. LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning

    Bowen Ping, Zijun Chen, Tingfeng Hui +4

    cs.LGcs.CLarXiv:2604.14922v12026
  46. Don't Retrieve, Navigate: Distilling Enterprise Knowledge into Navigable Agent Skills for QA and RAG

    Yiqun Sun, Pengfei Wei, Lawrence B. Hsieh

    cs.IRcs.AIcs.CLarXiv:2604.14572v32026
  47. LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories

    Zhanhao Liang, Tao Yang, Jie Wu +2

    cs.CVarXiv:2604.15311v22026
  48. Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models

    Haoyi Sun, Xiaoxiao Wang, Ning Mao +5

    cs.CVarXiv:2604.14629v12026
  49. DR$^{3}$-Eval: Towards Realistic and Reproducible Deep Research Evaluation

    Qianqian Xie, Qingheng Xiong, He Zhu +16

    cs.AIarXiv:2604.14683v12026
  50. EdgeDetect: Importance-Aware Gradient Compression with Homomorphic Aggregation for Federated Intrusion Detection

    Noor Islam S. Mohammad

    cs.CRarXiv:2604.14663v12026
  51. PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research

    Tingjia Miao, Wenkai Jin, Muhua Zhang +19

    cs.LGcs.AIphysics.data-anarXiv:2604.15411v12026
  52. Learning Adaptive Reasoning Paths for Efficient Visual Reasoning

    Yixu Huang, Tinghui Zhu, Muhao Chen

    cs.CVcs.CLarXiv:2604.14568v12026
  53. QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies

    Alexey Khoroshilov, Alexey Chernysh, Orkhan Ekhtibarov +2

    cs.CLarXiv:2604.15151v12026
  54. Maximal Brain Damage Without Data or Optimization: Disrupting Neural Networks via Sign-Bit Flips

    Ido Galil, Moshe Kimhi, Ran El-Yaniv

    cs.LGcs.AIcs.CVarXiv:2502.07408v22025
  55. PerturbRx: Learning Treatment-Conditioned Latent Transitions for Patient Drug Response Prediction

    Yoshitaka Inoue, Minoh Jeong, Alfred Hero +2

    q-bio.QMcs.LGarXiv:2608.21349v12026
  56. Hierarchical Codec Diffusion for Video-to-Speech Generation

    Jiaxin Ye, Gaoxiang Cong, Chenhui Wang +4

    cs.SDcs.CVarXiv:2604.15923v12026
  57. TwinTrack: Post-hoc Multi-Rater Calibration for Medical Image Segmentation

    Tristan Kirscher, Alexandra Ertl, Klaus Maier-Hein +3

    cs.LGarXiv:2604.15950v22026
  58. VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects

    Xiangbo Gao, Sicong Jiang, Bangya Liu +12

    cs.CVcs.AIcs.CLarXiv:2604.16272v22026
  59. ArtifactNet: Detecting AI-Generated Music via Forensic Residual Physics

    Heewon Oh

    cs.SDeess.ASarXiv:2604.16254v22026
  60. GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows

    Jize Wang, Xuanxuan Liu, Yining Li +7

    cs.CLcs.AIarXiv:2604.15715v12026