Every paper with a summary
Every arXiv paper Paperlayer has summarized so far. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
54,901 to 54,960 of 61,280
MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments
Han Wang, David Wan, Hyunji Lee +6
cs.CLcs.AIcs.CVarXiv:2604.13418v12026Cross-Tokenizer LLM Distillation through a Byte-Level Interface
Avyav Kumar Singh, Yen-Chen Wu, Alexandru Cioba +2
cs.CLarXiv:2604.07466v22026Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction
Efstathios Karypidis, Spyros Gidaris, Nikos Komodakis
cs.CVarXiv:2604.11707v12026Self-Evolving LLM Memory Extraction Across Heterogeneous Tasks
Yuqing Yang, Tengxiao Liu, Wang Bill Zhu +3
cs.CLarXiv:2604.11610v12026Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
Yinghui He, Simran Kaur, Adithya Bhaskar +7
cs.CLarXiv:2604.12002v22026Domain-Specific Latent Representations Improve the Fidelity of Diffusion-Based Medical Image Super-Resolution
Sebastian Cajas, Ashaba Judith, Rahul Gorijavolu +8
cs.CVcs.AIarXiv:2604.12152v12026Learning Versatile Humanoid Manipulation with Touch Dreaming
Yaru Niu, Zhenlong Fang, Binghong Chen +8
cs.ROarXiv:2604.13015v32026VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
Andrei Atanov, Jesse Allardice, Roman Bachmann +6
cs.CVcs.LGarXiv:2604.12887v12026Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness
Tomer Ashuach, Shai Gretz, Yoav Katz +2
cs.CLarXiv:2604.12373v52026Towards Long-horizon Agentic Multimodal Search
Yifan Du, Zikang Liu, Jinbiao Peng +5
cs.CVcs.AIarXiv:2604.12890v22026Toward Autonomous Long-Horizon Engineering for ML Research
Guoxin Chen, Jie Chen, Lei Chen +7
cs.CLarXiv:2604.13018v22026Exploration and Exploitation Errors Are Measurable for Language Model Agents
Jaden Park, Jungtaek Kim, Jongwon Jeong +3
cs.AIarXiv:2604.13151v12026Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
Jiacheng Liu, Xiaohan Zhao, Xinyi Shang +1
cs.SEcs.AIcs.CLarXiv:2604.14228v22026InfiniteScienceGym: An Unbounded, Procedurally-Generated Benchmark for Scientific Analysis
Oliver Bentham, Vivek Srikumar
cs.CLcs.AIarXiv:2604.13201v22026Boosting Visual Instruction Tuning with Self-Supervised Guidance
Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc +2
cs.CVarXiv:2604.12966v12026Forge-UGC: FX optimization and register-graph engine for universal graph compiler
Satyam Kumar, Saurabh Jha
cs.ARcs.AIcs.DCarXiv:2604.16498v12026Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding
Eun Woo Im, Dhruv Madhwal, Vivek Gupta
cs.LGarXiv:2604.13313v12026UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding
Fei Tang, Bofan Chen, Zhengxi Lu +8
cs.CVcs.AIcs.CLarXiv:2604.14113v12026Memory Transfer Learning: How Memories are Transferred Across Domains in Coding Agents
Kangsan Kim, Minki Kang, Taeil Kim +3
cs.AIcs.CLarXiv:2604.14004v12026TIP: Token Importance in On-Policy Distillation
Yuanda Xu, Hejian Sang, Zhengze Zhou +3
cs.LGcs.AIarXiv:2604.14084v42026From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space
Yuqiao Tan, Minzheng Wang, Bo Liu +5
cs.LGcs.AIcs.CLarXiv:2604.14142v12026Geometric Context Transformer for Streaming 3D Reconstruction
Lin-Zhuo Chen, Jian Gao, Yihang Chen +8
cs.CVarXiv:2604.14141v22026C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences
Akira Kawabata, Saku Sugawara
cs.CLcs.LGarXiv:2604.13618v12026Three-Phase Transformer
Mohammad R. Abu Ayyash
cs.CLcs.AIcs.LGarXiv:2604.14430v12026(1D) Ordered Tokens Enable Efficient Test-Time Search
Zhitong Gao, Parham Rezaei, Ali Cy +7
cs.CVcs.AIcs.LGarXiv:2604.15453v12026HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
Team HY-World, Chenjie Cao, Xuhui Zuo +42
cs.CVarXiv:2604.14268v12026RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography
Mélanie Roschewitz, Kenneth Styppa, Yitian Tao +10
cs.AIarXiv:2604.15231v22026Target-Oriented Pretraining Data Selection via Neuron-Activated Graph
Zijun Wang, Haoqin Tu, Weidong Zhou +7
cs.CLarXiv:2604.15706v12026HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System
Tianshuo Yang, Guanyu Chen, Yutian Chen +8
cs.CVcs.AIcs.ROarXiv:2604.14125v22026AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization
Genghan Zhang, Shaowei Zhu, Anjiang Wei +6
cs.LGcs.CLarXiv:2511.15915v22025Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Xiaohua Wang, Muzhao Tian, Yuqi Zeng +20
cs.LGarXiv:2604.13602v12026WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training
Yifu Chen, Shengpeng Ji, Qian Chen +9
cs.AIarXiv:2604.14932v12026OneHOI: Unifying Human-Object Interaction Generation and Editing
Jiun Tian Hoe, Weipeng Hu, Xudong Jiang +2
cs.CVcs.MMarXiv:2604.14062v12026MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
Yan Li, Zezi Zeng, Yifan Yang +12
cs.CVcs.AIcs.CLarXiv:2604.15309v12026GlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene Tokens
Roni Itkin, Noam Issachar, Yehonatan Keypur +3
cs.CVarXiv:2604.15284v22026TRACER: Trace-Based Adaptive Cost-Efficient Routing for LLM Classification
Adam Rida
cs.AIarXiv:2604.14531v12026OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis
Kanzhi Cheng, Zehao Li, Zheng Ma +11
cs.AIcs.CLcs.CVarXiv:2604.15093v12026Where does output diversity collapse in post-training?
Constantinos Karouzos, Xingwei Tan, Nikolaos Aletras
cs.CLcs.AIcs.LGarXiv:2604.16027v12026Don't Make Models Guess Security and Safety: Symbolic Guardrails for Domain-Specific AI Agents
Yining Hong, Yining She, Eunsuk Kang +2
cs.SEcs.AIcs.CRarXiv:2604.15579v22026Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
Sai Srinivas Kancheti, Aditya Sanjiv Kanade, Vineeth N. Balasubramanian +1
cs.CVcs.AIarXiv:2604.16060v12026Qwen3.5-Omni Technical Report
Qwen Team
cs.CLeess.ASarXiv:2604.15804v22026GFT: From Imitation to Reward Fine-Tuning with Unbiased Group Advantages and Dynamic Coefficient Rectification
Wangjie Gan, Miao Pan, Linbo Xi +4
cs.AIcs.LGarXiv:2604.14258v32026Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes
Victoria Yue Chen, Emery Pierson, Léopold Maillard +1
cs.CVarXiv:2604.14914v22026An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental Learning
Quyen Tran, Hai Nguyen, Hoang Phan +6
cs.LGcs.CVarXiv:2211.16780v42022LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning
Bowen Ping, Zijun Chen, Tingfeng Hui +4
cs.LGcs.CLarXiv:2604.14922v12026Don't Retrieve, Navigate: Distilling Enterprise Knowledge into Navigable Agent Skills for QA and RAG
Yiqun Sun, Pengfei Wei, Lawrence B. Hsieh
cs.IRcs.AIcs.CLarXiv:2604.14572v32026LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories
Zhanhao Liang, Tao Yang, Jie Wu +2
cs.CVarXiv:2604.15311v22026Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models
Haoyi Sun, Xiaoxiao Wang, Ning Mao +5
cs.CVarXiv:2604.14629v12026DR$^{3}$-Eval: Towards Realistic and Reproducible Deep Research Evaluation
Qianqian Xie, Qingheng Xiong, He Zhu +16
cs.AIarXiv:2604.14683v12026EdgeDetect: Importance-Aware Gradient Compression with Homomorphic Aggregation for Federated Intrusion Detection
Noor Islam S. Mohammad
cs.CRarXiv:2604.14663v12026PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research
Tingjia Miao, Wenkai Jin, Muhua Zhang +19
cs.LGcs.AIphysics.data-anarXiv:2604.15411v12026Learning Adaptive Reasoning Paths for Efficient Visual Reasoning
Yixu Huang, Tinghui Zhu, Muhao Chen
cs.CVcs.CLarXiv:2604.14568v12026QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies
Alexey Khoroshilov, Alexey Chernysh, Orkhan Ekhtibarov +2
cs.CLarXiv:2604.15151v12026Maximal Brain Damage Without Data or Optimization: Disrupting Neural Networks via Sign-Bit Flips
Ido Galil, Moshe Kimhi, Ran El-Yaniv
cs.LGcs.AIcs.CVarXiv:2502.07408v22025PerturbRx: Learning Treatment-Conditioned Latent Transitions for Patient Drug Response Prediction
Yoshitaka Inoue, Minoh Jeong, Alfred Hero +2
q-bio.QMcs.LGarXiv:2608.21349v12026Hierarchical Codec Diffusion for Video-to-Speech Generation
Jiaxin Ye, Gaoxiang Cong, Chenhui Wang +4
cs.SDcs.CVarXiv:2604.15923v12026TwinTrack: Post-hoc Multi-Rater Calibration for Medical Image Segmentation
Tristan Kirscher, Alexandra Ertl, Klaus Maier-Hein +3
cs.LGarXiv:2604.15950v22026VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects
Xiangbo Gao, Sicong Jiang, Bangya Liu +12
cs.CVcs.AIcs.CLarXiv:2604.16272v22026ArtifactNet: Detecting AI-Generated Music via Forensic Residual Physics
Heewon Oh
cs.SDeess.ASarXiv:2604.16254v22026GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows
Jize Wang, Xuanxuan Liu, Yining Li +7
cs.CLcs.AIarXiv:2604.15715v12026