Every paper with a summary
Every arXiv paper Paperlayer has summarized so far. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
54,841 to 54,900 of 61,428
InCoder-32B-Thinking: Industrial Code World Model for Thinking
Jian Yang, Wei Zhang, Jiajun Wu +22
cs.ARcs.AIcs.CLarXiv:2604.03144v12026Do Audio-Visual Large Language Models Really See and Hear?
Ramaneswaran Selvakumar, Kaousheik Jayakumar, S Sakshi +3
cs.AIcs.SDarXiv:2604.02605v12026Beyond the Assistant Turn: User Turn Generation as a Probe of Interaction Awareness in Language Models
Sarath Shekkizhar, Romain Cosentino, Adam Earle
cs.AIarXiv:2604.02315v22026LightThinker++: From Reasoning Compression to Memory Management
Yuqi Zhu, Jintian Zhang, Zhenjie Wan +7
cs.CLcs.AIcs.IRarXiv:2604.03679v12026POEMetric: The Last Stanza of Humanity
Bingru Li, Han Wang, Hazel Wilkinson
cs.CLarXiv:2604.03695v12026Training a Student Expert via Semi-Supervised Foundation Model Distillation
Pardis Taghavi, Tian Liu, Renjie Li +2
cs.CVarXiv:2604.03841v12026Can Natural Image Autoencoders Compactly Tokenize fMRI Volumes for Long-Range Dynamics Modeling?
Peter Yongho Kim, Juhyeon Park, Jungwoo Park +4
cs.CVarXiv:2604.03619v12026Squeez: Task-Conditioned Tool-Output Pruning for Coding Agents
Ádám Kovács
cs.SEcs.AIarXiv:2604.04979v12026Can LLMs Learn to Reason Robustly under Noisy Supervision?
Shenzhi Yang, Guangcheng Zhu, Bowen Song +7
cs.LGcs.AIarXiv:2604.03993v12026AURA: Always-On Understanding and Real-Time Assistance via Video Streams
Xudong Lu, Yang Bo, Jinpeng Chen +9
cs.CVarXiv:2604.04184v12026DARE: Diffusion Large Language Models Alignment and Reinforcement Executor
Jingyi Yang, Yuxian Jiang, Xuhao Hu +3
cs.CLarXiv:2604.04215v12026A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning
Tianle Chen, Deepti Ghadiyaram
cs.CVcs.SDarXiv:2604.03995v12026HDP: A Lightweight Cryptographic Protocol for Human Delegation Provenance in Agentic AI Systems
Asiri Dalugoda
cs.CRcs.MAarXiv:2604.04522v12026CLEAR: Unlocking Generative Potential for Degraded Image Understanding in Unified Multimodal Models
Xiangzhao Hao, Zefeng Zhang, Zhenyu Zhang +6
cs.CVarXiv:2604.04780v12026AvatarPointillist: AutoRegressive 4D Gaussian Avatarization
Hongyu Liu, Xuan Wang, Zijian Wu +7
cs.CVarXiv:2604.04787v22026Vero: An Open RL Recipe for General Visual Reasoning
Gabriel Sarch, Linrong Cai, Qunzhong Wang +3
cs.CVcs.AIcs.CLarXiv:2604.04917v32026Paper Espresso: From Paper Overload to Research Insight
Mingzhe Du, Luu Anh Tuan, Dong Huang +1
cs.DLcs.AIarXiv:2604.04562v12026Less Detail, Better Answers: Degradation-Driven Prompting for VQA
Haoxuan Han, Weijie Wang, Zeyu Zhang +2
cs.CVarXiv:2604.04838v22026Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw
Zijun Wang, Haoqin Tu, Letian Zhang +11
cs.CRcs.AIcs.CLarXiv:2604.04759v12026FileGram: Grounding Agent Personalization in File-System Behavioral Traces
Shuai Liu, Shulin Tian, Kairui Hu +6
cs.CVcs.AIarXiv:2604.04901v12026OpenWorldLib: A Unified Codebase and Definition of Advanced World Models
DataFlow Team, Bohan Zeng, Daili Hua +39
cs.CVarXiv:2604.04707v22026REAM: Merging Improves Pruning of Experts in LLMs
Saurav Jha, Maryam Hashemzadeh, Ali Saheb Pasand +3
cs.AIcs.CLcs.LGarXiv:2604.04356v12026MedGemma 1.5 Technical Report
Andrew Sellergren, Chufan Gao, Fereshteh Mahvar +39
cs.AIarXiv:2604.05081v22026MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale
Bin Wang, Tianyao He, Linke Ouyang +40
cs.CVcs.CLarXiv:2604.04771v22026ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
Xiangyi Li, Kyoung Whan Choe, Yimin Liu +12
cs.AIarXiv:2604.05172v22026MedConclusion: A Benchmark for Biomedical Conclusion Generation from Structured Abstracts
Weiyue Li, Ruizhi Qian, Yi Li +5
cs.CLcs.AIarXiv:2604.06505v12026GenLCA: 3D Diffusion for Full-Body Avatars from In-the-Wild Videos
Yiqian Wu, Rawal Khirodkar, Egor Zakharov +6
cs.CVarXiv:2604.07273v22026Neural Computers
Mingchen Zhuge, Changsheng Zhao, Haozhe Liu +16
cs.LGcs.AIarXiv:2604.06425v22026How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
Yujian Liu, Jiabao Ji, Li An +3
cs.CLarXiv:2604.04323v12026Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision
Hyunsoo Cha, Wonjung Woo, Byungjun Kim +1
cs.CVarXiv:2604.04934v22026R3PM-Net: Real-time, Robust, Real-world Point Matching Network
Yasaman Kashefbahrami, Erkut Akdag, Panagiotis Meletis +3
cs.CVcs.LGarXiv:2604.05060v22026FORGE: Fine-grained Multimodal Evaluation for Manufacturing Scenarios
Xiangru Jian, Hao Xu, Wei Pang +13
cs.CVcs.AIcs.LGarXiv:2604.07413v22026SuperLocalMemory V3.3: The Living Brain -- Biologically-Inspired Forgetting, Cognitive Quantization, and Multi-Channel Retrieval for Zero-LLM Agent Memory Systems
Varun Pratap Bhardwaj
cs.AIcs.CLcs.IRarXiv:2604.04514v12026CUE-R: Beyond the Final Answer in Retrieval-Augmented Generation
Siddharth Jain, Venkat Narayan Vedam
cs.IRcs.CLcs.LGarXiv:2604.05467v12026DeonticBench: A Benchmark for Reasoning over Rules
Guangyao Dou, Luis Brena, Akhil Deo +4
cs.CLarXiv:2604.04443v12026Can Large Language Models Reinvent Foundational Algorithms?
Jian Zhao, Haoren Luo, Yu Wang +3
cs.AIarXiv:2604.05716v12026MMEmb-R1: Reasoning-Enhanced Multimodal Embedding with Pair-Aware Selection and Adaptive Control
Yuchi Wang, Haiyang Yu, Weikang Bian +4
cs.CVcs.AIcs.CLarXiv:2604.06156v12026FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification
Ling Yue, Chaoqian Ouyang, Hang Xu +7
cs.AIcs.LGarXiv:2604.04074v42026Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
Qihan Ren, Peng Wang, Ruikun Cai +8
cs.AIarXiv:2604.06628v22026Qualixar OS: A Universal Operating System for AI Agent Orchestration
Varun Pratap Bhardwaj
cs.AIcs.MAcs.SEarXiv:2604.06392v12026UniSpace: Unified Visual Representation and Scalable Multimodal Modeling
Jinbo Yan, Limeng Qiao, Jie Qin +3
cs.CVcs.AIarXiv:2608.08676v12026HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents
Tencent Robotics X, HY Vision Team, : +20
cs.CVarXiv:2604.07430v12026OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence
Jianhui Liu, Haoze Sun, Wenbo Li +11
cs.CLarXiv:2604.07296v22026Personalizing Text-to-Image Generation to Individual Taste
Anne-Sofie Maerten, Juliane Verwiebe, Shyamgopal Karthik +3
cs.CVarXiv:2604.07427v12026MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU
Zhengqing Yuan, Hanchi Sun, Lichao Sun +1
cs.CLcs.DCcs.OSarXiv:2604.05091v12026Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
Chaoyou Fu, Haozhi Yuan, Yuhao Dong +16
cs.CVarXiv:2604.05015v12026A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
Tommie Kerssies, Gabriele Berton, Ju He +5
cs.CVarXiv:2604.04913v12026Beyond Hard Negatives: The Importance of Score Distribution in Knowledge Distillation for Dense Retrieval
Youngjoon Jang, Seongtae Hong, Hyeonseok Moon +1
cs.IRarXiv:2604.04734v22026STEER: Structured Event Evidence for Video Reasoning via Multi-Objective Reinforcement Learning
Zinuo Li, Yongxin Guo, Jun Liu +7
cs.CLarXiv:2604.04415v32026SkVM: Revisiting Language VM for Skills across Heterogenous LLMs and Harnesses
Le Chen, Erhu Feng, Yubin Xia +1
cs.SEcs.LGarXiv:2604.03088v32026Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning
Juekai Lin, Yun Zhu, Honglin Lin +6
cs.CVcs.AIarXiv:2604.06079v12026RAGEN-2: Reasoning Collapse in Agentic RL
Zihan Wang, Chi Gui, Xing Jin +13
cs.LGarXiv:2604.06268v12026Target Policy Optimization
Jean Kaddour
cs.LGarXiv:2604.06159v12026TRACE: Capability-Targeted Agentic Training
Hangoo Kang, Tarun Suresh, Jon Saad-Falcon +1
cs.AIarXiv:2604.05336v22026The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment
Rishab Balasubramanian, Pin-Jie Lin, Rituraj Sharma +6
cs.LGcs.AIarXiv:2604.06377v32026Action Images: End-to-End Policy Learning via Multiview Video Generation
Haoyu Zhen, Zixian Gao, Qiao Sun +7
cs.CVcs.ROarXiv:2604.06168v22026Experience Transfer for Multimodal LLM Agents in Minecraft Game
Chenghao Li, Jun Liu, Songbo Zhang +7
cs.AIarXiv:2604.05533v12026The Depth Ceiling: On the Limits of Large Language Models in Discovering Latent Planning
Yi Xu, Philipp Jettkant, Laura Ruis
cs.LGcs.AIcs.CLarXiv:2604.06427v12026In-Place Test-Time Training
Guhao Feng, Shengjie Luo, Kai Hua +4
cs.LGcs.AIcs.CLarXiv:2604.06169v12026Beyond Accuracy: Unveiling Inefficiency Patterns in Tool-Integrated Reasoning
Qisheng Su, Shiting Huang, Zhen Fang +3
cs.PFcs.SEarXiv:2604.05404v22026