Every paper with a summary
Every arXiv paper Paperlayer has summarized so far. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
54,721 to 54,780 of 61,351
T5Gemma-TTS Technical Report
Chihiro Arata, Kiyoshi Kurihara
eess.ASarXiv:2604.01760v12026Tex3D: Objects as Attack Surfaces via Adversarial 3D Textures for Vision-Language-Action Models
Jiawei Chen, Simin Huang, Jiawei Du +5
cs.CVcs.AIarXiv:2604.01618v12026Omni123: Exploring 3D Native Foundation Models with Limited 3D Data by Unifying Text to 2D and 3D Generation
Chongjie Ye, Cheng Cao, Chuanyu Pan +4
cs.CVcs.AIarXiv:2604.02289v12026UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
Yongkang Li, Lijun Zhou, Sixu Yan +11
cs.CVcs.ROarXiv:2604.02190v12026Omni-SimpleMem: Autoresearch-Guided Discovery of Lifelong Multimodal Agent Memory
Jiaqi Liu, Zipeng Ling, Shi Qiu +9
cs.AIarXiv:2604.01007v22026GPA: Learning GUI Process Automation from Demonstrations
Zirui Zhao, Jun Hao Liew, Yan Yang +5
cs.CVcs.AIcs.SEarXiv:2604.01676v22026LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model
Jiachun Jin, Zetong Zhou, Xiao Yang +4
cs.CVcs.LGarXiv:2604.02097v12026Therefore I am. I Think
Esakkivel Esakkiraja, Sai Rajeswar, Denis Akhiyarov +1
cs.AIarXiv:2604.01202v32026Steerable Visual Representations
Jona Ruthardt, Manu Gaur, Deva Ramanan +2
cs.CVcs.AIarXiv:2604.02327v22026The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook
Xinlei Yu, Zhangquan Chen, Yongbo He +36
cs.AIarXiv:2604.02029v22026DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning
Yang Zhou, Xiaofeng Wang, Hao Shao +8
cs.CVcs.AIcs.ROarXiv:2604.01765v12026SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization
Zhengxi Lu, Zhiyuan Yao, Jinyang Wu +7
cs.LGarXiv:2604.02268v22026Swift-SVD: Theoretical Optimality Meets Practical Efficiency in Low-Rank LLM Compression
Ruoling Qi, Yirui Liu, Xuaner Wu +6
cs.CLarXiv:2604.01609v22026A Simple Baseline for Streaming Video Understanding
Yujiao Shen, Shulin Tian, Jingkang Yang +1
cs.CVarXiv:2604.02317v12026Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
Gengsheng Li, Tianyu Yang, Junfeng Fang +6
cs.LGcs.AIarXiv:2604.02288v12026ThinkTwice: Jointly Optimizing Large Language Models for Reasoning and Self-Refinement
Difan Jiao, Qianfeng Wen, Blair Yang +2
cs.AIarXiv:2604.01591v22026Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation
Xingtong Ge, Yi Zhang, Yushi Huang +6
cs.CVeess.IVarXiv:2604.03118v22026Large Language Models Align with the Human Brain during Creative Thinking
Mete Ismayilzada, Simone A. Luchini, Abdulkadir Gokce +5
q-bio.NCcs.AIcs.CLarXiv:2604.03480v22026BidirLM: From Text to Omnimodal Bidirectional Encoders by Adapting and Composing Causal LLMs
Nicolas Boizard, Théo Deschamps-Berger, Hippolyte Gisserot-Boukhlef +2
cs.CLcs.AIarXiv:2604.02045v12026Adam's Law: Textual Frequency Law on Large Language Models
Hongyuan Adam Lu, Z. L., Victor Wei +5
cs.CLarXiv:2604.02176v32026Expert-Choice Routing Enables Adaptive Computation in Diffusion Language Models
Shuibai Zhang, Caspian Zhuang, Chihan Cui +8
cs.LGcs.CLarXiv:2604.01622v22026Token Warping Helps MLLMs Look from Nearby Viewpoints
Phillip Y. Lee, Chanho Park, Mingue Park +3
cs.CVarXiv:2604.02870v12026AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
Yunhao Feng, Yifan Ding, Yingshui Tan +6
cs.AIarXiv:2604.02947v12026The Geometric Alignment Tax: Tokenization vs. Continuous Geometry in Scientific Foundation Models
Prashant C. Raju
cs.LGcs.ITq-bio.QMarXiv:2604.04155v12026Significance and Stability Analysis of Genotype-Environment Interaction using GxEStat
Meng'en Qin, Zhe Li, Hui Huang +1
cs.CVstat.AParXiv:2604.03337v32026Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization
Paul Hyunbin Cho, Jinhyuk Jang, SeokYoung Lee +7
cs.CVarXiv:2606.11180v12026Disentangling by Factorising
Hyunjik Kim, Andriy Mnih
stat.MLcs.LGarXiv:1802.05983v32018Addressable Memory for Video World Models
Xindi Wu, Sven Elflein, James Lucas +5
cs.CVcs.LGarXiv:2608.07408v12026Certifying and removing disparate impact
Michael Feldman, Sorelle Friedler, John Moeller +2
stat.MLcs.CYarXiv:1412.3756v32014Born Again Neural Networks
Tommaso Furlanello, Zachary C. Lipton, Michael Tschannen +2
stat.MLcs.AIcs.LGarXiv:1805.04770v22018Lost in Translation? Exploring the Shift in Grammatical Gender from Latin to Occitan
Ahan Chatterjee, Matthias Schöffel, Matthias Aßenmacher +2
cs.CLcs.AIarXiv:2605.09156v22026PLUME: Latent Reasoning Based Universal Multimodal Embedding
Chenwei He, Xiangzhao Hao, Tianyu Yang +6
cs.CVarXiv:2604.02073v32026Summaries:한국어ClawArena: Benchmarking AI Agents in Evolving Information Environments
Haonian Ji, Kaiwen Xiong, Siwei Han +9
cs.LGcs.AIcs.CLarXiv:2604.04202v22026GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning
Ornith Team, Xiaoya Li, Guoyin Wang +3
cs.AIarXiv:2604.02721v32026Watch Before You Answer: Learning from Visually Grounded Post-Training
Yuxuan Zhang, EunJeong Hwang, Huaisong Zhang +8
cs.CVcs.AIcs.CLarXiv:2604.05117v12026SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
Yicheng Xiao, Wenhu Zhang, Lin Song +10
cs.CVarXiv:2604.04911v22026General Multimodal Protein Design Enables DNA-Encoding of Chemistry
Jarrid Rector-Brooks, Théophile Lambert, Marta Skreta +15
cs.LGarXiv:2604.05181v12026SciLT: Long-tailed Image Classification under Scientific Image Domains
Jiahao Chen, Bing Su
cs.CVarXiv:2604.03687v32026QEIL v2: Heterogeneous Computing for Edge Intelligence via Roofline-Derived Pareto-Optimal Energy Modeling and Multi-Objective Orchestration
Satyam Kumar, Saurabh Jha
cs.DCarXiv:2602.06057v32026Cog-DRIFT: Exploration on Adaptively Reformulated Instances Enables Learning from Hard Reasoning Problems
Justin Chih-Yao Chen, Archiki Prasad, Zaid Khan +4
cs.LGcs.AIcs.CLarXiv:2604.04767v12026SkillX: Automatically Constructing Skill Knowledge Bases for Agents
Chenxi Wang, Zhuoyun Yu, Xin Xie +8
cs.CLcs.AIcs.IRarXiv:2604.04804v22026TriAttention: Efficient Long Reasoning with Trigonometric KV Compression
Weian Mao, Xi Lin, Wei Huang +5
cs.CLcs.CVarXiv:2604.04921v12026Synthetic Sandbox for Training Machine Learning Engineering Agents
Yuhang Zhou, Lizhu Zhang, Yifan Wu +4
cs.CLcs.LGarXiv:2604.04872v12026InCoder-32B-Thinking: Industrial Code World Model for Thinking
Jian Yang, Wei Zhang, Jiajun Wu +22
cs.ARcs.AIcs.CLarXiv:2604.03144v12026Do Audio-Visual Large Language Models Really See and Hear?
Ramaneswaran Selvakumar, Kaousheik Jayakumar, S Sakshi +3
cs.AIcs.SDarXiv:2604.02605v12026Beyond the Assistant Turn: User Turn Generation as a Probe of Interaction Awareness in Language Models
Sarath Shekkizhar, Romain Cosentino, Adam Earle
cs.AIarXiv:2604.02315v22026LightThinker++: From Reasoning Compression to Memory Management
Yuqi Zhu, Jintian Zhang, Zhenjie Wan +7
cs.CLcs.AIcs.IRarXiv:2604.03679v12026POEMetric: The Last Stanza of Humanity
Bingru Li, Han Wang, Hazel Wilkinson
cs.CLarXiv:2604.03695v12026Training a Student Expert via Semi-Supervised Foundation Model Distillation
Pardis Taghavi, Tian Liu, Renjie Li +2
cs.CVarXiv:2604.03841v12026Can Natural Image Autoencoders Compactly Tokenize fMRI Volumes for Long-Range Dynamics Modeling?
Peter Yongho Kim, Juhyeon Park, Jungwoo Park +4
cs.CVarXiv:2604.03619v12026Squeez: Task-Conditioned Tool-Output Pruning for Coding Agents
Ádám Kovács
cs.SEcs.AIarXiv:2604.04979v12026Can LLMs Learn to Reason Robustly under Noisy Supervision?
Shenzhi Yang, Guangcheng Zhu, Bowen Song +7
cs.LGcs.AIarXiv:2604.03993v12026AURA: Always-On Understanding and Real-Time Assistance via Video Streams
Xudong Lu, Yang Bo, Jinpeng Chen +9
cs.CVarXiv:2604.04184v12026DARE: Diffusion Large Language Models Alignment and Reinforcement Executor
Jingyi Yang, Yuxian Jiang, Xuhao Hu +3
cs.CLarXiv:2604.04215v12026A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning
Tianle Chen, Deepti Ghadiyaram
cs.CVcs.SDarXiv:2604.03995v12026HDP: A Lightweight Cryptographic Protocol for Human Delegation Provenance in Agentic AI Systems
Asiri Dalugoda
cs.CRcs.MAarXiv:2604.04522v12026CLEAR: Unlocking Generative Potential for Degraded Image Understanding in Unified Multimodal Models
Xiangzhao Hao, Zefeng Zhang, Zhenyu Zhang +6
cs.CVarXiv:2604.04780v12026AvatarPointillist: AutoRegressive 4D Gaussian Avatarization
Hongyu Liu, Xuan Wang, Zijian Wu +7
cs.CVarXiv:2604.04787v22026Vero: An Open RL Recipe for General Visual Reasoning
Gabriel Sarch, Linrong Cai, Qunzhong Wang +3
cs.CVcs.AIcs.CLarXiv:2604.04917v32026Paper Espresso: From Paper Overload to Research Insight
Mingzhe Du, Luu Anh Tuan, Dong Huang +1
cs.DLcs.AIarXiv:2604.04562v12026