Every paper with a summary

Every arXiv paper Paperlayer has summarized so far. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

54,661 to 54,720 of 61,210

  1. FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification

    Ling Yue, Chaoqian Ouyang, Hang Xu +7

    cs.AIcs.LGarXiv:2604.04074v42026
  2. Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability

    Qihan Ren, Peng Wang, Ruikun Cai +8

    cs.AIarXiv:2604.06628v22026
  3. Qualixar OS: A Universal Operating System for AI Agent Orchestration

    Varun Pratap Bhardwaj

    cs.AIcs.MAcs.SEarXiv:2604.06392v12026
  4. UniSpace: Unified Visual Representation and Scalable Multimodal Modeling

    Jinbo Yan, Limeng Qiao, Jie Qin +3

    cs.CVcs.AIarXiv:2608.08676v12026
  5. HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents

    Tencent Robotics X, HY Vision Team, : +20

    cs.CVarXiv:2604.07430v12026
  6. OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence

    Jianhui Liu, Haoze Sun, Wenbo Li +11

    cs.CLarXiv:2604.07296v22026
  7. Personalizing Text-to-Image Generation to Individual Taste

    Anne-Sofie Maerten, Juliane Verwiebe, Shyamgopal Karthik +3

    cs.CVarXiv:2604.07427v12026
  8. MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU

    Zhengqing Yuan, Hanchi Sun, Lichao Sun +1

    cs.CLcs.DCcs.OSarXiv:2604.05091v12026
  9. Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding

    Chaoyou Fu, Haozhi Yuan, Yuhao Dong +16

    cs.CVarXiv:2604.05015v12026
  10. A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens

    Tommie Kerssies, Gabriele Berton, Ju He +5

    cs.CVarXiv:2604.04913v12026
  11. Beyond Hard Negatives: The Importance of Score Distribution in Knowledge Distillation for Dense Retrieval

    Youngjoon Jang, Seongtae Hong, Hyeonseok Moon +1

    cs.IRarXiv:2604.04734v22026
  12. STEER: Structured Event Evidence for Video Reasoning via Multi-Objective Reinforcement Learning

    Zinuo Li, Yongxin Guo, Jun Liu +7

    cs.CLarXiv:2604.04415v32026
  13. SkVM: Revisiting Language VM for Skills across Heterogenous LLMs and Harnesses

    Le Chen, Erhu Feng, Yubin Xia +1

    cs.SEcs.LGarXiv:2604.03088v32026
  14. Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning

    Juekai Lin, Yun Zhu, Honglin Lin +6

    cs.CVcs.AIarXiv:2604.06079v12026
  15. RAGEN-2: Reasoning Collapse in Agentic RL

    Zihan Wang, Chi Gui, Xing Jin +13

    cs.LGarXiv:2604.06268v12026
  16. Target Policy Optimization

    Jean Kaddour

    cs.LGarXiv:2604.06159v12026
  17. TRACE: Capability-Targeted Agentic Training

    Hangoo Kang, Tarun Suresh, Jon Saad-Falcon +1

    cs.AIarXiv:2604.05336v22026
  18. The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment

    Rishab Balasubramanian, Pin-Jie Lin, Rituraj Sharma +6

    cs.LGcs.AIarXiv:2604.06377v32026
  19. Action Images: End-to-End Policy Learning via Multiview Video Generation

    Haoyu Zhen, Zixian Gao, Qiao Sun +7

    cs.CVcs.ROarXiv:2604.06168v22026
  20. Experience Transfer for Multimodal LLM Agents in Minecraft Game

    Chenghao Li, Jun Liu, Songbo Zhang +7

    cs.AIarXiv:2604.05533v12026
  21. The Depth Ceiling: On the Limits of Large Language Models in Discovering Latent Planning

    Yi Xu, Philipp Jettkant, Laura Ruis

    cs.LGcs.AIcs.CLarXiv:2604.06427v12026
  22. In-Place Test-Time Training

    Guhao Feng, Shengjie Luo, Kai Hua +4

    cs.LGcs.AIcs.CLarXiv:2604.06169v12026
  23. Beyond Accuracy: Unveiling Inefficiency Patterns in Tool-Integrated Reasoning

    Qisheng Su, Shiting Huang, Zhen Fang +3

    cs.PFcs.SEarXiv:2604.05404v22026
  24. Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework

    Komal Kumar, Aman Chadha, Salman Khan +2

    cs.CLarXiv:2604.06170v12026
  25. Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents

    Bowen Ye, Rang Li, Qibin Yang +10

    cs.AIarXiv:2604.06132v32026
  26. Graph-Based Chain-of-Thought Pruning for Reducing Redundant Reflections in Reasoning LLMs

    Hongyuan Yuan, Xinran He, Run Shao +6

    cs.CLarXiv:2604.05643v22026
  27. QiMeng-PRepair: Precise Code Repair via Edit-Aware Reward Optimization

    Changxin Ke, Rui Zhang, Jiaming Guo +10

    cs.SEcs.LGarXiv:2604.05963v12026
  28. Improving Semantic Proximity in Information Retrieval through Cross-Lingual Alignment

    Seongtae Hong, Youngjoon Jang, Jungseob Lee +2

    cs.IRarXiv:2604.05684v12026
  29. AgentGL: Towards Agentic Graph Learning with LLMs via Reinforcement Learning

    Yuanfu Sun, Kang Li, Dongzhe Fan +2

    cs.CLarXiv:2604.05846v22026
  30. Spec Kit Agents: Context-Grounded Agentic Workflows

    Pardis Taghavi, Santosh Bhavani

    cs.SEcs.AIcs.MAarXiv:2604.05278v12026
  31. On the Step Length Confounding in LLM Reasoning Data Selection

    Bing Wang, Rui Miao, Chen Shen +7

    cs.CLcs.AIarXiv:2604.06834v12026
  32. An LLM agent for end-to-end computational materials discovery

    Chen Yuntong, Huang Ju, Liu Yu +6

    cond-mat.mtrl-scics.AIarXiv:2608.20434v12026
  33. MoRight: Motion Control Done Right

    Shaowei Liu, Xuanchi Ren, Tianchang Shen +5

    cs.CVcs.AIcs.GRarXiv:2604.07348v12026
  34. Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models

    Yuheng Shi, Xiaohuan Pei, Linfeng Wen +2

    cs.CVcs.AIarXiv:2604.06912v12026
  35. Fast Spatial Memory with Elastic Test-Time Training

    Ziqiao Ma, Xueyang Yu, Haoyu Zhen +3

    cs.CVcs.GRcs.LGarXiv:2604.07350v12026
  36. FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching

    Junchao Yi, Rui Zhao, Jiahao Tang +7

    cs.CVarXiv:2604.06757v32026
  37. Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization

    Qiyao Ma, Dechen Gao, Rui Cai +4

    cs.CLcs.LGarXiv:2604.07343v22026
  38. FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling

    Yitong Li, Junsong Chen, Shuchen Xue +8

    cs.LGcs.AIcs.CVarXiv:2604.06916v12026
  39. TC-AE: Unlocking Token Capacity for Deep Compression Autoencoders

    Teng Li, Ziyuan Huang, Cong Chen +5

    cs.CVarXiv:2604.07340v12026
  40. MARS: Enabling Autoregressive Models Multi-Token Generation

    Ziqi Jin, Lei Wang, Ziwei Luo +1

    cs.CLarXiv:2604.07023v12026
  41. INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling

    InSpatio Team, Donghui Shen, Guofeng Zhang +20

    cs.CVarXiv:2604.07209v22026
  42. CylinderDepth: Cylindrical Spatial Attention for Multi-View Consistent Self-Supervised Surround Depth Estimation

    Samer Abualhanud, Christian Grannemann, Max Mehltretter

    cs.CVarXiv:2511.16428v32025
  43. Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images

    Yuechen Jiang, Enze Zhang, Md Mohsinul Kabir +4

    cs.CVcs.CLcs.MMarXiv:2604.07338v12026
  44. Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference

    Quantong Qiu, Zhiyi Hong, Yi Yang +5

    cs.LGcs.CLarXiv:2604.07394v12026
  45. RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details

    Dewei Zhou, You Li, Zongxin Yang +1

    cs.CVarXiv:2604.06870v12026
  46. UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards

    Jun Wang, Shuo Tan, Zelong Sun +5

    cs.CVcs.AIarXiv:2604.14967v22026
  47. Small Vision-Language Models are Smart Compressors for Long Video Understanding

    Junjie Fei, Jun Chen, Zechun Liu +13

    cs.CVcs.AIcs.CLarXiv:2604.08120v12026
  48. Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics

    Ying Shen, Jerry Xiong, Tianjiao Yu +1

    cs.CVarXiv:2604.08503v32026
  49. Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces

    Jiawei Chen, Ruoxi Xu, Boxi Cao +11

    cs.CLcs.AIcs.LGarXiv:2604.08362v22026
  50. MolmoWeb: Open Visual Web Agent and Open Data for the Open Web

    Tanmay Gupta, Piper Wolters, Zixian Ma +13

    cs.CVarXiv:2604.08516v12026
  51. Structured Distillation of Web Agent Capabilities Enables Generalization

    Xing Han Lù, Siva Reddy

    cs.LGarXiv:2604.07776v12026
  52. SIM1: Physics-Aligned Simulator as Zero-Shot Data Scaler in Deformable Worlds

    Yunsong Zhou, Hangxu Liu, Xuekun Jiang +12

    cs.ROcs.AIcs.CVarXiv:2604.08544v22026
  53. AnomalyVFM -- Transforming Vision Foundation Models into Zero-Shot Anomaly Detectors

    Matic Fučka, Vitjan Zavrtanik, Danijel Skočaj

    cs.CVarXiv:2601.20524v22026
  54. Lighting-grounded Video Generation with Renderer-based Agent Reasoning

    Ziqi Cai, Taoyu Yang, Zheng Chang +4

    cs.CVarXiv:2604.07966v12026
  55. What do Language Models Learn and When? The Implicit Curriculum Hypothesis

    Emmy Liu, Kaiser Sun, Millicent Li +4

    cs.CLarXiv:2604.08510v12026
  56. Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models

    Shilin Yan, Jintao Tong, Hongwei Xue +6

    cs.CVcs.AIarXiv:2604.08545v12026
  57. Meta-learning In-Context Enables Training-Free Cross Subject Brain Decoding

    Mu Nan, Muquan Yu, Weijian Mai +12

    cs.LGq-bio.NCarXiv:2604.08537v12026
  58. On Semiotic-Grounded Interpretive Evaluation of Generative Art

    Ruixiang Jiang, Changwen Chen

    cs.CVcs.AIcs.HCarXiv:2604.08641v12026
  59. KnowU-Bench: Towards Interactive, Proactive, and Personalized Mobile Agent Evaluation

    Tongbo Chen, Zhengxi Lu, Zhan Xu +13

    cs.AIarXiv:2604.08455v12026
  60. WildDet3D: Scaling Promptable 3D Detection in the Wild

    Weikai Huang, Jieyu Zhang, Sijun Li +15

    cs.CVarXiv:2604.08626v22026