Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
6,481 to 6,540 of 18,795
Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
Yuanyang Yin, Gongxuan Wang, Yifan Zhan +3
cs.CVarXiv:2608.13546v12026UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos
Yuxuan Zhang, Haozhong Xiong, Jiayi Song +5
cs.CVcs.SDarXiv:2608.11752v22026LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time
Yuxuan Zhang, Haozhong Xiong, Yubo Huang +6
cs.CVarXiv:2608.11745v22026Gaze Target Estimation Anywhere with Concepts
Xu Cao, Houze Yang, Vipin Gunda +5
cs.CVcs.AIarXiv:2608.11367v12026AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss
Mingju Gao, Jingkai Zhou, Kun Gai +2
cs.CVarXiv:2608.11205v12026Beyond Pixels: From Video Priors to 4D Worlds
Zihao Liu, Xiaolong Shen, Zhenglin Zhou +2
cs.CVarXiv:2608.10744v12026Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence
Haoyu Zhang, Zhipeng Li, Xiaoying Tang +2
cs.AIcs.CLcs.CVarXiv:2608.10720v12026UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations
Dvir Samuel, Guy Bar-Shalom, Fabrizio Frasca +4
cs.CVcs.LGarXiv:2608.10835v12026DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation
Zhuchenyang Liu, Ziyi Wang, Yao Zhang +1
cs.IRcs.CLcs.CVarXiv:2608.10636v12026Multimodal Model Diffing for Feature Discovery and Control
Hunar Batra, Lachin Naghashyar, Ashkan Khakzar +4
cs.CVcs.AIcs.CLarXiv:2608.09928v12026Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
Diandian Zhang, Tingyu Song, Lin Fu +2
cs.CVcs.AIarXiv:2608.09873v12026RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance
Dongchi Huang, Hongyin Zhang, Bohan Hou +12
cs.ROcs.CVcs.LGarXiv:2608.09853v12026Vision-Language Grounding as Bidirectional Concept Correspondence
Jieyu Zhang, Ziqi Gao, Luke Zettlemoyer +1
cs.CVcs.AIcs.CLarXiv:2608.07886v12026YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family
Xu Lin, WenJie Nie, Jinlong Peng +4
cs.CVarXiv:2608.07051v12026EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
Ryan Hoque, Peide Huang, David J. Yoon +2
cs.CVcs.LGcs.ROarXiv:2505.11709v32025StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding
Xichen Zhang, Guankai Li, Yinghao Zhu +6
cs.CVarXiv:2608.05703v12026Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
Junlin Han, Shengbang Tong, David Fan +4
cs.CVcs.LGcs.MMarXiv:2608.05000v22026Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning
Kejian Zhu, Zhuoran Jin, Dongqi Huang +4
cs.CVarXiv:2608.03571v22026Quo Vadis, World Modeling?
Yu Yang, Xuemeng Yang, Licheng Wen +17
cs.CVcs.AIcs.ROarXiv:2608.02713v12026Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
Yuqi Liu, Bohao Peng, Zhisheng Zhong +4
cs.CVcs.MMarXiv:2503.06520v32025SULAND v2: A Refined RGB Dataset and Deep Learning Object Detection Benchmark for UAV/UGV-Based SUrface LANDmine Detection Under Domain Shift
Sagar Lekhak, Prasanna Reddy Pulakurthi, Lalit Joshi +2
cs.CVarXiv:2607.28996v12026MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing
Jiajia Lin, Mingxuan Du, Tuowen Zhou +2
cs.CVarXiv:2607.27616v12026Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Xiaomi Robotics Team, Jun Guo, Piaopiao Jin +31
cs.ROcs.CVarXiv:2607.15330v22026Summaries:한국어OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models
Jinsen Su, Yongdong Luo, Yuexiao Ma +3
cs.CVarXiv:2607.23193v32026VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
Xinhao Li, Yuhan Zhu, Xiangyu Zeng +24
cs.CVarXiv:2607.14935v12026Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
Zhishan Zou, Guoyan Sun, Zhiwei Wei +4
cs.CVarXiv:2607.12477v22026RoarNet: A Robust 3D Object Detection based on RegiOn Approximation Refinement
Kiwoo Shin, Youngwook Paul Kwon, Masayoshi Tomizuka
cs.CVarXiv:1811.03818v12018Self-Forcing++: Towards Minute-Scale High-Quality Video Generation
Justin Cui, Jie Wu, Ming Li +6
cs.CVcs.AIarXiv:2510.02283v12025advertorch v0.1: An Adversarial Robustness Toolbox based on PyTorch
Gavin Weiguang Ding, Luyu Wang, Xiaomeng Jin
cs.LGcs.CRcs.CVarXiv:1902.07623v12019HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
Gengluo Li, Xingyu Wan, Shangpin Peng +20
cs.CVarXiv:2607.04884v22026Representation Distribution Matching for One-Step Visual Generation
Lan Feng, Wuyang Li, Eloi Zablocki +2
cs.CVarXiv:2607.02375v12026WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory
Hanlin Wang, Hao Ouyang, Qiuyu Wang +10
cs.CVarXiv:2607.02517v12026AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models
Rintaro Otsubo, Ryo Fujii, Reina Ishikawa +6
cs.CVcs.AIarXiv:2607.02269v12026Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling
Xingyu Zheng, Xianglong Liu, Yifu Ding +4
cs.CVarXiv:2607.01642v12026Summaries:한국어MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering
Dang Quang Thien Tran, Quang V. Dang, Vinamra Tyagi +7
cs.CLcs.AIcs.CVarXiv:2607.01420v32026Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning
Hongxing Li, Xiufeng Huang, Dingming Li +11
cs.CVarXiv:2607.01191v12026SpheRoPE: Zero-Shot Optimization-Free 360 Panorama Generation with Spherical RoPE
Or Hirschorn, Aaron Olender, Eli Alshan +3
cs.CVarXiv:2606.32033v12026Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning
Shijie Li, Yilin Gao, Siyuan Yang +7
cs.CVarXiv:2607.00461v12026MuSViT: A Foundation Vision Model for Sheet Music Representation
Carlos Penarrubia, Antonio Rios-Vila, Eliseo Fuentes-Martinez +4
cs.CVarXiv:2606.31811v12026DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation
Siyu Yan, Yizhen Gao, Yilin Wang +2
cs.CVcs.MAarXiv:2606.31537v12026Physics-Informed Machine Learning: A Survey on Problems, Methods and Applications
Zhongkai Hao, Songming Liu, Yichi Zhang +4
cs.LGcs.AIcs.CVarXiv:2211.08064v22022Comprehensive Review of Deep Learning-Based 3D Point Cloud Completion Processing and Analysis
Ben Fei, Weidong Yang, Wenming Chen +5
cs.CVarXiv:2203.03311v32022The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction
Yuxi Wang, Chengkai Jin, Yufei Liu +6
cs.CVarXiv:2606.30308v22026Summaries:한국어PhotoQuilt: Training-Free Arbitrary-Resolution Photomosaics via Bootstrapped Tiled Denoising
Koorosh Roohi, Javad Rajabi, Andrew Fleet +1
cs.CVarXiv:2606.30968v12026LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing
Xinyu Wang, Chongbo Zhao, Fangneng Zhan +1
cs.CVarXiv:2606.26740v22026Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models
Lianghua Huang, Zhi-Fan Wu, Wei Wang +22
cs.CVcs.AIcs.GRarXiv:2606.25041v32026SkyReels-V2: Infinite-length Film Generative Model
Guibin Chen, Dixuan Lin, Jiangping Yang +22
cs.CVarXiv:2504.13074v32025FlowBender: Feedback-Aware Training for Self-Correcting Conditional Flows
Daniel Gilo, Sven Elflein, Ido Sobol +1
cs.CVarXiv:2606.20404v12026EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies
Ganlin Yang, Zhangzheng Tu, Yuqiang Yang +10
cs.CVarXiv:2606.20092v22026BadWorld: Adversarial Attacks on World Models
Linghui Shen, Mingyue Cui, Xingyi Yang
cs.CVarXiv:2606.16519v12026SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control
Zhengyi Luo, Ye Yuan, Tingwu Wang +25
cs.ROcs.AIcs.CVarXiv:2511.07820v42025GridVQA-X: A Framework for Evaluating Multimodal Explainability Methods
Sujay Belsare, Sudarshan Nikhil, Sushant Kumar +2
cs.CVarXiv:2606.14740v12026Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusion
Weichen Fan, Haiwen Diao, Penghao Wu +1
cs.CVarXiv:2606.15236v22026Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions
Xin Jin, Huanqia Cai, Zhen Li +9
cs.CVarXiv:2606.09076v32026AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding
Qize Yu, Jiadi You, Yuran Wang +10
cs.ROcs.CVcs.MMarXiv:2606.06155v12026MAOAM: Unified Object and Material Selection with Vision-Language Models
Jaden Park, Valentin Deschaintre, Jason Kuen +5
cs.CVarXiv:2606.04880v12026SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models
Olaf Dünkel, Basavaraj Sunagad, Haoran Wang +3
cs.CVarXiv:2605.31597v32026VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
Hidir Yesiltepe, Jiazhen Hu, Tuna Han Salih Meral +4
cs.CVcs.AIarXiv:2605.30351v12026AdaState: Self-Evolving Anchors for Streaming Video Generation
Yusuf Dalva, Pinar Yanardag
cs.CVarXiv:2605.30349v12026Geo-Align: Video Generation Alignment via Metric Geometry Reward
Zizun Li, Haoyu Guo, Runzhe Teng +2
cs.CVarXiv:2605.23903v12026