Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
16,561 to 16,620 of 18,815
R3PM-Net: Real-time, Robust, Real-world Point Matching Network
Yasaman Kashefbahrami, Erkut Akdag, Panagiotis Meletis +3
cs.CVcs.LGarXiv:2604.05060v22026FORGE: Fine-grained Multimodal Evaluation for Manufacturing Scenarios
Xiangru Jian, Hao Xu, Wei Pang +13
cs.CVcs.AIcs.LGarXiv:2604.07413v22026MMEmb-R1: Reasoning-Enhanced Multimodal Embedding with Pair-Aware Selection and Adaptive Control
Yuchi Wang, Haiyang Yu, Weikang Bian +4
cs.CVcs.AIcs.CLarXiv:2604.06156v12026UniSpace: Unified Visual Representation and Scalable Multimodal Modeling
Jinbo Yan, Limeng Qiao, Jie Qin +3
cs.CVcs.AIarXiv:2608.08676v12026HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents
Tencent Robotics X, HY Vision Team, : +20
cs.CVarXiv:2604.07430v12026Personalizing Text-to-Image Generation to Individual Taste
Anne-Sofie Maerten, Juliane Verwiebe, Shyamgopal Karthik +3
cs.CVarXiv:2604.07427v12026Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
Chaoyou Fu, Haozhi Yuan, Yuhao Dong +16
cs.CVarXiv:2604.05015v12026A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
Tommie Kerssies, Gabriele Berton, Ju He +5
cs.CVarXiv:2604.04913v12026Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning
Juekai Lin, Yun Zhu, Honglin Lin +6
cs.CVcs.AIarXiv:2604.06079v12026Action Images: End-to-End Policy Learning via Multiview Video Generation
Haoyu Zhen, Zixian Gao, Qiao Sun +7
cs.CVcs.ROarXiv:2604.06168v22026MoRight: Motion Control Done Right
Shaowei Liu, Xuanchi Ren, Tianchang Shen +5
cs.CVcs.AIcs.GRarXiv:2604.07348v12026Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models
Yuheng Shi, Xiaohuan Pei, Linfeng Wen +2
cs.CVcs.AIarXiv:2604.06912v12026Fast Spatial Memory with Elastic Test-Time Training
Ziqiao Ma, Xueyang Yu, Haoyu Zhen +3
cs.CVcs.GRcs.LGarXiv:2604.07350v12026FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching
Junchao Yi, Rui Zhao, Jiahao Tang +7
cs.CVarXiv:2604.06757v32026FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling
Yitong Li, Junsong Chen, Shuchen Xue +8
cs.LGcs.AIcs.CVarXiv:2604.06916v12026TC-AE: Unlocking Token Capacity for Deep Compression Autoencoders
Teng Li, Ziyuan Huang, Cong Chen +5
cs.CVarXiv:2604.07340v12026INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling
InSpatio Team, Donghui Shen, Guofeng Zhang +20
cs.CVarXiv:2604.07209v22026CylinderDepth: Cylindrical Spatial Attention for Multi-View Consistent Self-Supervised Surround Depth Estimation
Samer Abualhanud, Christian Grannemann, Max Mehltretter
cs.CVarXiv:2511.16428v32025Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images
Yuechen Jiang, Enze Zhang, Md Mohsinul Kabir +4
cs.CVcs.CLcs.MMarXiv:2604.07338v12026RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details
Dewei Zhou, You Li, Zongxin Yang +1
cs.CVarXiv:2604.06870v12026UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards
Jun Wang, Shuo Tan, Zelong Sun +5
cs.CVcs.AIarXiv:2604.14967v22026Small Vision-Language Models are Smart Compressors for Long Video Understanding
Junjie Fei, Jun Chen, Zechun Liu +13
cs.CVcs.AIcs.CLarXiv:2604.08120v12026Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics
Ying Shen, Jerry Xiong, Tianjiao Yu +1
cs.CVarXiv:2604.08503v32026MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
Tanmay Gupta, Piper Wolters, Zixian Ma +13
cs.CVarXiv:2604.08516v12026SIM1: Physics-Aligned Simulator as Zero-Shot Data Scaler in Deformable Worlds
Yunsong Zhou, Hangxu Liu, Xuekun Jiang +12
cs.ROcs.AIcs.CVarXiv:2604.08544v22026AnomalyVFM -- Transforming Vision Foundation Models into Zero-Shot Anomaly Detectors
Matic Fučka, Vitjan Zavrtanik, Danijel Skočaj
cs.CVarXiv:2601.20524v22026Lighting-grounded Video Generation with Renderer-based Agent Reasoning
Ziqi Cai, Taoyu Yang, Zheng Chang +4
cs.CVarXiv:2604.07966v12026Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models
Shilin Yan, Jintao Tong, Hongwei Xue +6
cs.CVcs.AIarXiv:2604.08545v12026On Semiotic-Grounded Interpretive Evaluation of Generative Art
Ruixiang Jiang, Changwen Chen
cs.CVcs.AIcs.HCarXiv:2604.08641v12026WildDet3D: Scaling Promptable 3D Detection in the Wild
Weikai Huang, Jieyu Zhang, Sijun Li +15
cs.CVarXiv:2604.08626v22026Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
Zile Wang, Zexiang Liu, Jiaxing Li +20
cs.CVarXiv:2604.08995v22026Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator
Luozheng Qin, Jia Gong, Qian Qiao +6
cs.CVcs.AIarXiv:2604.08121v12026GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
Mingyu Ouyang, Siyuan Hu, Kevin Qinghong Lin +2
cs.CVcs.AIcs.HCarXiv:2604.07429v12026Envisioning the Future, One Step at a Time
Stefan Andreas Baumann, Jannik Wiese, Tommaso Martorella +2
cs.CVcs.AIcs.LGarXiv:2604.09527v12026MixFlow: Mixed Source Distributions Improve Rectified Flows
Nazir Nayal, Christopher Wewer, Jan Eric Lenssen
cs.CVcs.LGarXiv:2604.09181v12026VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
Guanyu Zhou, Yida Yin, Wenhao Chai +3
cs.CVcs.AIcs.CLarXiv:2604.09531v12026OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks
Wenbo Hu, Xin Chen, Yan Gao-Tian +3
cs.CVcs.AIcs.CLarXiv:2604.08539v22026LPM 1.0: Video-based Character Performance Model
Ailing Zeng, Casper Yang, Chauncey Ge +22
cs.CVcs.AIcs.MMarXiv:2604.07823v22026MegaStyle: Constructing Diverse and Scalable Style Dataset via Consistent Text-to-Image Style Mapping
Junyao Gao, Sibo Liu, Jiaxing Li +6
cs.CVarXiv:2604.08364v22026When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
Zhengyang Sun, Yu Chen, Xin Zhou +4
cs.CVarXiv:2604.08546v12026AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation
Ziwei Zhou, Zeyuan Lai, Rui Wang +6
cs.CVcs.AIcs.CLarXiv:2604.08540v12026Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Video
Chanhyuk Choi, Taesoo Kim, Donggyu Lee +2
cs.CVcs.LGarXiv:2604.07786v22026ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video
Boyuan Wang, Xiaofeng Wang, Yongkang Li +9
cs.CVarXiv:2604.07882v12026Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization
Sai Srinivas Kancheti, Aditya Kanade, Rohit Sinha +2
cs.CVcs.AIarXiv:2604.08476v120263D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding
Makanjuola Ogunleye, Eman Abdelrahman, Ismini Lourentzou
cs.CVcs.AIcs.LGarXiv:2604.08645v12026TAIHRI: Task-Aware 3D Human Keypoints Localization for Close-Range Human-Robot Interaction
Ao Li, Yonggen Ling, Yiyang Lin +3
cs.CVarXiv:2604.08921v12026Strips as Tokens: Artist Mesh Generation with Native UV Segmentation
Rui Xu, Dafei Qin, Kaichun Qiao +8
cs.CVcs.CGcs.GRarXiv:2604.09132v22026Learning Long-term Motion Embeddings for Efficient Kinematics Generation
Nick Stracke, Kolja Bauer, Stefan Andreas Baumann +3
cs.CVarXiv:2604.11737v12026Beyond Perception Errors: Semantic Fixation in Large Vision-Language Models
Md Tanvirul Alam
cs.CVcs.LGarXiv:2604.12119v12026Seeing Through Touch: Tactile-Driven Visual Localization of Material Regions
Seongyu Kim, Seungwoo Lee, Hyeonggon Ryu +2
cs.CVarXiv:2604.11579v12026Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
Zeyue Tian, Binxin Yang, Zhaoyang Liu +8
cs.SDcs.AIcs.CVarXiv:2604.10708v22026Continuous Adversarial Flow Models
Shanchuan Lin, Ceyuan Yang, Zhijie Lin +2
cs.LGcs.CVarXiv:2604.11521v12026Solving Physics Olympiad via Reinforcement Learning on Physics Simulators
Mihir Prabhudesai, Aryan Satpathy, Yangmin Li +6
cs.LGcs.AIcs.CVarXiv:2604.11805v12026LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment
Dujun Nie, Fengjiao Chen, Qi Lv +4
cs.CVcs.ROarXiv:2604.11689v12026Anthropogenic Regional Adaptation in Multimodal Vision-Language Model
Samuel Cahyawijaya, Peerat Limkonchotiwat, Tack Hwa Wong +45
cs.AIcs.CLcs.CVarXiv:2604.11490v22026HDR Video Generation via Latent Alignment with Logarithmic Encoding
Naomi Ken Korem, Mohamed Oumoumad, Harel Cain +6
cs.CVarXiv:2604.11788v12026Zero-shot World Models Are Developmentally Efficient Learners
Khai Loong Aw, Klemen Kotar, Wanhee Lee +6
cs.AIcs.CVarXiv:2604.10333v12026Counting to Four is still a Chore for VLMs
Duy Le Dinh Anh, Patrick Amadeus Irawan, Tuan Van Vo
cs.CVarXiv:2604.10039v12026Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation
Gordon Chen, Ziqi Huang, Ziwei Liu
cs.CVarXiv:2604.10030v12026EditCrafter: Tuning-free High-Resolution Image Editing via Pretrained Diffusion Model
Kunho Kim, Sumin Seo, Yongjun Cho +1
cs.CVarXiv:2604.10268v12026