Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

16,561 to 16,620 of 18,815

  1. R3PM-Net: Real-time, Robust, Real-world Point Matching Network

    Yasaman Kashefbahrami, Erkut Akdag, Panagiotis Meletis +3

    cs.CVcs.LGarXiv:2604.05060v22026
  2. FORGE: Fine-grained Multimodal Evaluation for Manufacturing Scenarios

    Xiangru Jian, Hao Xu, Wei Pang +13

    cs.CVcs.AIcs.LGarXiv:2604.07413v22026
  3. MMEmb-R1: Reasoning-Enhanced Multimodal Embedding with Pair-Aware Selection and Adaptive Control

    Yuchi Wang, Haiyang Yu, Weikang Bian +4

    cs.CVcs.AIcs.CLarXiv:2604.06156v12026
  4. UniSpace: Unified Visual Representation and Scalable Multimodal Modeling

    Jinbo Yan, Limeng Qiao, Jie Qin +3

    cs.CVcs.AIarXiv:2608.08676v12026
  5. HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents

    Tencent Robotics X, HY Vision Team, : +20

    cs.CVarXiv:2604.07430v12026
  6. Personalizing Text-to-Image Generation to Individual Taste

    Anne-Sofie Maerten, Juliane Verwiebe, Shyamgopal Karthik +3

    cs.CVarXiv:2604.07427v12026
  7. Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding

    Chaoyou Fu, Haozhi Yuan, Yuhao Dong +16

    cs.CVarXiv:2604.05015v12026
  8. A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens

    Tommie Kerssies, Gabriele Berton, Ju He +5

    cs.CVarXiv:2604.04913v12026
  9. Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning

    Juekai Lin, Yun Zhu, Honglin Lin +6

    cs.CVcs.AIarXiv:2604.06079v12026
  10. Action Images: End-to-End Policy Learning via Multiview Video Generation

    Haoyu Zhen, Zixian Gao, Qiao Sun +7

    cs.CVcs.ROarXiv:2604.06168v22026
  11. MoRight: Motion Control Done Right

    Shaowei Liu, Xuanchi Ren, Tianchang Shen +5

    cs.CVcs.AIcs.GRarXiv:2604.07348v12026
  12. Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models

    Yuheng Shi, Xiaohuan Pei, Linfeng Wen +2

    cs.CVcs.AIarXiv:2604.06912v12026
  13. Fast Spatial Memory with Elastic Test-Time Training

    Ziqiao Ma, Xueyang Yu, Haoyu Zhen +3

    cs.CVcs.GRcs.LGarXiv:2604.07350v12026
  14. FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching

    Junchao Yi, Rui Zhao, Jiahao Tang +7

    cs.CVarXiv:2604.06757v32026
  15. FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling

    Yitong Li, Junsong Chen, Shuchen Xue +8

    cs.LGcs.AIcs.CVarXiv:2604.06916v12026
  16. TC-AE: Unlocking Token Capacity for Deep Compression Autoencoders

    Teng Li, Ziyuan Huang, Cong Chen +5

    cs.CVarXiv:2604.07340v12026
  17. INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling

    InSpatio Team, Donghui Shen, Guofeng Zhang +20

    cs.CVarXiv:2604.07209v22026
  18. CylinderDepth: Cylindrical Spatial Attention for Multi-View Consistent Self-Supervised Surround Depth Estimation

    Samer Abualhanud, Christian Grannemann, Max Mehltretter

    cs.CVarXiv:2511.16428v32025
  19. Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images

    Yuechen Jiang, Enze Zhang, Md Mohsinul Kabir +4

    cs.CVcs.CLcs.MMarXiv:2604.07338v12026
  20. RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details

    Dewei Zhou, You Li, Zongxin Yang +1

    cs.CVarXiv:2604.06870v12026
  21. UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards

    Jun Wang, Shuo Tan, Zelong Sun +5

    cs.CVcs.AIarXiv:2604.14967v22026
  22. Small Vision-Language Models are Smart Compressors for Long Video Understanding

    Junjie Fei, Jun Chen, Zechun Liu +13

    cs.CVcs.AIcs.CLarXiv:2604.08120v12026
  23. Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics

    Ying Shen, Jerry Xiong, Tianjiao Yu +1

    cs.CVarXiv:2604.08503v32026
  24. MolmoWeb: Open Visual Web Agent and Open Data for the Open Web

    Tanmay Gupta, Piper Wolters, Zixian Ma +13

    cs.CVarXiv:2604.08516v12026
  25. SIM1: Physics-Aligned Simulator as Zero-Shot Data Scaler in Deformable Worlds

    Yunsong Zhou, Hangxu Liu, Xuekun Jiang +12

    cs.ROcs.AIcs.CVarXiv:2604.08544v22026
  26. AnomalyVFM -- Transforming Vision Foundation Models into Zero-Shot Anomaly Detectors

    Matic Fučka, Vitjan Zavrtanik, Danijel Skočaj

    cs.CVarXiv:2601.20524v22026
  27. Lighting-grounded Video Generation with Renderer-based Agent Reasoning

    Ziqi Cai, Taoyu Yang, Zheng Chang +4

    cs.CVarXiv:2604.07966v12026
  28. Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models

    Shilin Yan, Jintao Tong, Hongwei Xue +6

    cs.CVcs.AIarXiv:2604.08545v12026
  29. On Semiotic-Grounded Interpretive Evaluation of Generative Art

    Ruixiang Jiang, Changwen Chen

    cs.CVcs.AIcs.HCarXiv:2604.08641v12026
  30. WildDet3D: Scaling Promptable 3D Detection in the Wild

    Weikai Huang, Jieyu Zhang, Sijun Li +15

    cs.CVarXiv:2604.08626v22026
  31. Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory

    Zile Wang, Zexiang Liu, Jiaxing Li +20

    cs.CVarXiv:2604.08995v22026
  32. Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator

    Luozheng Qin, Jia Gong, Qian Qiao +6

    cs.CVcs.AIarXiv:2604.08121v12026
  33. GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents

    Mingyu Ouyang, Siyuan Hu, Kevin Qinghong Lin +2

    cs.CVcs.AIcs.HCarXiv:2604.07429v12026
  34. Envisioning the Future, One Step at a Time

    Stefan Andreas Baumann, Jannik Wiese, Tommaso Martorella +2

    cs.CVcs.AIcs.LGarXiv:2604.09527v12026
  35. MixFlow: Mixed Source Distributions Improve Rectified Flows

    Nazir Nayal, Christopher Wewer, Jan Eric Lenssen

    cs.CVcs.LGarXiv:2604.09181v12026
  36. VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images

    Guanyu Zhou, Yida Yin, Wenhao Chai +3

    cs.CVcs.AIcs.CLarXiv:2604.09531v12026
  37. OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks

    Wenbo Hu, Xin Chen, Yan Gao-Tian +3

    cs.CVcs.AIcs.CLarXiv:2604.08539v22026
  38. LPM 1.0: Video-based Character Performance Model

    Ailing Zeng, Casper Yang, Chauncey Ge +22

    cs.CVcs.AIcs.MMarXiv:2604.07823v22026
  39. MegaStyle: Constructing Diverse and Scalable Style Dataset via Consistent Text-to-Image Style Mapping

    Junyao Gao, Sibo Liu, Jiaxing Li +6

    cs.CVarXiv:2604.08364v22026
  40. When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models

    Zhengyang Sun, Yu Chen, Xin Zhou +4

    cs.CVarXiv:2604.08546v12026
  41. AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

    Ziwei Zhou, Zeyuan Lai, Rui Wang +6

    cs.CVcs.AIcs.CLarXiv:2604.08540v12026
  42. Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Video

    Chanhyuk Choi, Taesoo Kim, Donggyu Lee +2

    cs.CVcs.LGarXiv:2604.07786v22026
  43. ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video

    Boyuan Wang, Xiaofeng Wang, Yongkang Li +9

    cs.CVarXiv:2604.07882v12026
  44. Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization

    Sai Srinivas Kancheti, Aditya Kanade, Rohit Sinha +2

    cs.CVcs.AIarXiv:2604.08476v12026
  45. 3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding

    Makanjuola Ogunleye, Eman Abdelrahman, Ismini Lourentzou

    cs.CVcs.AIcs.LGarXiv:2604.08645v12026
  46. TAIHRI: Task-Aware 3D Human Keypoints Localization for Close-Range Human-Robot Interaction

    Ao Li, Yonggen Ling, Yiyang Lin +3

    cs.CVarXiv:2604.08921v12026
  47. Strips as Tokens: Artist Mesh Generation with Native UV Segmentation

    Rui Xu, Dafei Qin, Kaichun Qiao +8

    cs.CVcs.CGcs.GRarXiv:2604.09132v22026
  48. Learning Long-term Motion Embeddings for Efficient Kinematics Generation

    Nick Stracke, Kolja Bauer, Stefan Andreas Baumann +3

    cs.CVarXiv:2604.11737v12026
  49. Beyond Perception Errors: Semantic Fixation in Large Vision-Language Models

    Md Tanvirul Alam

    cs.CVcs.LGarXiv:2604.12119v12026
  50. Seeing Through Touch: Tactile-Driven Visual Localization of Material Regions

    Seongyu Kim, Seungwoo Lee, Hyeonggon Ryu +2

    cs.CVarXiv:2604.11579v12026
  51. Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing

    Zeyue Tian, Binxin Yang, Zhaoyang Liu +8

    cs.SDcs.AIcs.CVarXiv:2604.10708v22026
  52. Continuous Adversarial Flow Models

    Shanchuan Lin, Ceyuan Yang, Zhijie Lin +2

    cs.LGcs.CVarXiv:2604.11521v12026
  53. Solving Physics Olympiad via Reinforcement Learning on Physics Simulators

    Mihir Prabhudesai, Aryan Satpathy, Yangmin Li +6

    cs.LGcs.AIcs.CVarXiv:2604.11805v12026
  54. LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment

    Dujun Nie, Fengjiao Chen, Qi Lv +4

    cs.CVcs.ROarXiv:2604.11689v12026
  55. Anthropogenic Regional Adaptation in Multimodal Vision-Language Model

    Samuel Cahyawijaya, Peerat Limkonchotiwat, Tack Hwa Wong +45

    cs.AIcs.CLcs.CVarXiv:2604.11490v22026
  56. HDR Video Generation via Latent Alignment with Logarithmic Encoding

    Naomi Ken Korem, Mohamed Oumoumad, Harel Cain +6

    cs.CVarXiv:2604.11788v12026
  57. Zero-shot World Models Are Developmentally Efficient Learners

    Khai Loong Aw, Klemen Kotar, Wanhee Lee +6

    cs.AIcs.CVarXiv:2604.10333v12026
  58. Counting to Four is still a Chore for VLMs

    Duy Le Dinh Anh, Patrick Amadeus Irawan, Tuan Van Vo

    cs.CVarXiv:2604.10039v12026
  59. Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation

    Gordon Chen, Ziqi Huang, Ziwei Liu

    cs.CVarXiv:2604.10030v12026
  60. EditCrafter: Tuning-free High-Resolution Image Editing via Pretrained Diffusion Model

    Kunho Kim, Sumin Seo, Yongjun Cho +1

    cs.CVarXiv:2604.10268v12026