Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

18,241 to 18,300 of 18,795

  1. How LoRA Remembers? A Parametric Memory Law for LLM Finetuning

    Ziwen Xu, Haiwen Hong, Linsong Yu +4

    cs.CLcs.AIcs.CVarXiv:2605.30260v12026
  2. Why Far Looks Up: Probing Spatial Representation in Vision-Language Models

    Cheolhong Min, Jaeyun Jung, Daeun Lee +5

    cs.CVarXiv:2605.30161v12026
  3. SurGe: Improved Surface Geometry in Point Maps

    Karim Knaebel, Gonzalo Martin Garcia, Christian Schmidt +4

    cs.CVarXiv:2605.31577v12026
  4. PEEK: Picking Essential frames via Efficient Knowledge distillation

    Killian Steunou, Anas Filali Razzouki, Khalil Guetari +2

    cs.CVarXiv:2605.31029v12026
  5. Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?

    Liyang Li, Muzhi Zhu, Zhiyue Zhao +5

    cs.CVarXiv:2606.01247v12026
  6. SmartDirector: Keyframe-Conditioned Cinematic Video Generation with Narrative Pacing Control

    Zhida Zhang, Jie Ma, Zhan Peng +5

    cs.CVcs.AIarXiv:2605.27891v12026
  7. From Pixels to Words -- Towards Native One-Vision Models at Scale

    Haiwen Diao, Jiahao Wang, Penghao Wu +18

    cs.CVarXiv:2605.28820v12026
  8. Colored Noise Diffusion Sampling

    Hadar Davidson, Noam Issachar, Sagie Benaim

    cs.CVarXiv:2605.30332v12026
  9. Linearizing Vision Transformer with Test-Time Training

    Yining Li, Dongchen Han, Zeyu Liu +3

    cs.CVarXiv:2605.02772v22026
  10. VLM3: Vision Language Models Are Native 3D Learners

    Zhipeng Cai, Zhuang Liu, Yunyang Xiong +3

    cs.CVcs.AIarXiv:2605.30561v12026
  11. LVSA: Training-Free Sparse Attention for Long Video Diffusion

    Gael Glorian, Ioannis Lamprou, Zhen Zhang +2

    cs.CVcs.LGarXiv:2605.31057v12026
  12. Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models

    Guangzhao He, Rundong Luo, Wei-Chiu Ma +1

    cs.CVarXiv:2606.02580v12026
  13. Distribution-free false-alarm calibration and chance-corrected spatial evaluation for industrial anomaly detection

    Jie Deng

    cs.CVcs.LGarXiv:2608.15090v12026
  14. Shape Operator PCA: Curvature-Aware Projections for Geometric Machine Learning

    Alexandre L. M. Levada

    cs.LGcs.AIcs.CVarXiv:2608.15313v12026
  15. IP Protection in the Era of Visual Generative AI: A Survey

    Zhuan Shi, Shunchang Liu, Alireza Dehghanpour Farashah +8

    cs.CVcs.CRcs.LGarXiv:2608.14730v12026
  16. Catching Hallucinated Citations in Video-LLM Question Answering: A Self-Verification Pipeline and Verifier Ablation Study

    Yogesh Kumar

    cs.CVcs.AIarXiv:2608.15574v12026
  17. MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding

    Bonan Zhang, Shiyu Dong, Quan Hung Tran +9

    cs.CVarXiv:2608.17402v12026
  18. CETalk: Continuous Valence-Arousal Control for Audio-Driven 3D Talking Head Generation

    Peng Jia, Li Dai, Zhen Xiao +2

    cs.CVcs.AIarXiv:2608.15110v12026
  19. Synthesizing Post-Acetazolamide Cerebral Blood Flow Maps from Baseline MRI in Moyamoya Using 3D Generative AI

    Julia Huang, Camila Gonzalez, Rydham Goyal +5

    eess.IVcs.AIcs.CVarXiv:2608.14758v12026
  20. Earth Observation Foundation Models for Terrestrial Ecohydrology: From Representation Learning to Process Inference

    Yi Yu, Jian Peng, Yucheng Lin +2

    cs.LGcs.CVphysics.bio-pharXiv:2608.15282v12026
  21. Zero-Shot Adaptation of Medical Vision Foundation Models for High-Frequency Micro-Ultrasound Prostate Segmentation

    Ayusha Abbas, Saram Abbas, Kabita Adhikari

    cs.CVcs.LGarXiv:2608.14796v12026
  22. Cross-Modal Ultrasound-MRI Learning for Fetal Brain Ventricular Volumetry and Abnormality Screening

    Yuhao Huang, Yuanji Zhang, Yuhuan Lu +3

    eess.IVcs.AIcs.CVarXiv:2608.14763v12026
  23. Uncertainty Identifies Difficult Samples Across Methods: A Multi-Task Study on a Heterogeneous Skin Lesion Dataset

    Leon Koole, Jiapan Guo, Matias Valdenegro-Toro

    cs.CVcs.LGarXiv:2608.14768v12026
  24. Teach and Grow: An Agent-Centered Architecture for General Robot Learning

    Chang Nie, Zhe Liu, Hesheng Wang

    cs.ROcs.AIcs.CVarXiv:2608.17209v12026
  25. The 10th AI City Challenge

    Zheng Tang, Shuo Wang, David C. Anastasiu +34

    cs.CVcs.AIarXiv:2608.17044v12026
  26. YILDIZ-VPR: A Novel Dataset with Dense Coverage Under Diverse Environmental Conditions for Visual Place Recognition

    Serdar Yildiz, Abbas Memiş, Songül Varli

    cs.CVcs.AIarXiv:2608.17033v12026
  27. MeshWeaver: Sparse-Voxel-Guided Surface Weaving for Autoregressive Mesh Generation

    Jiale Xu, Wang Zhao, Ying Shan

    cs.CVarXiv:2606.04688v12026
  28. Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?

    Jiaqi Tang, Jianmin Chen, Youyang Zhai +6

    cs.CVcs.AIcs.CLarXiv:2606.08063v12026
  29. CoVEBench: Can Video Editing Models Handle Complex Instructions?

    Jiangtao Wu, Jiaming Wang, Yiwen He +7

    cs.CVcs.AIarXiv:2606.08415v22026
  30. DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning

    Hangui Lin, Yan Shu, Zhengyang Liang +6

    cs.CVarXiv:2606.08035v12026
  31. ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations

    Junke Wang, Xiao Wang, Jiacheng Pan +16

    cs.CVarXiv:2606.11188v12026
  32. VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization

    Junhao Cheng, Liang Hou, Tianxiong Zhong +4

    cs.CVarXiv:2606.02564v32026
  33. Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching

    Hao Zhong, Muzhi Zhu, Shenyan Zeng +8

    cs.CVarXiv:2606.03577v12026
  34. VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding

    Lin Fu, Zheyuan Yang, Yang Wang +3

    cs.CVarXiv:2606.05259v12026
  35. WaveDiT: Distribution-Aware Wavelet Flow Matching for Efficient 3D Brain MRI Synthesis

    Danilo Danese, Angela Lombardi, Giuseppe Fasano +2

    cs.CVarXiv:2606.08670v22026
  36. Echo-Memory: A Controlled Study of Memory in Action World Models

    Wayne King, Zeyue Xue, Yuxuan Bian +13

    cs.CVcs.GRcs.LGarXiv:2606.09803v12026
  37. A Cookbook of 3D Vision: Data, Learning Paradigms, and Application

    Hongyang Du, Zongxia Li, Dawei Liu +8

    cs.CVarXiv:2606.04291v12026
  38. LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation

    Qixin Hu, Shuai Yang, Wei Huang +2

    cs.CVarXiv:2606.02553v12026
  39. WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark

    Yida Yin, Harish Krishnakumar, Chung Peng Lee +9

    cs.CVarXiv:2606.06538v12026
  40. Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback

    Huaisong Zhang, Hao Yu, Yuxuan Zhang +7

    cs.CVarXiv:2606.06113v22026
  41. IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder

    Yitong Chen, Zijie Diao, Junke Wang +5

    cs.CVarXiv:2606.11096v12026
  42. Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning

    Chaofan Ma, Zhenjie Mao, Yuhuan Yang +5

    cs.CVcs.AIarXiv:2606.11683v12026
  43. Native Active Perception as Reasoning for Omni-Modal Understanding

    Zhenghao Xing, Ruiyang Xu, Yuxuan Wang +8

    cs.CVcs.CLcs.SDarXiv:2606.19341v22026
  44. InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning

    Ziang Yan, Sheng Xia, Jiashuo Yu +10

    cs.CVarXiv:2606.12195v12026
  45. HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

    Guozhen Zhang, Xuerui Qiu, Yutao Cui +11

    cs.CVcs.AIarXiv:2606.13289v12026
  46. ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning

    Sicheng Yang, Hangjie Yuan, Wenjun Zhang +5

    cs.CVcs.AIcs.CLarXiv:2606.14697v12026
  47. IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products

    Haonan Qi, Jin Cao, Yongqi Zhang +9

    cs.CVarXiv:2606.14383v22026
  48. Human Universal Grasping

    Kevin Yuanbo Wu, Tianxing Zhou, Isaac Tu +5

    cs.ROcs.AIcs.CVarXiv:2606.17054v12026
  49. Looped World Models

    Hongyuan Adam Lu, Z. L. Victor Wei, Qun Zhang +28

    cs.LGcs.AIcs.CLarXiv:2606.18208v12026
  50. BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation

    Max Van Puyvelde, Ibrahim Gulluk, Wim Van Criekinge +1

    cs.AIcs.CVcs.LGarXiv:2606.19651v22026
  51. CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing

    Fuchen Long, Cong Wang, Zitao Gao +8

    cs.CVarXiv:2608.17566v12026
  52. Physics-IQ Verified

    Tim Rädsch, Yuki M Asano, Hilde Kuehne +4

    cs.CVarXiv:2606.18943v12026
  53. Continuous-Time Distribution Matching for Few-Step Diffusion Distillation

    Tao Liu, Hao Yan, Mengting Chen +8

    cs.CVcs.AIarXiv:2605.06376v12026
  54. CPCANet: Deep Unfolding Common Principal Component Analysis for Domain Generalization

    Yu-Hsi Chen, Abd-Krim Seghouane

    cs.CVarXiv:2605.05136v32026
  55. RT-Splatting: Joint Reflection-Transmission Modeling with Gaussian Splatting

    Ji Shi, Xianghua Ying, Bowei Xing +2

    cs.CVarXiv:2605.18263v12026
  56. Swift Sampling: Selecting Temporal Surprises via Taylor Series

    Dahye Kim, Bhuvan Sachdeva, Karan Uppal +3

    cs.CVcs.AIarXiv:2605.22678v12026
  57. NeuROK: Generative 4D Neural Object Kinematics

    Chen Geng, Guangzhao He, Yue Gao +3

    cs.CVcs.GRarXiv:2605.30347v12026
  58. VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies

    Mingjian Gao, Wenqiao Zhang, Yuqian Yuan +9

    cs.CVcs.AIarXiv:2605.30011v12026
  59. ABot-Earth 0.5: Generative 3D Earth Model

    Ming Qian, Tianjian Ouyang, Mingchao Sun +25

    cs.CVarXiv:2606.09967v12026
  60. Memento: Reconstruct to Remember for Consistent Long Video Generation

    Xuan Wei, Longbin Ji, Guan Wang +5

    cs.CVarXiv:2606.14667v12026