Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
18,241 to 18,300 of 18,795
How LoRA Remembers? A Parametric Memory Law for LLM Finetuning
Ziwen Xu, Haiwen Hong, Linsong Yu +4
cs.CLcs.AIcs.CVarXiv:2605.30260v12026Why Far Looks Up: Probing Spatial Representation in Vision-Language Models
Cheolhong Min, Jaeyun Jung, Daeun Lee +5
cs.CVarXiv:2605.30161v12026SurGe: Improved Surface Geometry in Point Maps
Karim Knaebel, Gonzalo Martin Garcia, Christian Schmidt +4
cs.CVarXiv:2605.31577v12026PEEK: Picking Essential frames via Efficient Knowledge distillation
Killian Steunou, Anas Filali Razzouki, Khalil Guetari +2
cs.CVarXiv:2605.31029v12026Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?
Liyang Li, Muzhi Zhu, Zhiyue Zhao +5
cs.CVarXiv:2606.01247v12026SmartDirector: Keyframe-Conditioned Cinematic Video Generation with Narrative Pacing Control
Zhida Zhang, Jie Ma, Zhan Peng +5
cs.CVcs.AIarXiv:2605.27891v12026From Pixels to Words -- Towards Native One-Vision Models at Scale
Haiwen Diao, Jiahao Wang, Penghao Wu +18
cs.CVarXiv:2605.28820v12026Colored Noise Diffusion Sampling
Hadar Davidson, Noam Issachar, Sagie Benaim
cs.CVarXiv:2605.30332v12026Linearizing Vision Transformer with Test-Time Training
Yining Li, Dongchen Han, Zeyu Liu +3
cs.CVarXiv:2605.02772v22026VLM3: Vision Language Models Are Native 3D Learners
Zhipeng Cai, Zhuang Liu, Yunyang Xiong +3
cs.CVcs.AIarXiv:2605.30561v12026LVSA: Training-Free Sparse Attention for Long Video Diffusion
Gael Glorian, Ioannis Lamprou, Zhen Zhang +2
cs.CVcs.LGarXiv:2605.31057v12026Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models
Guangzhao He, Rundong Luo, Wei-Chiu Ma +1
cs.CVarXiv:2606.02580v12026Distribution-free false-alarm calibration and chance-corrected spatial evaluation for industrial anomaly detection
Jie Deng
cs.CVcs.LGarXiv:2608.15090v12026Shape Operator PCA: Curvature-Aware Projections for Geometric Machine Learning
Alexandre L. M. Levada
cs.LGcs.AIcs.CVarXiv:2608.15313v12026IP Protection in the Era of Visual Generative AI: A Survey
Zhuan Shi, Shunchang Liu, Alireza Dehghanpour Farashah +8
cs.CVcs.CRcs.LGarXiv:2608.14730v12026Catching Hallucinated Citations in Video-LLM Question Answering: A Self-Verification Pipeline and Verifier Ablation Study
Yogesh Kumar
cs.CVcs.AIarXiv:2608.15574v12026MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding
Bonan Zhang, Shiyu Dong, Quan Hung Tran +9
cs.CVarXiv:2608.17402v12026CETalk: Continuous Valence-Arousal Control for Audio-Driven 3D Talking Head Generation
Peng Jia, Li Dai, Zhen Xiao +2
cs.CVcs.AIarXiv:2608.15110v12026Synthesizing Post-Acetazolamide Cerebral Blood Flow Maps from Baseline MRI in Moyamoya Using 3D Generative AI
Julia Huang, Camila Gonzalez, Rydham Goyal +5
eess.IVcs.AIcs.CVarXiv:2608.14758v12026Earth Observation Foundation Models for Terrestrial Ecohydrology: From Representation Learning to Process Inference
Yi Yu, Jian Peng, Yucheng Lin +2
cs.LGcs.CVphysics.bio-pharXiv:2608.15282v12026Zero-Shot Adaptation of Medical Vision Foundation Models for High-Frequency Micro-Ultrasound Prostate Segmentation
Ayusha Abbas, Saram Abbas, Kabita Adhikari
cs.CVcs.LGarXiv:2608.14796v12026Cross-Modal Ultrasound-MRI Learning for Fetal Brain Ventricular Volumetry and Abnormality Screening
Yuhao Huang, Yuanji Zhang, Yuhuan Lu +3
eess.IVcs.AIcs.CVarXiv:2608.14763v12026Uncertainty Identifies Difficult Samples Across Methods: A Multi-Task Study on a Heterogeneous Skin Lesion Dataset
Leon Koole, Jiapan Guo, Matias Valdenegro-Toro
cs.CVcs.LGarXiv:2608.14768v12026Teach and Grow: An Agent-Centered Architecture for General Robot Learning
Chang Nie, Zhe Liu, Hesheng Wang
cs.ROcs.AIcs.CVarXiv:2608.17209v12026The 10th AI City Challenge
Zheng Tang, Shuo Wang, David C. Anastasiu +34
cs.CVcs.AIarXiv:2608.17044v12026YILDIZ-VPR: A Novel Dataset with Dense Coverage Under Diverse Environmental Conditions for Visual Place Recognition
Serdar Yildiz, Abbas Memiş, Songül Varli
cs.CVcs.AIarXiv:2608.17033v12026MeshWeaver: Sparse-Voxel-Guided Surface Weaving for Autoregressive Mesh Generation
Jiale Xu, Wang Zhao, Ying Shan
cs.CVarXiv:2606.04688v12026Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?
Jiaqi Tang, Jianmin Chen, Youyang Zhai +6
cs.CVcs.AIcs.CLarXiv:2606.08063v12026CoVEBench: Can Video Editing Models Handle Complex Instructions?
Jiangtao Wu, Jiaming Wang, Yiwen He +7
cs.CVcs.AIarXiv:2606.08415v22026DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning
Hangui Lin, Yan Shu, Zhengyang Liang +6
cs.CVarXiv:2606.08035v12026ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations
Junke Wang, Xiao Wang, Jiacheng Pan +16
cs.CVarXiv:2606.11188v12026VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization
Junhao Cheng, Liang Hou, Tianxiong Zhong +4
cs.CVarXiv:2606.02564v32026Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching
Hao Zhong, Muzhi Zhu, Shenyan Zeng +8
cs.CVarXiv:2606.03577v12026VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding
Lin Fu, Zheyuan Yang, Yang Wang +3
cs.CVarXiv:2606.05259v12026WaveDiT: Distribution-Aware Wavelet Flow Matching for Efficient 3D Brain MRI Synthesis
Danilo Danese, Angela Lombardi, Giuseppe Fasano +2
cs.CVarXiv:2606.08670v22026Echo-Memory: A Controlled Study of Memory in Action World Models
Wayne King, Zeyue Xue, Yuxuan Bian +13
cs.CVcs.GRcs.LGarXiv:2606.09803v12026A Cookbook of 3D Vision: Data, Learning Paradigms, and Application
Hongyang Du, Zongxia Li, Dawei Liu +8
cs.CVarXiv:2606.04291v12026LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation
Qixin Hu, Shuai Yang, Wei Huang +2
cs.CVarXiv:2606.02553v12026WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark
Yida Yin, Harish Krishnakumar, Chung Peng Lee +9
cs.CVarXiv:2606.06538v12026Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback
Huaisong Zhang, Hao Yu, Yuxuan Zhang +7
cs.CVarXiv:2606.06113v22026IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder
Yitong Chen, Zijie Diao, Junke Wang +5
cs.CVarXiv:2606.11096v12026Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning
Chaofan Ma, Zhenjie Mao, Yuhuan Yang +5
cs.CVcs.AIarXiv:2606.11683v12026Native Active Perception as Reasoning for Omni-Modal Understanding
Zhenghao Xing, Ruiyang Xu, Yuxuan Wang +8
cs.CVcs.CLcs.SDarXiv:2606.19341v22026InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning
Ziang Yan, Sheng Xia, Jiashuo Yu +10
cs.CVarXiv:2606.12195v12026HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers
Guozhen Zhang, Xuerui Qiu, Yutao Cui +11
cs.CVcs.AIarXiv:2606.13289v12026ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning
Sicheng Yang, Hangjie Yuan, Wenjun Zhang +5
cs.CVcs.AIcs.CLarXiv:2606.14697v12026IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products
Haonan Qi, Jin Cao, Yongqi Zhang +9
cs.CVarXiv:2606.14383v22026Human Universal Grasping
Kevin Yuanbo Wu, Tianxing Zhou, Isaac Tu +5
cs.ROcs.AIcs.CVarXiv:2606.17054v12026Looped World Models
Hongyuan Adam Lu, Z. L. Victor Wei, Qun Zhang +28
cs.LGcs.AIcs.CLarXiv:2606.18208v12026BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation
Max Van Puyvelde, Ibrahim Gulluk, Wim Van Criekinge +1
cs.AIcs.CVcs.LGarXiv:2606.19651v22026CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing
Fuchen Long, Cong Wang, Zitao Gao +8
cs.CVarXiv:2608.17566v12026Physics-IQ Verified
Tim Rädsch, Yuki M Asano, Hilde Kuehne +4
cs.CVarXiv:2606.18943v12026Continuous-Time Distribution Matching for Few-Step Diffusion Distillation
Tao Liu, Hao Yan, Mengting Chen +8
cs.CVcs.AIarXiv:2605.06376v12026CPCANet: Deep Unfolding Common Principal Component Analysis for Domain Generalization
Yu-Hsi Chen, Abd-Krim Seghouane
cs.CVarXiv:2605.05136v32026RT-Splatting: Joint Reflection-Transmission Modeling with Gaussian Splatting
Ji Shi, Xianghua Ying, Bowei Xing +2
cs.CVarXiv:2605.18263v12026Swift Sampling: Selecting Temporal Surprises via Taylor Series
Dahye Kim, Bhuvan Sachdeva, Karan Uppal +3
cs.CVcs.AIarXiv:2605.22678v12026NeuROK: Generative 4D Neural Object Kinematics
Chen Geng, Guangzhao He, Yue Gao +3
cs.CVcs.GRarXiv:2605.30347v12026VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
Mingjian Gao, Wenqiao Zhang, Yuqian Yuan +9
cs.CVcs.AIarXiv:2605.30011v12026ABot-Earth 0.5: Generative 3D Earth Model
Ming Qian, Tianjian Ouyang, Mingchao Sun +25
cs.CVarXiv:2606.09967v12026Memento: Reconstruct to Remember for Consistent Long Video Generation
Xuan Wei, Longbin Ji, Guan Wang +5
cs.CVarXiv:2606.14667v12026