Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
3,781 to 3,840 of 18,852
Learning to Match Features with Seeded Graph Matching Network
Hongkai Chen, Zixin Luo, Jiahui Zhang +5
cs.CVarXiv:2108.08771v12021Learning What to Learn for Video Object Segmentation
Goutam Bhat, Felix Järemo Lawin, Martin Danelljan +4
cs.CVarXiv:2003.11540v22020Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning
Yihong Huang, Fei Ma, Yihua Shao +4
cs.CVcs.AIcs.CLarXiv:2602.02951v12026Matcher: Segment Anything with One Shot Using All-Purpose Feature Matching
Yang Liu, Muzhi Zhu, Hengtao Li +3
cs.CVarXiv:2305.13310v22023Learning to Ground Before Reading: Unified PCB Engineering Drawing Parsing with Compact Vision-Language Models
Jinghao Liu, Xingrun Liu, Gengchen Sun +3
cs.CVcs.MMarXiv:2608.29268v12026Oriented Edge Forests for Boundary Detection
Sam Hallman, Charless C. Fowlkes
cs.CVarXiv:1412.4181v22014Deep Semantic Segmentation for Automated Driving: Taxonomy, Roadmap and Challenges
Mennatullah Siam, Sara Elkerdawy, Martin Jagersand +1
stat.MLcs.CVarXiv:1707.02432v22017Extended depth-of-field in holographic image reconstruction using deep learning based auto-focusing and phase-recovery
Yichen Wu, Yair Rivenson, Yibo Zhang +4
cs.CVcs.LGphysics.opticsarXiv:1803.08138v12018Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation
Zengqun Zhao, Yanzuo Lu, Ziquan Liu +3
cs.CVarXiv:2603.21366v22026High throughput quantitative metallography for complex microstructures using deep learning: A case study in ultrahigh carbon steel
Brian L. DeCost, Bo Lei, Toby Francis +1
cs.CVarXiv:1805.08693v22018Action Transformer: A Self-Attention Model for Short-Time Pose-Based Human Action Recognition
Vittorio Mazzia, Simone Angarano, Francesco Salvetti +2
cs.CVcs.LGarXiv:2107.00606v62021Temperature-Adaptive Transformed Teacher Matching
Hiroaki Aizawa, Yoshikazu Hayashi
cs.LGcs.CVarXiv:2608.29099v12026BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks
Yixiang Chen, Peiyan Li, Jiabing Yang +8
cs.ROcs.CVarXiv:2602.03793v12026Enhancing Low-Cost Video Editing with Lightweight Adaptors and Temporal-Aware Inversion
Yangfan He, Sida Li, Jianhui Wang +11
cs.CVarXiv:2501.04606v42025Growing a Stand, Not a Tree: Joint Canopy Generation Reproduces Crown Shyness
Guang Yang, Fengchen Liu
cs.CVarXiv:2608.28692v12026FineCIR: Explicit Parsing of Fine-Grained Modification Semantics for Composed Image Retrieval
Zixu Li, Zhiheng Fu, Yupeng Hu +3
cs.CVcs.AIarXiv:2503.21309v12025Continuous Dice Coefficient: a Method for Evaluating Probabilistic Segmentations
Reuben R Shamir, Yuval Duchin, Jinyoung Kim +2
cs.CVeess.IVarXiv:1906.11031v12019Median K-flats for hybrid linear modeling with many outliers
Teng Zhang, Arthur Szlam, Gilad Lerman
cs.CVcs.LGarXiv:0909.3123v12009Knowledge-Guided Deep Fractal Neural Networks for Human Pose Estimation
Guanghan Ning, Zhi Zhang, Zhihai He
cs.CVarXiv:1705.02407v22017Unsupervised Image Translation using Adversarial Networks for Improved Plant Disease Recognition
Haseeb Nazki, Sook Yoon, Alvaro Fuentes +1
cs.CVcs.LGeess.IVarXiv:1909.11915v12019The Nearest Target Is the Wrong One: Target Separation in Arc2Face Identity Unlearning
Zeynel Tok
cs.CVarXiv:2608.30087v12026SynCrash: A Multi-Stage Pipeline for Zero-Shot Accident Detection and Localization in Traffic Surveillance Video
Arkya Jyoti Bagchi, Ritul Jangir, Varun Raskar
cs.CVcs.AIarXiv:2608.29759v12026CenterCLIP: Token Clustering for Efficient Text-Video Retrieval
Shuai Zhao, Linchao Zhu, Xiaohan Wang +1
cs.CVcs.IRarXiv:2205.00823v12022A survey of face recognition techniques under occlusion
Dan Zeng, Raymond Veldhuis, Luuk Spreeuwers
cs.CVarXiv:2006.11366v12020Face Image Quality Assessment: A Literature Survey
Torsten Schlett, Christian Rathgeb, Olaf Henniger +3
cs.CVarXiv:2009.01103v32020DriveVA: Video Action Models are Zero-Shot Drivers
Mengmeng Liu, Diankun Zhang, Jiuming Liu +7
cs.CVcs.ROarXiv:2604.04198v22026Detection and Localization of Image Forgeries using Resampling Features and Deep Learning
Jason Bunk, Jawadul H. Bappy, Tajuddin Manhar Mohammed +6
cs.CVarXiv:1707.00433v12017Improved Mixed-Example Data Augmentation
Cecilia Summers, Michael J. Dinneen
cs.CVcs.LGarXiv:1805.11272v42018VideoMemory: Toward Consistent Video Generation via Memory Integration
Jinsong Zhou, Yihua Du, Xinli Xu +7
cs.CVarXiv:2601.03655v12026Complete Dictionary Recovery over the Sphere I: Overview and the Geometric Picture
Ju Sun, Qing Qu, John Wright
cs.ITcs.CVmath.OCarXiv:1511.03607v32015VehicleNet: Learning Robust Visual Representation for Vehicle Re-identification
Zhedong Zheng, Tao Ruan, Yunchao Wei +2
cs.CVarXiv:2004.06305v22020Towards Best Practice in Explaining Neural Network Decisions with LRP
Maximilian Kohlbrenner, Alexander Bauer, Shinichi Nakajima +3
cs.LGcs.CVstat.MLarXiv:1910.09840v32019Defending Wearable VLMs Against Private Attribute Inference
Zhimin Li, Pan Wang, Jingxian Chen +4
cs.CVcs.AIarXiv:2608.28691v12026Multi-source Domain Adaptation for Semantic Segmentation
Sicheng Zhao, Bo Li, Xiangyu Yue +5
cs.CVcs.LGeess.IVarXiv:1910.12181v12019SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding
Zhenyu Yang, Yuhang Hu, Zemin Du +6
cs.CVarXiv:2502.10810v22025Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
Wenqi Zhang, Mengna Wang, Gangao Liu +10
cs.CLcs.CVarXiv:2503.21696v22025A Graph-CNN for 3D Point Cloud Classification
Yingxue Zhang, Michael Rabbat
cs.CVcs.LGstat.MLarXiv:1812.01711v12018DSNet: Automatic Dermoscopic Skin Lesion Segmentation
Md. Kamrul Hasan, Lavsen Dahal, Prasad N. Samarakoon +2
eess.IVcs.CVarXiv:1907.04305v22019Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World
Yuzhi Huang, Kairun Wen, Rongxin Gao +14
cs.CVarXiv:2603.12746v12026Can Language Models Encode Perceptual Structure Without Grounding? A Case Study in Color
Mostafa Abdou, Artur Kulmizev, Daniel Hershcovich +3
cs.CVcs.CLarXiv:2109.06129v22021Real-time Multi-Class Helmet Violation Detection Using Few-Shot Data Sampling Technique and YOLOv8
Armstrong Aboah, Bin Wang, Ulas Bagci +1
cs.CVarXiv:2304.08256v12023End-to-End Learning of Driving Models with Surround-View Cameras and Route Planners
Simon Hecker, Dengxin Dai, Luc Van Gool
cs.CVarXiv:1803.10158v22018MeanFuser: Fast One-Step Multi-Modal Trajectory Generation and Adaptive Reconstruction via MeanFlow for End-to-End Autonomous Driving
Junli Wang, Yinan Zheng, Xueyi Liu +9
cs.CVcs.ROarXiv:2602.20060v22026Faster Inference of Flow-Based Generative Models via Improved Data-Noise Coupling
Aram Davtyan, Leello Tadesse Dadi, Volkan Cevher +1
cs.LGcs.CVarXiv:2603.15279v12026Stochastic Liquid Deformation Fields: An SDE Generalisation of Closed-Form Continuous-Time Cells for Dynamic 3D Gaussian Splatting
Mingzhao Li, Arghya Pal
cs.CVarXiv:2608.28702v12026Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video
Zeren Jiang, Chuanxia Zheng, Iro Laina +2
cs.CVarXiv:2601.05251v12026Motion Prompting: Controlling Video Generation with Motion Trajectories
Daniel Geng, Charles Herrmann, Junhwa Hur +11
cs.CVarXiv:2412.02700v22024Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs
Xiang Fang, Wanlong Fang, Changshuo Wang +4
cs.CVarXiv:2605.27894v12026Text-Driven Artistic Staging: Pose, Lighting, and Camera References from Paintings
Yunge Wen
cs.CVcs.AIarXiv:2608.28823v12026Post-Training Piecewise Linear Quantization for Deep Neural Networks
Jun Fang, Ali Shafiee, Hamzah Abdel-Aziz +3
cs.CVcs.LGarXiv:2002.00104v22020Exploring scalable medical image encoders beyond text supervision
Fernando Pérez-García, Harshita Sharma, Sam Bond-Taylor +12
cs.CVarXiv:2401.10815v32024A Comprehensive Study of Knowledge Editing for Large Language Models
Ningyu Zhang, Yunzhi Yao, Bozhong Tian +19
cs.CLcs.AIcs.CVarXiv:2401.01286v52024Distributed Semantic Segmentation With Improved Rate-Distortion Trade-Off
Danish Nazir, Timo Bartels, Thorsten Bagdonat +1
cs.CVcs.LGarXiv:2608.28684v12026SemMAE: Semantic-Guided Masking for Learning Masked Autoencoders
Gang Li, Heliang Zheng, Daqing Liu +3
cs.CVarXiv:2206.10207v32022Revealing the Dark Secrets of Masked Image Modeling
Zhenda Xie, Zigang Geng, Jingcheng Hu +3
cs.CVcs.AIcs.LGarXiv:2205.13543v22022Gate-Shift Networks for Video Action Recognition
Swathikiran Sudhakaran, Sergio Escalera, Oswald Lanz
cs.CVarXiv:1912.00381v22019Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety
Xingjun Ma, Yifeng Gao, Yixu Wang +45
cs.CRcs.AIcs.CLarXiv:2502.05206v62025SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Estimation
Jiehong Lin, Lihua Liu, Dekun Lu +1
cs.CVarXiv:2311.15707v22023DisCo: Disentangled Control for Realistic Human Dance Generation
Tan Wang, Linjie Li, Kevin Lin +6
cs.CVcs.AIarXiv:2307.00040v32023Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation
Zaijing Li, Bing Hu, Rui Shao +5
cs.ROcs.AIcs.CVarXiv:2602.20200v22026