Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,781 to 3,840 of 18,852

  1. Learning to Match Features with Seeded Graph Matching Network

    Hongkai Chen, Zixin Luo, Jiahui Zhang +5

    cs.CVarXiv:2108.08771v12021
  2. Learning What to Learn for Video Object Segmentation

    Goutam Bhat, Felix Järemo Lawin, Martin Danelljan +4

    cs.CVarXiv:2003.11540v22020
  3. Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning

    Yihong Huang, Fei Ma, Yihua Shao +4

    cs.CVcs.AIcs.CLarXiv:2602.02951v12026
  4. Matcher: Segment Anything with One Shot Using All-Purpose Feature Matching

    Yang Liu, Muzhi Zhu, Hengtao Li +3

    cs.CVarXiv:2305.13310v22023
  5. Learning to Ground Before Reading: Unified PCB Engineering Drawing Parsing with Compact Vision-Language Models

    Jinghao Liu, Xingrun Liu, Gengchen Sun +3

    cs.CVcs.MMarXiv:2608.29268v12026
  6. Oriented Edge Forests for Boundary Detection

    Sam Hallman, Charless C. Fowlkes

    cs.CVarXiv:1412.4181v22014
  7. Deep Semantic Segmentation for Automated Driving: Taxonomy, Roadmap and Challenges

    Mennatullah Siam, Sara Elkerdawy, Martin Jagersand +1

    stat.MLcs.CVarXiv:1707.02432v22017
  8. Extended depth-of-field in holographic image reconstruction using deep learning based auto-focusing and phase-recovery

    Yichen Wu, Yair Rivenson, Yibo Zhang +4

    cs.CVcs.LGphysics.opticsarXiv:1803.08138v12018
  9. Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation

    Zengqun Zhao, Yanzuo Lu, Ziquan Liu +3

    cs.CVarXiv:2603.21366v22026
  10. High throughput quantitative metallography for complex microstructures using deep learning: A case study in ultrahigh carbon steel

    Brian L. DeCost, Bo Lei, Toby Francis +1

    cs.CVarXiv:1805.08693v22018
  11. Action Transformer: A Self-Attention Model for Short-Time Pose-Based Human Action Recognition

    Vittorio Mazzia, Simone Angarano, Francesco Salvetti +2

    cs.CVcs.LGarXiv:2107.00606v62021
  12. Temperature-Adaptive Transformed Teacher Matching

    Hiroaki Aizawa, Yoshikazu Hayashi

    cs.LGcs.CVarXiv:2608.29099v12026
  13. BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks

    Yixiang Chen, Peiyan Li, Jiabing Yang +8

    cs.ROcs.CVarXiv:2602.03793v12026
  14. Enhancing Low-Cost Video Editing with Lightweight Adaptors and Temporal-Aware Inversion

    Yangfan He, Sida Li, Jianhui Wang +11

    cs.CVarXiv:2501.04606v42025
  15. Growing a Stand, Not a Tree: Joint Canopy Generation Reproduces Crown Shyness

    Guang Yang, Fengchen Liu

    cs.CVarXiv:2608.28692v12026
  16. FineCIR: Explicit Parsing of Fine-Grained Modification Semantics for Composed Image Retrieval

    Zixu Li, Zhiheng Fu, Yupeng Hu +3

    cs.CVcs.AIarXiv:2503.21309v12025
  17. Continuous Dice Coefficient: a Method for Evaluating Probabilistic Segmentations

    Reuben R Shamir, Yuval Duchin, Jinyoung Kim +2

    cs.CVeess.IVarXiv:1906.11031v12019
  18. Median K-flats for hybrid linear modeling with many outliers

    Teng Zhang, Arthur Szlam, Gilad Lerman

    cs.CVcs.LGarXiv:0909.3123v12009
  19. Knowledge-Guided Deep Fractal Neural Networks for Human Pose Estimation

    Guanghan Ning, Zhi Zhang, Zhihai He

    cs.CVarXiv:1705.02407v22017
  20. Unsupervised Image Translation using Adversarial Networks for Improved Plant Disease Recognition

    Haseeb Nazki, Sook Yoon, Alvaro Fuentes +1

    cs.CVcs.LGeess.IVarXiv:1909.11915v12019
  21. The Nearest Target Is the Wrong One: Target Separation in Arc2Face Identity Unlearning

    Zeynel Tok

    cs.CVarXiv:2608.30087v12026
  22. SynCrash: A Multi-Stage Pipeline for Zero-Shot Accident Detection and Localization in Traffic Surveillance Video

    Arkya Jyoti Bagchi, Ritul Jangir, Varun Raskar

    cs.CVcs.AIarXiv:2608.29759v12026
  23. CenterCLIP: Token Clustering for Efficient Text-Video Retrieval

    Shuai Zhao, Linchao Zhu, Xiaohan Wang +1

    cs.CVcs.IRarXiv:2205.00823v12022
  24. A survey of face recognition techniques under occlusion

    Dan Zeng, Raymond Veldhuis, Luuk Spreeuwers

    cs.CVarXiv:2006.11366v12020
  25. Face Image Quality Assessment: A Literature Survey

    Torsten Schlett, Christian Rathgeb, Olaf Henniger +3

    cs.CVarXiv:2009.01103v32020
  26. DriveVA: Video Action Models are Zero-Shot Drivers

    Mengmeng Liu, Diankun Zhang, Jiuming Liu +7

    cs.CVcs.ROarXiv:2604.04198v22026
  27. Detection and Localization of Image Forgeries using Resampling Features and Deep Learning

    Jason Bunk, Jawadul H. Bappy, Tajuddin Manhar Mohammed +6

    cs.CVarXiv:1707.00433v12017
  28. Improved Mixed-Example Data Augmentation

    Cecilia Summers, Michael J. Dinneen

    cs.CVcs.LGarXiv:1805.11272v42018
  29. VideoMemory: Toward Consistent Video Generation via Memory Integration

    Jinsong Zhou, Yihua Du, Xinli Xu +7

    cs.CVarXiv:2601.03655v12026
  30. Complete Dictionary Recovery over the Sphere I: Overview and the Geometric Picture

    Ju Sun, Qing Qu, John Wright

    cs.ITcs.CVmath.OCarXiv:1511.03607v32015
  31. VehicleNet: Learning Robust Visual Representation for Vehicle Re-identification

    Zhedong Zheng, Tao Ruan, Yunchao Wei +2

    cs.CVarXiv:2004.06305v22020
  32. Towards Best Practice in Explaining Neural Network Decisions with LRP

    Maximilian Kohlbrenner, Alexander Bauer, Shinichi Nakajima +3

    cs.LGcs.CVstat.MLarXiv:1910.09840v32019
  33. Defending Wearable VLMs Against Private Attribute Inference

    Zhimin Li, Pan Wang, Jingxian Chen +4

    cs.CVcs.AIarXiv:2608.28691v12026
  34. Multi-source Domain Adaptation for Semantic Segmentation

    Sicheng Zhao, Bo Li, Xiangyu Yue +5

    cs.CVcs.LGeess.IVarXiv:1910.12181v12019
  35. SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding

    Zhenyu Yang, Yuhang Hu, Zemin Du +6

    cs.CVarXiv:2502.10810v22025
  36. Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks

    Wenqi Zhang, Mengna Wang, Gangao Liu +10

    cs.CLcs.CVarXiv:2503.21696v22025
  37. A Graph-CNN for 3D Point Cloud Classification

    Yingxue Zhang, Michael Rabbat

    cs.CVcs.LGstat.MLarXiv:1812.01711v12018
  38. DSNet: Automatic Dermoscopic Skin Lesion Segmentation

    Md. Kamrul Hasan, Lavsen Dahal, Prasad N. Samarakoon +2

    eess.IVcs.CVarXiv:1907.04305v22019
  39. Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World

    Yuzhi Huang, Kairun Wen, Rongxin Gao +14

    cs.CVarXiv:2603.12746v12026
  40. Can Language Models Encode Perceptual Structure Without Grounding? A Case Study in Color

    Mostafa Abdou, Artur Kulmizev, Daniel Hershcovich +3

    cs.CVcs.CLarXiv:2109.06129v22021
  41. Real-time Multi-Class Helmet Violation Detection Using Few-Shot Data Sampling Technique and YOLOv8

    Armstrong Aboah, Bin Wang, Ulas Bagci +1

    cs.CVarXiv:2304.08256v12023
  42. End-to-End Learning of Driving Models with Surround-View Cameras and Route Planners

    Simon Hecker, Dengxin Dai, Luc Van Gool

    cs.CVarXiv:1803.10158v22018
  43. MeanFuser: Fast One-Step Multi-Modal Trajectory Generation and Adaptive Reconstruction via MeanFlow for End-to-End Autonomous Driving

    Junli Wang, Yinan Zheng, Xueyi Liu +9

    cs.CVcs.ROarXiv:2602.20060v22026
  44. Faster Inference of Flow-Based Generative Models via Improved Data-Noise Coupling

    Aram Davtyan, Leello Tadesse Dadi, Volkan Cevher +1

    cs.LGcs.CVarXiv:2603.15279v12026
  45. Stochastic Liquid Deformation Fields: An SDE Generalisation of Closed-Form Continuous-Time Cells for Dynamic 3D Gaussian Splatting

    Mingzhao Li, Arghya Pal

    cs.CVarXiv:2608.28702v12026
  46. Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video

    Zeren Jiang, Chuanxia Zheng, Iro Laina +2

    cs.CVarXiv:2601.05251v12026
  47. Motion Prompting: Controlling Video Generation with Motion Trajectories

    Daniel Geng, Charles Herrmann, Junhwa Hur +11

    cs.CVarXiv:2412.02700v22024
  48. Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs

    Xiang Fang, Wanlong Fang, Changshuo Wang +4

    cs.CVarXiv:2605.27894v12026
  49. Text-Driven Artistic Staging: Pose, Lighting, and Camera References from Paintings

    Yunge Wen

    cs.CVcs.AIarXiv:2608.28823v12026
  50. Post-Training Piecewise Linear Quantization for Deep Neural Networks

    Jun Fang, Ali Shafiee, Hamzah Abdel-Aziz +3

    cs.CVcs.LGarXiv:2002.00104v22020
  51. Exploring scalable medical image encoders beyond text supervision

    Fernando Pérez-García, Harshita Sharma, Sam Bond-Taylor +12

    cs.CVarXiv:2401.10815v32024
  52. A Comprehensive Study of Knowledge Editing for Large Language Models

    Ningyu Zhang, Yunzhi Yao, Bozhong Tian +19

    cs.CLcs.AIcs.CVarXiv:2401.01286v52024
  53. Distributed Semantic Segmentation With Improved Rate-Distortion Trade-Off

    Danish Nazir, Timo Bartels, Thorsten Bagdonat +1

    cs.CVcs.LGarXiv:2608.28684v12026
  54. SemMAE: Semantic-Guided Masking for Learning Masked Autoencoders

    Gang Li, Heliang Zheng, Daqing Liu +3

    cs.CVarXiv:2206.10207v32022
  55. Revealing the Dark Secrets of Masked Image Modeling

    Zhenda Xie, Zigang Geng, Jingcheng Hu +3

    cs.CVcs.AIcs.LGarXiv:2205.13543v22022
  56. Gate-Shift Networks for Video Action Recognition

    Swathikiran Sudhakaran, Sergio Escalera, Oswald Lanz

    cs.CVarXiv:1912.00381v22019
  57. Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety

    Xingjun Ma, Yifeng Gao, Yixu Wang +45

    cs.CRcs.AIcs.CLarXiv:2502.05206v62025
  58. SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Estimation

    Jiehong Lin, Lihua Liu, Dekun Lu +1

    cs.CVarXiv:2311.15707v22023
  59. DisCo: Disentangled Control for Realistic Human Dance Generation

    Tan Wang, Linjie Li, Kevin Lin +6

    cs.CVcs.AIarXiv:2307.00040v32023
  60. Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation

    Zaijing Li, Bing Hu, Rui Shao +5

    cs.ROcs.AIcs.CVarXiv:2602.20200v22026