Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

7,681 to 7,740 of 18,855

  1. DiffuMask: Synthesizing Images with Pixel-level Annotations for Semantic Segmentation Using Diffusion Models

    Weijia Wu, Yuzhong Zhao, Mike Zheng Shou +2

    cs.CVarXiv:2303.11681v42023
  2. S4NN: temporal backpropagation for spiking neural networks with one spike per neuron

    Saeed Reza Kheradpisheh, Timothée Masquelier

    cs.NEcs.CVcs.LGarXiv:1910.09495v42019
  3. Patch Slimming for Efficient Vision Transformers

    Yehui Tang, Kai Han, Yunhe Wang +4

    cs.CVcs.LGarXiv:2106.02852v22021
  4. Semantic Hierarchy Emerges in Deep Generative Representations for Scene Synthesis

    Ceyuan Yang, Yujun Shen, Bolei Zhou

    cs.CVcs.GRcs.LGarXiv:1911.09267v32019
  5. Learning Image Matching by Simply Watching Video

    Gucan Long, Laurent Kneip, Jose M. Alvarez +1

    cs.CVarXiv:1603.06041v22016
  6. MV-dVRK: A Multi-Viewpoint Benchmark for Spatial Surgical Perception

    Guido Caccianiga, Sergey Prokudin, Yutong Chen +9

    cs.CVcs.ROarXiv:2609.02717v12026
  7. Camera Relocalization by Computing Pairwise Relative Poses Using Convolutional Neural Network

    Zakaria Laskar, Iaroslav Melekhov, Surya Kalia +1

    cs.CVarXiv:1707.09733v22017
  8. Generating Useful Accident-Prone Driving Scenarios via a Learned Traffic Prior

    Davis Rempe, Jonah Philion, Leonidas J. Guibas +2

    cs.CVcs.LGcs.ROarXiv:2112.05077v22021
  9. UnCapsTSR: An Unsupervised Transformer-based Image Super-Resolution Approach for Capsule Endoscopy Images

    Anjali Sarvaiya, Shubh Kawa, Lalit Agrawal +3

    cs.CVarXiv:2609.02476v12026
  10. Semi-Supervised Semantic Segmentation via Adaptive Equalization Learning

    Hanzhe Hu, Fangyun Wei, Han Hu +3

    cs.CVarXiv:2110.05474v12021
  11. PRISM: An Agentic Multi-Model Architecture for Proactive Safety in Autonomous Transportation Systems

    Joyjit Roy, Samaresh Kumar Singh, Sushanta Das

    cs.MAcs.CVcs.ETarXiv:2609.01623v12026
  12. The Devil of Face Recognition is in the Noise

    Fei Wang, Liren Chen, Cheng Li +4

    cs.CVarXiv:1807.11649v12018
  13. Towards Long-Form Video Understanding

    Chao-Yuan Wu, Philipp Krähenbühl

    cs.CVarXiv:2106.11310v12021
  14. Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models

    Junyu Chen, Han Cai, Junsong Chen +6

    cs.CVcs.AIarXiv:2410.10733v82024
  15. Clockwork Convnets for Video Semantic Segmentation

    Evan Shelhamer, Kate Rakelly, Judy Hoffman +1

    cs.CVarXiv:1608.03609v12016
  16. Person Search via A Mask-Guided Two-Stream CNN Model

    Di Chen, Shanshan Zhang, Wanli Ouyang +2

    cs.CVarXiv:1807.08107v12018
  17. Person-in-WiFi: Fine-grained Person Perception using WiFi

    Fei Wang, Sanping Zhou, Stanislav Panev +2

    cs.CVarXiv:1904.00276v12019
  18. Highly Efficient Salient Object Detection with 100K Parameters

    Shang-Hua Gao, Yong-Qiang Tan, Ming-Ming Cheng +3

    cs.CVarXiv:2003.05643v22020
  19. Benchmarking RAW and RGB Restoration in Image Signal Processors

    Zihao Lu, Radu Timofte, Marcos V. Conde

    cs.CVarXiv:2609.02831v12026
  20. Clustering on Multi-Layer Graphs via Subspace Analysis on Grassmann Manifolds

    Xiaowen Dong, Pascal Frossard, Pierre Vandergheynst +1

    cs.LGcs.CVcs.SIarXiv:1303.2221v12013
  21. Balancing Frequencies and Pixels in Flow Matching

    Lucas Degeorge, Paul Couairon, Arijit Ghosh +3

    cs.CVarXiv:2609.02748v12026
  22. Learning Spatial Attention for Face Super-Resolution

    Chaofeng Chen, Dihong Gong, Hao Wang +2

    cs.CVarXiv:2012.01211v22020
  23. Dual Contrastive Learning for Unsupervised Image-to-Image Translation

    Junlin Han, Mehrdad Shoeiby, Lars Petersson +1

    cs.CVeess.IVarXiv:2104.07689v12021
  24. SEA-RAFT: Simple, Efficient, Accurate RAFT for Optical Flow

    Yihan Wang, Lahav Lipson, Jia Deng

    cs.CVarXiv:2405.14793v12024
  25. Confidence Propagation through CNNs for Guided Sparse Depth Regression

    Abdelrahman Eldesokey, Michael Felsberg, Fahad Shahbaz Khan

    cs.CVarXiv:1811.01791v22018
  26. HDR-GAN: HDR Image Reconstruction from Multi-Exposed LDR Images with Large Motions

    Yuzhen Niu, Jianbin Wu, Wenxi Liu +2

    eess.IVcs.CVarXiv:2007.01628v12020
  27. FoundYou: A Unified Model for Personalized Segmentation and Retrieval

    Gabriele Trivigno, Marcos Alfaro, Claudia Cuttano +3

    cs.CVarXiv:2608.29917v12026
  28. Parsing-based View-aware Embedding Network for Vehicle Re-Identification

    Dechao Meng, Liang Li, Xuejing Liu +6

    cs.CVarXiv:2004.05021v12020
  29. A Complete Survey on Generative AI (AIGC): Is ChatGPT from GPT-4 to GPT-5 All You Need?

    Chaoning Zhang, Chenshuang Zhang, Sheng Zheng +14

    cs.AIcs.CVcs.LGarXiv:2303.11717v12023
  30. Land Cover Classification via Multi-temporal Spatial Data by Recurrent Neural Networks

    Dino Ienco, Raffaele Gaetano, Claire Dupaquier +1

    cs.CVcs.LGarXiv:1704.04055v12017
  31. Foundation and Multimodal Large Language Models for Face Presentation and Morph Attack Detection

    Hatef Otroshi Shahreza, Asif Hussain Khan, Peter Lorenz +2

    cs.CVarXiv:2608.29802v12026
  32. Deeply-Supervised CNN for Prostate Segmentation

    Qikui Zhu, Bo Du, Baris Turkbey +2

    cs.CVarXiv:1703.07523v32017
  33. Learning to Parse Wireframes in Images of Man-Made Environments

    Kun Huang, Yifan Wang, Zihan Zhou +3

    cs.CVarXiv:2007.07527v12020
  34. Remote Sensing Image Super-resolution and Object Detection: Benchmark and State of the Art

    Yi Wang, Syed Muhammad Arsalan Bashir, Mahrukh Khan +5

    cs.CVcs.AIcs.LGarXiv:2111.03260v12021
  35. Deep Residual Learning for Compressed Sensing CT Reconstruction via Persistent Homology Analysis

    Yo Seob Han, Jaejun Yoo, Jong Chul Ye

    cs.CVarXiv:1611.06391v22016
  36. Image Data Augmentation Approaches: A Comprehensive Survey and Future directions

    Teerath Kumar, Alessandra Mileo, Rob Brennan +1

    cs.CVcs.AIcs.LGarXiv:2301.02830v42023
  37. Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models

    Aditi Sarker, Nazreen Shah, Rafi Ibn Sultan +3

    cs.CVcs.AIcs.LGarXiv:2608.29996v12026
  38. Pretraining is All You Need for Image-to-Image Translation

    Tengfei Wang, Ting Zhang, Bo Zhang +4

    cs.CVarXiv:2205.12952v12022
  39. Multi-view Consistency as Supervisory Signal for Learning Shape and Pose Prediction

    Shubham Tulsiani, Alexei A. Efros, Jitendra Malik

    cs.CVarXiv:1801.03910v22018
  40. WonderWorld: Interactive 3D Scene Generation from a Single Image

    Hong-Xing Yu, Haoyi Duan, Charles Herrmann +2

    cs.CVcs.GRarXiv:2406.09394v42024
  41. A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss

    Suryaansh Jain, Rahasya Barkur, Vishal G +8

    cs.CVarXiv:2609.00591v12026
  42. Sanity Checks for Saliency Metrics

    Richard Tomsett, Dan Harborne, Supriyo Chakraborty +2

    cs.LGcs.CVeess.IVarXiv:1912.01451v12019
  43. TPCN: Temporal Point Cloud Networks for Motion Forecasting

    Maosheng Ye, Tongyi Cao, Qifeng Chen

    cs.CVcs.ROarXiv:2103.03067v12021
  44. RADNET: Radiologist Level Accuracy using Deep Learning for HEMORRHAGE detection in CT Scans

    Monika Grewal, Muktabh Mayank Srivastava, Pulkit Kumar +1

    cs.CVstat.MLarXiv:1710.04934v22017
  45. Benchmarking the Robustness of Semantic Segmentation Models

    Christoph Kamann, Carsten Rother

    cs.CVeess.IVarXiv:1908.05005v32019
  46. Robust Multimodal Brain Tumor Segmentation via Feature Disentanglement and Gated Fusion

    Cheng Chen, Qi Dou, Yueming Jin +3

    cs.CVarXiv:2002.09708v12020
  47. NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference

    Aurélien Lac, Tony Wu

    cs.IRcs.AIcs.CVarXiv:2609.01657v12026
  48. Feature Aggregation and Propagation Network for Camouflaged Object Detection

    Tao Zhou, Yi Zhou, Chen Gong +2

    cs.CVarXiv:2212.00990v12022
  49. FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos

    Maya Moriya, Sigal Raab, Yael Vinker +1

    cs.CVcs.AIcs.GRarXiv:2609.00377v12026
  50. Toolflows for Mapping Convolutional Neural Networks on FPGAs: A Survey and Future Directions

    Stylianos I. Venieris, Alexandros Kouris, Christos-Savvas Bouganis

    cs.CVcs.ARcs.LGarXiv:1803.05900v12018
  51. CAS-CNN: A Deep Convolutional Neural Network for Image Compression Artifact Suppression

    Lukas Cavigelli, Pascal Hager, Luca Benini

    cs.CVcs.AIcs.GRarXiv:1611.07233v12016
  52. Text Flow: A Unified Text Detection System in Natural Scene Images

    Shangxuan Tian, Yifeng Pan, Chang Huang +3

    cs.CVarXiv:1604.06877v12016
  53. Kirin: Animal Motion Generation from In-the-Wild Video

    Brian Nlong Zhao, Zhuoyang Pan, James M. Rehg +2

    cs.CVarXiv:2609.01823v12026
  54. Hybrid Models for Open Set Recognition

    Hongjie Zhang, Ang Li, Jie Guo +1

    cs.CVarXiv:2003.12506v22020
  55. Falling Things: A Synthetic Dataset for 3D Object Detection and Pose Estimation

    Jonathan Tremblay, Thang To, Stan Birchfield

    cs.CVarXiv:1804.06534v22018
  56. Uncertainty-guided Continual Learning with Bayesian Neural Networks

    Sayna Ebrahimi, Mohamed Elhoseiny, Trevor Darrell +1

    cs.LGcs.AIcs.CVarXiv:1906.02425v22019
  57. Towards the Limit of Network Quantization

    Yoojin Choi, Mostafa El-Khamy, Jungwon Lee

    cs.CVcs.LGcs.NEarXiv:1612.01543v22016
  58. RedCaps: web-curated image-text data created by the people, for the people

    Karan Desai, Gaurav Kaul, Zubin Aysola +1

    cs.CVcs.CLarXiv:2111.11431v12021
  59. Group Sparsity: The Hinge Between Filter Pruning and Decomposition for Network Compression

    Yawei Li, Shuhang Gu, Christoph Mayer +2

    cs.CVarXiv:2003.08935v12020
  60. ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes

    Mingda Lin, Weijie Wang, Zeyu Zhang +7

    cs.CVarXiv:2609.01740v12026