Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

16,981 to 17,040 of 18,848

  1. Quantizing deep convolutional networks for efficient inference: A whitepaper

    Raghuraman Krishnamoorthi

    cs.LGcs.CVstat.MLarXiv:1806.08342v12018
  2. MaskGAN: Towards Diverse and Interactive Facial Image Manipulation

    Cheng-Han Lee, Ziwei Liu, Lingyun Wu +1

    cs.CVcs.GRcs.LGarXiv:1907.11922v22019
  3. MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

    Kunchang Li, Yali Wang, Yinan He +9

    cs.CVarXiv:2311.17005v42023
  4. VizWiz Grand Challenge: Answering Visual Questions from Blind People

    Danna Gurari, Qing Li, Abigale J. Stangl +5

    cs.CVcs.CLcs.HCarXiv:1802.08218v42018
  5. HyperFace: A Deep Multi-task Learning Framework for Face Detection, Landmark Localization, Pose Estimation, and Gender Recognition

    Rajeev Ranjan, Vishal M. Patel, Rama Chellappa

    cs.CVarXiv:1603.01249v32016
  6. Deep Joint Rain Detection and Removal from a Single Image

    Wenhan Yang, Robby T. Tan, Jiashi Feng +3

    cs.CVarXiv:1609.07769v32016
  7. Salient Object Detection: A Discriminative Regional Feature Integration Approach

    Huaizu Jiang, Zejian Yuan, Ming-Ming Cheng +3

    cs.CVarXiv:1410.5926v12014
  8. Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling

    Yuan Wang, Ouxiang Li, Yulong Xu +8

    cs.CVarXiv:2605.05922v22026
  9. Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance

    Ziyun Zeng, Yiqi Lin, Guoqiang Liang +1

    cs.CVcs.AIarXiv:2605.06535v12026
  10. Part-based R-CNNs for Fine-grained Category Detection

    Ning Zhang, Jeff Donahue, Ross Girshick +1

    cs.CVarXiv:1407.3867v12014
  11. EDVR: Video Restoration with Enhanced Deformable Convolutional Networks

    Xintao Wang, Kelvin C. K. Chan, Ke Yu +2

    cs.CVarXiv:1905.02716v12019
  12. 4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding

    Zhangquan Chen, Manyuan Zhang, Xinlei Yu +9

    cs.CVarXiv:2605.05997v22026
  13. Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study

    Hao Dong, Hongzhao Li, Shupan Li +3

    cs.CVcs.AIcs.LGarXiv:2605.06643v12026
  14. BAM: Bottleneck Attention Module

    Jongchan Park, Sanghyun Woo, Joon-Young Lee +1

    cs.CVarXiv:1807.06514v22018
  15. CURL: Contrastive Unsupervised Representations for Reinforcement Learning

    Aravind Srinivas, Michael Laskin, Pieter Abbeel

    cs.LGcs.CVstat.MLarXiv:2004.04136v42020
  16. Age Progression/Regression by Conditional Adversarial Autoencoder

    Zhifei Zhang, Yang Song, Hairong Qi

    cs.CVarXiv:1702.08423v22017
  17. RemoteZero: Geospatial Reasoning with Zero Labels

    Liang Yao, Fan Liu, Shengxiang Xu +4

    cs.CVarXiv:2605.04451v22026
  18. Local Light Field Fusion: Practical View Synthesis with Prescriptive Sampling Guidelines

    Ben Mildenhall, Pratul P. Srinivasan, Rodrigo Ortiz-Cayon +4

    cs.CVcs.GRarXiv:1905.00889v12019
  19. FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation

    Yuanzhi Wang, Xuhua Ren, Jiaxiang Cheng +7

    cs.CVcs.AIarXiv:2605.04702v12026
  20. Adversarial Patch

    Tom B. Brown, Dandelion Mané, Aurko Roy +2

    cs.CVarXiv:1712.09665v22017
  21. MoCoGAN: Decomposing Motion and Content for Video Generation

    Sergey Tulyakov, Ming-Yu Liu, Xiaodong Yang +1

    cs.CVarXiv:1707.04993v22017
  22. GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMs

    Pranav Mantini, Shishir K. Shah

    cs.CVarXiv:2605.06477v12026
  23. SegNeXt: Rethinking Convolutional Attention Design for Semantic Segmentation

    Meng-Hao Guo, Cheng-Ze Lu, Qibin Hou +3

    cs.CVarXiv:2209.08575v12022
  24. Object Detectors Emerge in Deep Scene CNNs

    Bolei Zhou, Aditya Khosla, Agata Lapedriza +2

    cs.CVcs.NEarXiv:1412.6856v22014
  25. Trainable Nonlinear Reaction Diffusion: A Flexible Framework for Fast and Effective Image Restoration

    Yunjin Chen, Thomas Pock

    cs.CVarXiv:1508.02848v22015
  26. UNetFormer: A UNet-like Transformer for Efficient Semantic Segmentation of Remote Sensing Urban Scene Imagery

    Libo Wang, Rui Li, Ce Zhang +4

    cs.CVarXiv:2109.08937v42021
  27. PlenOctrees for Real-time Rendering of Neural Radiance Fields

    Alex Yu, Ruilong Li, Matthew Tancik +3

    cs.CVcs.GRarXiv:2103.14024v22021
  28. The Secrets of Salient Object Segmentation

    Yin Li, Xiaodi Hou, Christof Koch +2

    cs.CVarXiv:1406.2807v22014
  29. Fast Online Object Tracking and Segmentation: A Unifying Approach

    Qiang Wang, Li Zhang, Luca Bertinetto +2

    cs.CVarXiv:1812.05050v22018
  30. Empirical Evidence for Simply Connected Decision Regions in Image Classifiers

    Arjhun Swaminathan, Mete Akgün

    cs.CVcs.LGarXiv:2605.06380v12026
  31. Uncovering Entity Identity Confusion in Multimodal Knowledge Editing

    Shu Wu, Xiaotian Ye, Xinyu Mou +3

    cs.CLcs.CVarXiv:2605.06096v12026
  32. Visual Saliency Based on Multiscale Deep Features

    Guanbin Li, Yizhou Yu

    cs.CVarXiv:1503.08663v32015
  33. A Survey on Object Detection in Optical Remote Sensing Images

    Gong Cheng, Junwei Han

    cs.CVarXiv:1603.06201v22016
  34. Twins: Revisiting the Design of Spatial Attention in Vision Transformers

    Xiangxiang Chu, Zhi Tian, Yuqing Wang +5

    cs.CVcs.AIcs.LGarXiv:2104.13840v42021
  35. Learning Discriminative Model Prediction for Tracking

    Goutam Bhat, Martin Danelljan, Luc Van Gool +1

    cs.CVarXiv:1904.07220v22019
  36. ATOM: Accurate Tracking by Overlap Maximization

    Martin Danelljan, Goutam Bhat, Fahad Shahbaz Khan +1

    cs.CVarXiv:1811.07628v22018
  37. AtlasNet: A Papier-Mâché Approach to Learning 3D Surface Generation

    Thibault Groueix, Matthew Fisher, Vladimir G. Kim +2

    cs.CVarXiv:1802.05384v32018
  38. FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling

    Bowen Zhang, Yidong Wang, Wenxin Hou +4

    cs.LGcs.CVarXiv:2110.08263v32021
  39. End-to-End Incremental Learning

    Francisco M. Castro, Manuel J. Marín-Jiménez, Nicolás Guil +2

    cs.CVarXiv:1807.09536v22018
  40. Simultaneous Detection and Segmentation

    Bharath Hariharan, Pablo Arbeláez, Ross Girshick +1

    cs.CVarXiv:1407.1808v12014
  41. Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation

    Golnaz Ghiasi, Yin Cui, Aravind Srinivas +5

    cs.CVarXiv:2012.07177v22020
  42. Relation Networks for Object Detection

    Han Hu, Jiayuan Gu, Zheng Zhang +2

    cs.CVarXiv:1711.11575v22017
  43. From Captions to Visual Concepts and Back

    Hao Fang, Saurabh Gupta, Forrest Iandola +9

    cs.CVcs.CLarXiv:1411.4952v32014
  44. Learning to Prompt for Continual Learning

    Zifeng Wang, Zizhao Zhang, Chen-Yu Lee +7

    cs.LGcs.CVarXiv:2112.08654v22021
  45. Learning Fine-grained Image Similarity with Deep Ranking

    Jiang Wang, Yang song, Thomas Leung +5

    cs.CVarXiv:1404.4661v12014
  46. Going deeper with Image Transformers

    Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles +2

    cs.CVarXiv:2103.17239v22021
  47. Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

    Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo +1

    cs.CVcs.AIcs.LGarXiv:2104.13921v32021
  48. Stand-Alone Self-Attention in Vision Models

    Prajit Ramachandran, Niki Parmar, Ashish Vaswani +3

    cs.CVarXiv:1906.05909v12019
  49. LIFT: Learned Invariant Feature Transform

    Kwang Moo Yi, Eduard Trulls, Vincent Lepetit +1

    cs.CVarXiv:1603.09114v22016
  50. Denoising Diffusion Restoration Models

    Bahjat Kawar, Michael Elad, Stefano Ermon +1

    eess.IVcs.CVcs.LGarXiv:2201.11793v32022
  51. Delta-Adapter: Scalable Exemplar-Based Image Editing with Single-Pair Supervision

    Jiacheng Chen, Songze Li, Han Fu +5

    cs.CVarXiv:2605.07940v12026
  52. ResUNet++: An Advanced Architecture for Medical Image Segmentation

    Debesh Jha, Pia H. Smedsrud, Michael A. Riegler +4

    eess.IVcs.CVarXiv:1911.07067v12019
  53. MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

    Weihao Yu, Zhengyuan Yang, Linjie Li +5

    cs.AIcs.CLcs.CVarXiv:2308.02490v42023
  54. Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs

    Hao Wang, Yiqun Sun, Pengfei Wei +2

    cs.CVcs.AIcs.CLarXiv:2605.07447v12026
  55. Transformer Tracking

    Xin Chen, Bin Yan, Jiawen Zhu +3

    cs.CVarXiv:2103.15436v12021
  56. BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning

    Shaokai Ye, Vasileios Saveris, Yihao Qian +3

    cs.CVcs.AIarXiv:2605.07394v12026
  57. Dynamic Edge-Conditioned Filters in Convolutional Neural Networks on Graphs

    Martin Simonovsky, Nikos Komodakis

    cs.CVcs.LGcs.NEarXiv:1704.02901v32017
  58. Context Encoding for Semantic Segmentation

    Hang Zhang, Kristin Dana, Jianping Shi +4

    cs.CVarXiv:1803.08904v12018
  59. Validation, comparison, and combination of algorithms for automatic detection of pulmonary nodules in computed tomography images: the LUNA16 challenge

    Arnaud Arindra Adiyoso Setio, Alberto Traverso, Thomas de Bel +29

    cs.CVarXiv:1612.08012v42016
  60. Implicit Preference Alignment for Human Image Animation

    Yuanzhi Wang, Xuhua Ren, Jiaxiang Cheng +5

    cs.CVcs.AIarXiv:2605.07545v12026