Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
16,981 to 17,040 of 18,848
Quantizing deep convolutional networks for efficient inference: A whitepaper
Raghuraman Krishnamoorthi
cs.LGcs.CVstat.MLarXiv:1806.08342v12018MaskGAN: Towards Diverse and Interactive Facial Image Manipulation
Cheng-Han Lee, Ziwei Liu, Lingyun Wu +1
cs.CVcs.GRcs.LGarXiv:1907.11922v22019MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Kunchang Li, Yali Wang, Yinan He +9
cs.CVarXiv:2311.17005v42023VizWiz Grand Challenge: Answering Visual Questions from Blind People
Danna Gurari, Qing Li, Abigale J. Stangl +5
cs.CVcs.CLcs.HCarXiv:1802.08218v42018HyperFace: A Deep Multi-task Learning Framework for Face Detection, Landmark Localization, Pose Estimation, and Gender Recognition
Rajeev Ranjan, Vishal M. Patel, Rama Chellappa
cs.CVarXiv:1603.01249v32016Deep Joint Rain Detection and Removal from a Single Image
Wenhan Yang, Robby T. Tan, Jiashi Feng +3
cs.CVarXiv:1609.07769v32016Salient Object Detection: A Discriminative Regional Feature Integration Approach
Huaizu Jiang, Zejian Yuan, Ming-Ming Cheng +3
cs.CVarXiv:1410.5926v12014Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling
Yuan Wang, Ouxiang Li, Yulong Xu +8
cs.CVarXiv:2605.05922v22026Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance
Ziyun Zeng, Yiqi Lin, Guoqiang Liang +1
cs.CVcs.AIarXiv:2605.06535v12026Part-based R-CNNs for Fine-grained Category Detection
Ning Zhang, Jeff Donahue, Ross Girshick +1
cs.CVarXiv:1407.3867v12014EDVR: Video Restoration with Enhanced Deformable Convolutional Networks
Xintao Wang, Kelvin C. K. Chan, Ke Yu +2
cs.CVarXiv:1905.02716v120194DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding
Zhangquan Chen, Manyuan Zhang, Xinlei Yu +9
cs.CVarXiv:2605.05997v22026Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study
Hao Dong, Hongzhao Li, Shupan Li +3
cs.CVcs.AIcs.LGarXiv:2605.06643v12026BAM: Bottleneck Attention Module
Jongchan Park, Sanghyun Woo, Joon-Young Lee +1
cs.CVarXiv:1807.06514v22018CURL: Contrastive Unsupervised Representations for Reinforcement Learning
Aravind Srinivas, Michael Laskin, Pieter Abbeel
cs.LGcs.CVstat.MLarXiv:2004.04136v42020Age Progression/Regression by Conditional Adversarial Autoencoder
Zhifei Zhang, Yang Song, Hairong Qi
cs.CVarXiv:1702.08423v22017RemoteZero: Geospatial Reasoning with Zero Labels
Liang Yao, Fan Liu, Shengxiang Xu +4
cs.CVarXiv:2605.04451v22026Local Light Field Fusion: Practical View Synthesis with Prescriptive Sampling Guidelines
Ben Mildenhall, Pratul P. Srinivasan, Rodrigo Ortiz-Cayon +4
cs.CVcs.GRarXiv:1905.00889v12019FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation
Yuanzhi Wang, Xuhua Ren, Jiaxiang Cheng +7
cs.CVcs.AIarXiv:2605.04702v12026Adversarial Patch
Tom B. Brown, Dandelion Mané, Aurko Roy +2
cs.CVarXiv:1712.09665v22017MoCoGAN: Decomposing Motion and Content for Video Generation
Sergey Tulyakov, Ming-Yu Liu, Xiaodong Yang +1
cs.CVarXiv:1707.04993v22017GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMs
Pranav Mantini, Shishir K. Shah
cs.CVarXiv:2605.06477v12026SegNeXt: Rethinking Convolutional Attention Design for Semantic Segmentation
Meng-Hao Guo, Cheng-Ze Lu, Qibin Hou +3
cs.CVarXiv:2209.08575v12022Object Detectors Emerge in Deep Scene CNNs
Bolei Zhou, Aditya Khosla, Agata Lapedriza +2
cs.CVcs.NEarXiv:1412.6856v22014Trainable Nonlinear Reaction Diffusion: A Flexible Framework for Fast and Effective Image Restoration
Yunjin Chen, Thomas Pock
cs.CVarXiv:1508.02848v22015UNetFormer: A UNet-like Transformer for Efficient Semantic Segmentation of Remote Sensing Urban Scene Imagery
Libo Wang, Rui Li, Ce Zhang +4
cs.CVarXiv:2109.08937v42021PlenOctrees for Real-time Rendering of Neural Radiance Fields
Alex Yu, Ruilong Li, Matthew Tancik +3
cs.CVcs.GRarXiv:2103.14024v22021The Secrets of Salient Object Segmentation
Yin Li, Xiaodi Hou, Christof Koch +2
cs.CVarXiv:1406.2807v22014Fast Online Object Tracking and Segmentation: A Unifying Approach
Qiang Wang, Li Zhang, Luca Bertinetto +2
cs.CVarXiv:1812.05050v22018Empirical Evidence for Simply Connected Decision Regions in Image Classifiers
Arjhun Swaminathan, Mete Akgün
cs.CVcs.LGarXiv:2605.06380v12026Uncovering Entity Identity Confusion in Multimodal Knowledge Editing
Shu Wu, Xiaotian Ye, Xinyu Mou +3
cs.CLcs.CVarXiv:2605.06096v12026Visual Saliency Based on Multiscale Deep Features
Guanbin Li, Yizhou Yu
cs.CVarXiv:1503.08663v32015A Survey on Object Detection in Optical Remote Sensing Images
Gong Cheng, Junwei Han
cs.CVarXiv:1603.06201v22016Twins: Revisiting the Design of Spatial Attention in Vision Transformers
Xiangxiang Chu, Zhi Tian, Yuqing Wang +5
cs.CVcs.AIcs.LGarXiv:2104.13840v42021Learning Discriminative Model Prediction for Tracking
Goutam Bhat, Martin Danelljan, Luc Van Gool +1
cs.CVarXiv:1904.07220v22019ATOM: Accurate Tracking by Overlap Maximization
Martin Danelljan, Goutam Bhat, Fahad Shahbaz Khan +1
cs.CVarXiv:1811.07628v22018AtlasNet: A Papier-Mâché Approach to Learning 3D Surface Generation
Thibault Groueix, Matthew Fisher, Vladimir G. Kim +2
cs.CVarXiv:1802.05384v32018FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling
Bowen Zhang, Yidong Wang, Wenxin Hou +4
cs.LGcs.CVarXiv:2110.08263v32021End-to-End Incremental Learning
Francisco M. Castro, Manuel J. Marín-Jiménez, Nicolás Guil +2
cs.CVarXiv:1807.09536v22018Simultaneous Detection and Segmentation
Bharath Hariharan, Pablo Arbeláez, Ross Girshick +1
cs.CVarXiv:1407.1808v12014Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation
Golnaz Ghiasi, Yin Cui, Aravind Srinivas +5
cs.CVarXiv:2012.07177v22020Relation Networks for Object Detection
Han Hu, Jiayuan Gu, Zheng Zhang +2
cs.CVarXiv:1711.11575v22017From Captions to Visual Concepts and Back
Hao Fang, Saurabh Gupta, Forrest Iandola +9
cs.CVcs.CLarXiv:1411.4952v32014Learning to Prompt for Continual Learning
Zifeng Wang, Zizhao Zhang, Chen-Yu Lee +7
cs.LGcs.CVarXiv:2112.08654v22021Learning Fine-grained Image Similarity with Deep Ranking
Jiang Wang, Yang song, Thomas Leung +5
cs.CVarXiv:1404.4661v12014Going deeper with Image Transformers
Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles +2
cs.CVarXiv:2103.17239v22021Open-vocabulary Object Detection via Vision and Language Knowledge Distillation
Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo +1
cs.CVcs.AIcs.LGarXiv:2104.13921v32021Stand-Alone Self-Attention in Vision Models
Prajit Ramachandran, Niki Parmar, Ashish Vaswani +3
cs.CVarXiv:1906.05909v12019LIFT: Learned Invariant Feature Transform
Kwang Moo Yi, Eduard Trulls, Vincent Lepetit +1
cs.CVarXiv:1603.09114v22016Denoising Diffusion Restoration Models
Bahjat Kawar, Michael Elad, Stefano Ermon +1
eess.IVcs.CVcs.LGarXiv:2201.11793v32022Delta-Adapter: Scalable Exemplar-Based Image Editing with Single-Pair Supervision
Jiacheng Chen, Songze Li, Han Fu +5
cs.CVarXiv:2605.07940v12026ResUNet++: An Advanced Architecture for Medical Image Segmentation
Debesh Jha, Pia H. Smedsrud, Michael A. Riegler +4
eess.IVcs.CVarXiv:1911.07067v12019MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Weihao Yu, Zhengyuan Yang, Linjie Li +5
cs.AIcs.CLcs.CVarXiv:2308.02490v42023Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs
Hao Wang, Yiqun Sun, Pengfei Wei +2
cs.CVcs.AIcs.CLarXiv:2605.07447v12026Transformer Tracking
Xin Chen, Bin Yan, Jiawen Zhu +3
cs.CVarXiv:2103.15436v12021BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning
Shaokai Ye, Vasileios Saveris, Yihao Qian +3
cs.CVcs.AIarXiv:2605.07394v12026Dynamic Edge-Conditioned Filters in Convolutional Neural Networks on Graphs
Martin Simonovsky, Nikos Komodakis
cs.CVcs.LGcs.NEarXiv:1704.02901v32017Context Encoding for Semantic Segmentation
Hang Zhang, Kristin Dana, Jianping Shi +4
cs.CVarXiv:1803.08904v12018Validation, comparison, and combination of algorithms for automatic detection of pulmonary nodules in computed tomography images: the LUNA16 challenge
Arnaud Arindra Adiyoso Setio, Alberto Traverso, Thomas de Bel +29
cs.CVarXiv:1612.08012v42016Implicit Preference Alignment for Human Image Animation
Yuanzhi Wang, Xuhua Ren, Jiaxiang Cheng +5
cs.CVcs.AIarXiv:2605.07545v12026