Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
17,041 to 17,100 of 18,866
From Captions to Visual Concepts and Back
Hao Fang, Saurabh Gupta, Forrest Iandola +9
cs.CVcs.CLarXiv:1411.4952v32014Learning to Prompt for Continual Learning
Zifeng Wang, Zizhao Zhang, Chen-Yu Lee +7
cs.LGcs.CVarXiv:2112.08654v22021Learning Fine-grained Image Similarity with Deep Ranking
Jiang Wang, Yang song, Thomas Leung +5
cs.CVarXiv:1404.4661v12014Going deeper with Image Transformers
Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles +2
cs.CVarXiv:2103.17239v22021Open-vocabulary Object Detection via Vision and Language Knowledge Distillation
Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo +1
cs.CVcs.AIcs.LGarXiv:2104.13921v32021Stand-Alone Self-Attention in Vision Models
Prajit Ramachandran, Niki Parmar, Ashish Vaswani +3
cs.CVarXiv:1906.05909v12019LIFT: Learned Invariant Feature Transform
Kwang Moo Yi, Eduard Trulls, Vincent Lepetit +1
cs.CVarXiv:1603.09114v22016Denoising Diffusion Restoration Models
Bahjat Kawar, Michael Elad, Stefano Ermon +1
eess.IVcs.CVcs.LGarXiv:2201.11793v32022Delta-Adapter: Scalable Exemplar-Based Image Editing with Single-Pair Supervision
Jiacheng Chen, Songze Li, Han Fu +5
cs.CVarXiv:2605.07940v12026ResUNet++: An Advanced Architecture for Medical Image Segmentation
Debesh Jha, Pia H. Smedsrud, Michael A. Riegler +4
eess.IVcs.CVarXiv:1911.07067v12019MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Weihao Yu, Zhengyuan Yang, Linjie Li +5
cs.AIcs.CLcs.CVarXiv:2308.02490v42023Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs
Hao Wang, Yiqun Sun, Pengfei Wei +2
cs.CVcs.AIcs.CLarXiv:2605.07447v12026Transformer Tracking
Xin Chen, Bin Yan, Jiawen Zhu +3
cs.CVarXiv:2103.15436v12021BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning
Shaokai Ye, Vasileios Saveris, Yihao Qian +3
cs.CVcs.AIarXiv:2605.07394v12026Dynamic Edge-Conditioned Filters in Convolutional Neural Networks on Graphs
Martin Simonovsky, Nikos Komodakis
cs.CVcs.LGcs.NEarXiv:1704.02901v32017Context Encoding for Semantic Segmentation
Hang Zhang, Kristin Dana, Jianping Shi +4
cs.CVarXiv:1803.08904v12018Validation, comparison, and combination of algorithms for automatic detection of pulmonary nodules in computed tomography images: the LUNA16 challenge
Arnaud Arindra Adiyoso Setio, Alberto Traverso, Thomas de Bel +29
cs.CVarXiv:1612.08012v42016Implicit Preference Alignment for Human Image Animation
Yuanzhi Wang, Xuhua Ren, Jiaxiang Cheng +5
cs.CVcs.AIarXiv:2605.07545v12026Single-Shot Refinement Neural Network for Object Detection
Shifeng Zhang, Longyin Wen, Xiao Bian +2
cs.CVarXiv:1711.06897v32017MC-RFM: Geometry-Aware Few-Shot Adaptation via Mixed-Curvature Riemannian Flow Matching
Salim Khazem, Ibrahim Mohamed Serouis, Zakaria Ezzahed
cs.CVcs.AIcs.LGarXiv:2605.08557v12026Contrastive Representation Distillation
Yonglong Tian, Dilip Krishnan, Phillip Isola
cs.LGcs.CVstat.MLarXiv:1910.10699v32019DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR
Shilong Liu, Feng Li, Hao Zhang +5
cs.CVarXiv:2201.12329v42022Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields Reconstruction
Cheng Sun, Min Sun, Hwann-Tzong Chen
cs.CVarXiv:2111.11215v22021Learning Temporal Regularity in Video Sequences
Mahmudul Hasan, Jonghyun Choi, Jan Neumann +2
cs.CVarXiv:1604.04574v12016Learning Robust Global Representations by Penalizing Local Predictive Power
Haohan Wang, Songwei Ge, Eric P. Xing +1
cs.CVarXiv:1905.13549v22019The highD Dataset: A Drone Dataset of Naturalistic Vehicle Trajectories on German Highways for Validation of Highly Automated Driving Systems
Robert Krajewski, Julian Bock, Laurent Kloeker +1
cs.CVcs.AIcs.IRarXiv:1810.05642v12018Human Motion Diffusion Model
Guy Tevet, Sigal Raab, Brian Gordon +3
cs.CVcs.GRarXiv:2209.14916v22022Towards Accurate Generative Models of Video: A New Metric & Challenges
Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach +3
cs.CVcs.AIcs.LGarXiv:1812.01717v22018From Pixels to Concepts: Do Segmentation Models Understand What They Segment?
Shuang Liang, Zeqing Wang, Yuxian Li +2
cs.CVarXiv:2605.09591v12026TOOD: Task-aligned One-stage Object Detection
Chengjian Feng, Yujie Zhong, Yu Gao +2
cs.CVarXiv:2108.07755v32021RigidFormer: Learning Rigid Dynamics using Transformers
Zhiyang Dou, Minghao Guo, Haixu Wu +3
cs.CVcs.AIcs.GRarXiv:2605.09196v12026Additive Margin Softmax for Face Verification
Feng Wang, Weiyang Liu, Haijun Liu +1
cs.CVarXiv:1801.05599v42018LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?
Kechen Fang, Yihua Qin, Chongyi Wang +3
cs.CVarXiv:2605.08985v12026Convolutional Neural Networks at Constrained Time Cost
Kaiming He, Jian Sun
cs.CVarXiv:1412.1710v12014Multi-Concept Customization of Text-to-Image Diffusion
Nupur Kumari, Bingliang Zhang, Richard Zhang +2
cs.CVcs.GRcs.LGarXiv:2212.04488v22022Flower: A Friendly Federated Learning Research Framework
Daniel J. Beutel, Taner Topal, Akhil Mathur +8
cs.LGcs.CVstat.MLarXiv:2007.14390v52020Image Deblurring and Super-resolution by Adaptive Sparse Domain Selection and Adaptive Regularization
Weisheng Dong, Lei Zhang, Guangming Shi +1
cs.CVcs.MMarXiv:1012.1184v12010Reinforcing Multimodal Reasoning Against Visual Degradation
Rui Liu, Dian Yu, Haolin Liu +6
cs.CVcs.CLarXiv:2605.09262v12026Attention to Scale: Scale-aware Semantic Image Segmentation
Liang-Chieh Chen, Yi Yang, Jiang Wang +2
cs.CVarXiv:1511.03339v22015VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning
Adrien Bardes, Jean Ponce, Yann LeCun
cs.CVcs.AIcs.LGarXiv:2105.04906v32021The Replica Dataset: A Digital Replica of Indoor Spaces
Julian Straub, Thomas Whelan, Lingni Ma +27
cs.CVcs.GReess.IVarXiv:1906.05797v12019A Review on Deep Learning Techniques Applied to Semantic Segmentation
Alberto Garcia-Garcia, Sergio Orts-Escolano, Sergiu Oprea +2
cs.CVcs.AIarXiv:1704.06857v12017Tracking Objects as Points
Xingyi Zhou, Vladlen Koltun, Philipp Krähenbühl
cs.CVarXiv:2004.01177v22020Taskonomy: Disentangling Task Transfer Learning
Amir Zamir, Alexander Sax, William Shen +3
cs.CVcs.AIcs.LGarXiv:1804.08328v12018HOTA: A Higher Order Metric for Evaluating Multi-Object Tracking
Jonathon Luiten, Aljosa Osep, Patrick Dendorfer +4
cs.CVarXiv:2009.07736v22020DivideMix: Learning with Noisy Labels as Semi-supervised Learning
Junnan Li, Richard Socher, Steven C. H. Hoi
cs.CVarXiv:2002.07394v12020MetaFormer Is Actually What You Need for Vision
Weihao Yu, Mi Luo, Pan Zhou +5
cs.CVcs.AIcs.LGarXiv:2111.11418v32021CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows
Xiaoyi Dong, Jianmin Bao, Dongdong Chen +5
cs.CVcs.LGarXiv:2107.00652v32021Scaling Egocentric Vision: The EPIC-KITCHENS Dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella +8
cs.CVarXiv:1804.02748v22018Big Transfer (BiT): General Visual Representation Learning
Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai +4
cs.CVcs.LGarXiv:1912.11370v32019MegaDepth: Learning Single-View Depth Prediction from Internet Photos
Zhengqi Li, Noah Snavely
cs.CVarXiv:1804.00607v42018Enhancing Geometric Factors in Model Learning and Inference for Object Detection and Instance Segmentation
Zhaohui Zheng, Ping Wang, Dongwei Ren +4
cs.CVarXiv:2005.03572v42020Volume Rendering of Neural Implicit Surfaces
Lior Yariv, Jiatao Gu, Yoni Kasten +1
cs.CVarXiv:2106.12052v22021Evaluating the visualization of what a Deep Neural Network has learned
Wojciech Samek, Alexander Binder, Grégoire Montavon +2
cs.CVarXiv:1509.06321v12015Stacked Cross Attention for Image-Text Matching
Kuang-Huei Lee, Xi Chen, Gang Hua +2
cs.CVcs.AIcs.LGarXiv:1803.08024v22018Black-box Adversarial Attacks with Limited Queries and Information
Andrew Ilyas, Logan Engstrom, Anish Athalye +1
cs.CVcs.CRstat.MLarXiv:1804.08598v32018End-to-End Multi-Task Learning with Attention
Shikun Liu, Edward Johns, Andrew J. Davison
cs.CVarXiv:1803.10704v22018Simultaneous Deep Transfer Across Domains and Tasks
Eric Tzeng, Judy Hoffman, Trevor Darrell +1
cs.CVarXiv:1510.02192v12015Convolutional Radio Modulation Recognition Networks
Timothy J O'Shea, Johnathan Corgan, T. Charles Clancy
cs.LGcs.CVarXiv:1602.04105v32016YouTube-8M: A Large-Scale Video Classification Benchmark
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee +4
cs.CVarXiv:1609.08675v12016