Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

17,041 to 17,100 of 18,866

  1. From Captions to Visual Concepts and Back

    Hao Fang, Saurabh Gupta, Forrest Iandola +9

    cs.CVcs.CLarXiv:1411.4952v32014
  2. Learning to Prompt for Continual Learning

    Zifeng Wang, Zizhao Zhang, Chen-Yu Lee +7

    cs.LGcs.CVarXiv:2112.08654v22021
  3. Learning Fine-grained Image Similarity with Deep Ranking

    Jiang Wang, Yang song, Thomas Leung +5

    cs.CVarXiv:1404.4661v12014
  4. Going deeper with Image Transformers

    Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles +2

    cs.CVarXiv:2103.17239v22021
  5. Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

    Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo +1

    cs.CVcs.AIcs.LGarXiv:2104.13921v32021
  6. Stand-Alone Self-Attention in Vision Models

    Prajit Ramachandran, Niki Parmar, Ashish Vaswani +3

    cs.CVarXiv:1906.05909v12019
  7. LIFT: Learned Invariant Feature Transform

    Kwang Moo Yi, Eduard Trulls, Vincent Lepetit +1

    cs.CVarXiv:1603.09114v22016
  8. Denoising Diffusion Restoration Models

    Bahjat Kawar, Michael Elad, Stefano Ermon +1

    eess.IVcs.CVcs.LGarXiv:2201.11793v32022
  9. Delta-Adapter: Scalable Exemplar-Based Image Editing with Single-Pair Supervision

    Jiacheng Chen, Songze Li, Han Fu +5

    cs.CVarXiv:2605.07940v12026
  10. ResUNet++: An Advanced Architecture for Medical Image Segmentation

    Debesh Jha, Pia H. Smedsrud, Michael A. Riegler +4

    eess.IVcs.CVarXiv:1911.07067v12019
  11. MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

    Weihao Yu, Zhengyuan Yang, Linjie Li +5

    cs.AIcs.CLcs.CVarXiv:2308.02490v42023
  12. Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs

    Hao Wang, Yiqun Sun, Pengfei Wei +2

    cs.CVcs.AIcs.CLarXiv:2605.07447v12026
  13. Transformer Tracking

    Xin Chen, Bin Yan, Jiawen Zhu +3

    cs.CVarXiv:2103.15436v12021
  14. BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning

    Shaokai Ye, Vasileios Saveris, Yihao Qian +3

    cs.CVcs.AIarXiv:2605.07394v12026
  15. Dynamic Edge-Conditioned Filters in Convolutional Neural Networks on Graphs

    Martin Simonovsky, Nikos Komodakis

    cs.CVcs.LGcs.NEarXiv:1704.02901v32017
  16. Context Encoding for Semantic Segmentation

    Hang Zhang, Kristin Dana, Jianping Shi +4

    cs.CVarXiv:1803.08904v12018
  17. Validation, comparison, and combination of algorithms for automatic detection of pulmonary nodules in computed tomography images: the LUNA16 challenge

    Arnaud Arindra Adiyoso Setio, Alberto Traverso, Thomas de Bel +29

    cs.CVarXiv:1612.08012v42016
  18. Implicit Preference Alignment for Human Image Animation

    Yuanzhi Wang, Xuhua Ren, Jiaxiang Cheng +5

    cs.CVcs.AIarXiv:2605.07545v12026
  19. Single-Shot Refinement Neural Network for Object Detection

    Shifeng Zhang, Longyin Wen, Xiao Bian +2

    cs.CVarXiv:1711.06897v32017
  20. MC-RFM: Geometry-Aware Few-Shot Adaptation via Mixed-Curvature Riemannian Flow Matching

    Salim Khazem, Ibrahim Mohamed Serouis, Zakaria Ezzahed

    cs.CVcs.AIcs.LGarXiv:2605.08557v12026
  21. Contrastive Representation Distillation

    Yonglong Tian, Dilip Krishnan, Phillip Isola

    cs.LGcs.CVstat.MLarXiv:1910.10699v32019
  22. DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR

    Shilong Liu, Feng Li, Hao Zhang +5

    cs.CVarXiv:2201.12329v42022
  23. Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields Reconstruction

    Cheng Sun, Min Sun, Hwann-Tzong Chen

    cs.CVarXiv:2111.11215v22021
  24. Learning Temporal Regularity in Video Sequences

    Mahmudul Hasan, Jonghyun Choi, Jan Neumann +2

    cs.CVarXiv:1604.04574v12016
  25. Learning Robust Global Representations by Penalizing Local Predictive Power

    Haohan Wang, Songwei Ge, Eric P. Xing +1

    cs.CVarXiv:1905.13549v22019
  26. The highD Dataset: A Drone Dataset of Naturalistic Vehicle Trajectories on German Highways for Validation of Highly Automated Driving Systems

    Robert Krajewski, Julian Bock, Laurent Kloeker +1

    cs.CVcs.AIcs.IRarXiv:1810.05642v12018
  27. Human Motion Diffusion Model

    Guy Tevet, Sigal Raab, Brian Gordon +3

    cs.CVcs.GRarXiv:2209.14916v22022
  28. Towards Accurate Generative Models of Video: A New Metric & Challenges

    Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach +3

    cs.CVcs.AIcs.LGarXiv:1812.01717v22018
  29. From Pixels to Concepts: Do Segmentation Models Understand What They Segment?

    Shuang Liang, Zeqing Wang, Yuxian Li +2

    cs.CVarXiv:2605.09591v12026
  30. TOOD: Task-aligned One-stage Object Detection

    Chengjian Feng, Yujie Zhong, Yu Gao +2

    cs.CVarXiv:2108.07755v32021
  31. RigidFormer: Learning Rigid Dynamics using Transformers

    Zhiyang Dou, Minghao Guo, Haixu Wu +3

    cs.CVcs.AIcs.GRarXiv:2605.09196v12026
  32. Additive Margin Softmax for Face Verification

    Feng Wang, Weiyang Liu, Haijun Liu +1

    cs.CVarXiv:1801.05599v42018
  33. LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?

    Kechen Fang, Yihua Qin, Chongyi Wang +3

    cs.CVarXiv:2605.08985v12026
  34. Convolutional Neural Networks at Constrained Time Cost

    Kaiming He, Jian Sun

    cs.CVarXiv:1412.1710v12014
  35. Multi-Concept Customization of Text-to-Image Diffusion

    Nupur Kumari, Bingliang Zhang, Richard Zhang +2

    cs.CVcs.GRcs.LGarXiv:2212.04488v22022
  36. Flower: A Friendly Federated Learning Research Framework

    Daniel J. Beutel, Taner Topal, Akhil Mathur +8

    cs.LGcs.CVstat.MLarXiv:2007.14390v52020
  37. Image Deblurring and Super-resolution by Adaptive Sparse Domain Selection and Adaptive Regularization

    Weisheng Dong, Lei Zhang, Guangming Shi +1

    cs.CVcs.MMarXiv:1012.1184v12010
  38. Reinforcing Multimodal Reasoning Against Visual Degradation

    Rui Liu, Dian Yu, Haolin Liu +6

    cs.CVcs.CLarXiv:2605.09262v12026
  39. Attention to Scale: Scale-aware Semantic Image Segmentation

    Liang-Chieh Chen, Yi Yang, Jiang Wang +2

    cs.CVarXiv:1511.03339v22015
  40. VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning

    Adrien Bardes, Jean Ponce, Yann LeCun

    cs.CVcs.AIcs.LGarXiv:2105.04906v32021
  41. The Replica Dataset: A Digital Replica of Indoor Spaces

    Julian Straub, Thomas Whelan, Lingni Ma +27

    cs.CVcs.GReess.IVarXiv:1906.05797v12019
  42. A Review on Deep Learning Techniques Applied to Semantic Segmentation

    Alberto Garcia-Garcia, Sergio Orts-Escolano, Sergiu Oprea +2

    cs.CVcs.AIarXiv:1704.06857v12017
  43. Tracking Objects as Points

    Xingyi Zhou, Vladlen Koltun, Philipp Krähenbühl

    cs.CVarXiv:2004.01177v22020
  44. Taskonomy: Disentangling Task Transfer Learning

    Amir Zamir, Alexander Sax, William Shen +3

    cs.CVcs.AIcs.LGarXiv:1804.08328v12018
  45. HOTA: A Higher Order Metric for Evaluating Multi-Object Tracking

    Jonathon Luiten, Aljosa Osep, Patrick Dendorfer +4

    cs.CVarXiv:2009.07736v22020
  46. DivideMix: Learning with Noisy Labels as Semi-supervised Learning

    Junnan Li, Richard Socher, Steven C. H. Hoi

    cs.CVarXiv:2002.07394v12020
  47. MetaFormer Is Actually What You Need for Vision

    Weihao Yu, Mi Luo, Pan Zhou +5

    cs.CVcs.AIcs.LGarXiv:2111.11418v32021
  48. CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows

    Xiaoyi Dong, Jianmin Bao, Dongdong Chen +5

    cs.CVcs.LGarXiv:2107.00652v32021
  49. Scaling Egocentric Vision: The EPIC-KITCHENS Dataset

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella +8

    cs.CVarXiv:1804.02748v22018
  50. Big Transfer (BiT): General Visual Representation Learning

    Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai +4

    cs.CVcs.LGarXiv:1912.11370v32019
  51. MegaDepth: Learning Single-View Depth Prediction from Internet Photos

    Zhengqi Li, Noah Snavely

    cs.CVarXiv:1804.00607v42018
  52. Enhancing Geometric Factors in Model Learning and Inference for Object Detection and Instance Segmentation

    Zhaohui Zheng, Ping Wang, Dongwei Ren +4

    cs.CVarXiv:2005.03572v42020
  53. Volume Rendering of Neural Implicit Surfaces

    Lior Yariv, Jiatao Gu, Yoni Kasten +1

    cs.CVarXiv:2106.12052v22021
  54. Evaluating the visualization of what a Deep Neural Network has learned

    Wojciech Samek, Alexander Binder, Grégoire Montavon +2

    cs.CVarXiv:1509.06321v12015
  55. Stacked Cross Attention for Image-Text Matching

    Kuang-Huei Lee, Xi Chen, Gang Hua +2

    cs.CVcs.AIcs.LGarXiv:1803.08024v22018
  56. Black-box Adversarial Attacks with Limited Queries and Information

    Andrew Ilyas, Logan Engstrom, Anish Athalye +1

    cs.CVcs.CRstat.MLarXiv:1804.08598v32018
  57. End-to-End Multi-Task Learning with Attention

    Shikun Liu, Edward Johns, Andrew J. Davison

    cs.CVarXiv:1803.10704v22018
  58. Simultaneous Deep Transfer Across Domains and Tasks

    Eric Tzeng, Judy Hoffman, Trevor Darrell +1

    cs.CVarXiv:1510.02192v12015
  59. Convolutional Radio Modulation Recognition Networks

    Timothy J O'Shea, Johnathan Corgan, T. Charles Clancy

    cs.LGcs.CVarXiv:1602.04105v32016
  60. YouTube-8M: A Large-Scale Video Classification Benchmark

    Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee +4

    cs.CVarXiv:1609.08675v12016