Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

18,061 to 18,120 of 18,837

  1. ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks

    Qilong Wang, Banggu Wu, Pengfei Zhu +3

    cs.CVarXiv:1910.03151v42019
  2. Long-term Recurrent Convolutional Networks for Visual Recognition and Description

    Jeff Donahue, Lisa Anne Hendricks, Marcus Rohrbach +4

    cs.CVarXiv:1411.4389v42014
  3. What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?

    Alex Kendall, Yarin Gal

    cs.CVarXiv:1703.04977v22017
  4. ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design

    Ningning Ma, Xiangyu Zhang, Hai-Tao Zheng +1

    cs.CVarXiv:1807.11164v12018
  5. Show and Tell: A Neural Image Caption Generator

    Oriol Vinyals, Alexander Toshev, Samy Bengio +1

    cs.CVarXiv:1411.4555v22014
  6. Zero-Shot Text-to-Image Generation

    Aditya Ramesh, Mikhail Pavlov, Gabriel Goh +5

    cs.CVcs.LGarXiv:2102.12092v22021
  7. Network In Network

    Min Lin, Qiang Chen, Shuicheng Yan

    cs.NEcs.CVcs.LGarXiv:1312.4400v32013
  8. Deformable Convolutional Networks

    Jifeng Dai, Haozhi Qi, Yuwen Xiong +4

    cs.CVarXiv:1703.06211v32017
  9. CARLA: An Open Urban Driving Simulator

    Alexey Dosovitskiy, German Ros, Felipe Codevilla +2

    cs.LGcs.AIcs.CVarXiv:1711.03938v12017
  10. Accurate Image Super-Resolution Using Very Deep Convolutional Networks

    Jiwon Kim, Jung Kwon Lee, Kyoung Mu Lee

    cs.CVcs.LGarXiv:1511.04587v22015
  11. Path Aggregation Network for Instance Segmentation

    Shu Liu, Lu Qi, Haifang Qin +2

    cs.CVarXiv:1803.01534v42018
  12. Analyzing and Improving the Image Quality of StyleGAN

    Tero Karras, Samuli Laine, Miika Aittala +3

    cs.CVcs.LGcs.NEarXiv:1912.04958v22019
  13. BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

    Junnan Li, Dongxu Li, Caiming Xiong +1

    cs.CVarXiv:2201.12086v22022
  14. EfficientDet: Scalable and Efficient Object Detection

    Mingxing Tan, Ruoming Pang, Quoc V. Le

    cs.CVcs.LGeess.IVarXiv:1911.09070v72019
  15. UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

    Khurram Soomro, Amir Roshan Zamir, Mubarak Shah

    cs.CVarXiv:1212.0402v12012
  16. Enhanced Deep Residual Networks for Single Image Super-Resolution

    Bee Lim, Sanghyun Son, Heewon Kim +2

    cs.CVarXiv:1707.02921v12017
  17. Adding Conditional Control to Text-to-Image Diffusion Models

    Lvmin Zhang, Anyi Rao, Maneesh Agrawala

    cs.CVcs.AIcs.GRarXiv:2302.05543v32023
  18. Learning both Weights and Connections for Efficient Neural Networks

    Song Han, Jeff Pool, John Tran +1

    cs.NEcs.CVcs.LGarXiv:1506.02626v32015
  19. Spatial Transformer Networks

    Max Jaderberg, Karen Simonyan, Andrew Zisserman +1

    cs.CVarXiv:1506.02025v32015
  20. Deformable DETR: Deformable Transformers for End-to-End Object Detection

    Xizhou Zhu, Weijie Su, Lewei Lu +3

    cs.CVarXiv:2010.04159v42020
  21. 3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation

    Özgün Çiçek, Ahmed Abdulkadir, Soeren S. Lienkamp +2

    cs.CVarXiv:1606.06650v12016
  22. Two-Stream Convolutional Networks for Action Recognition in Videos

    Karen Simonyan, Andrew Zisserman

    cs.CVarXiv:1406.2199v22014
  23. Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps

    Karen Simonyan, Andrea Vedaldi, Andrew Zisserman

    cs.CVarXiv:1312.6034v22013
  24. SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5MB model size

    Forrest N. Iandola, Song Han, Matthew W. Moskewicz +3

    cs.CVcs.AIarXiv:1602.07360v42016
  25. ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices

    Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin +1

    cs.CVarXiv:1707.01083v22017
  26. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

    Junnan Li, Dongxu Li, Silvio Savarese +1

    cs.CVarXiv:2301.12597v32023
  27. Wide Residual Networks

    Sergey Zagoruyko, Nikos Komodakis

    cs.CVcs.LGcs.NEarXiv:1605.07146v42016
  28. A ConvNet for the 2020s

    Zhuang Liu, Hanzi Mao, Chao-Yuan Wu +3

    cs.CVarXiv:2201.03545v22022
  29. SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers

    Enze Xie, Wenhai Wang, Zhiding Yu +3

    cs.CVcs.LGarXiv:2105.15203v32021
  30. Hierarchical Text-Conditional Image Generation with CLIP Latents

    Aditya Ramesh, Prafulla Dhariwal, Alex Nichol +2

    cs.CVarXiv:2204.06125v12022
  31. iMaC: Translating Actions into Motion and Contact Images for Embodied World Models

    Zhenyu Wu, Xiuwei Xu, Yukun Zhou +8

    cs.ROcs.CVarXiv:2606.09813v12026
  32. Surflo: Consistent 3D Surface Flow Model with Global State

    Antoine Guédon, Shu Nakamura, Nicolas Dufour +3

    cs.CVarXiv:2606.13644v12026
  33. Emerging Properties in Self-Supervised Vision Transformers

    Mathilde Caron, Hugo Touvron, Ishan Misra +4

    cs.CVarXiv:2104.14294v22021
  34. MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffold

    Yang Zhou, Ziheng Wang, Yuqin Lu +4

    cs.CVarXiv:2606.13376v22026
  35. Searching for MobileNetV3

    Andrew Howard, Mark Sandler, Grace Chu +9

    cs.CVarXiv:1905.02244v52019
  36. Training data-efficient image transformers & distillation through attention

    Hugo Touvron, Matthieu Cord, Matthijs Douze +3

    cs.CVarXiv:2012.12877v22020
  37. Hierarchical Data Selection via Manifold Coverage and Sparse Feature Coverage in LLM Post-training

    Peng Sun, Yi Yang, Antong Zhang +7

    cs.LGcs.CVarXiv:2608.16927v12026
  38. Image Super-Resolution Using Deep Convolutional Networks

    Chao Dong, Chen Change Loy, Kaiming He +1

    cs.CVcs.NEarXiv:1501.00092v32014
  39. UNet++: A Nested U-Net Architecture for Medical Image Segmentation

    Zongwei Zhou, Md Mahfuzur Rahman Siddiquee, Nima Tajbakhsh +1

    cs.CVcs.LGeess.IVarXiv:1807.10165v12018
  40. YOLO9000: Better, Faster, Stronger

    Joseph Redmon, Ali Farhadi

    cs.CVarXiv:1612.08242v12016
  41. MagViT: Interpretable Multi-Magnification Transformers with Patient-Level Model Selection for Breast Histopathology

    Nabil Ashab, Soumit Kumar Kundu, Saif Mahmud Parvez +3

    eess.IVcs.CVcs.LGarXiv:2608.16959v12026
  42. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset

    Joao Carreira, Andrew Zisserman

    cs.CVcs.LGarXiv:1705.07750v32017
  43. Deep Learning Face Attributes in the Wild

    Ziwei Liu, Ping Luo, Xiaogang Wang +1

    cs.CVarXiv:1411.7766v32014
  44. Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting

    Xingjian Shi, Zhourong Chen, Hao Wang +3

    cs.CVarXiv:1506.04214v22015
  45. VideoMDM: Towards 3D Human Motion Generation From 2D Supervision

    Amir Mann, Gal Michael Harari, Merav Keidar +1

    cs.LGcs.CVarXiv:2606.13364v12026
  46. InterleaveThinker: Reinforcing Agentic Interleaved Generation

    Dian Zheng, Harry Lee, Manyuan Zhang +4

    cs.CVarXiv:2606.13679v22026
  47. Mask R-CNN

    Kaiming He, Georgia Gkioxari, Piotr Dollár +1

    cs.CVarXiv:1703.06870v32017
  48. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

    Song Han, Huizi Mao, William J. Dally

    cs.CVcs.NEarXiv:1510.00149v52015
  49. V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation

    Fausto Milletari, Nassir Navab, Seyed-Ahmad Ahmadi

    cs.CVarXiv:1606.04797v12016
  50. Identity Mappings in Deep Residual Networks

    Kaiming He, Xiangyu Zhang, Shaoqing Ren +1

    cs.CVcs.LGarXiv:1603.05027v32016
  51. Lymphocyte Mimicry Correction via Region-Level Tissue Reasoning and Unbalanced Optimal Transport

    Xiang Li, Yuqi Wang, Casey C. Heirman +2

    cs.CVcs.LGarXiv:2608.17151v12026
  52. Rethinking Atrous Convolution for Semantic Image Segmentation

    Liang-Chieh Chen, George Papandreou, Florian Schroff +1

    cs.CVarXiv:1706.05587v32017
  53. Non-local Neural Networks

    Xiaolong Wang, Ross Girshick, Abhinav Gupta +1

    cs.CVarXiv:1711.07971v32017
  54. Visual Instruction Tuning

    Haotian Liu, Chunyuan Li, Qingyang Wu +1

    cs.CVcs.AIcs.CLarXiv:2304.08485v22023
  55. Learning Deep Features for Discriminative Localization

    Bolei Zhou, Aditya Khosla, Agata Lapedriza +2

    cs.CVarXiv:1512.04150v12015
  56. Show, Attend and Tell: Neural Image Caption Generation with Visual Attention

    Kelvin Xu, Jimmy Ba, Ryan Kiros +5

    cs.LGcs.CVarXiv:1502.03044v32015
  57. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors

    Chien-Yao Wang, Alexey Bochkovskiy, Hong-Yuan Mark Liao

    cs.CVarXiv:2207.02696v12022
  58. Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms

    Han Xiao, Kashif Rasul, Roland Vollgraf

    cs.LGcs.CVstat.MLarXiv:1708.07747v22017
  59. Track2View: 4D-Consistent Camera-Controlled Video Generation via Paired 3D Point Tracks

    Feng Qiao, Zhaochong An, Zhexiao Xiong +2

    cs.CVarXiv:2606.15534v12026
  60. Perceptual Losses for Real-Time Style Transfer and Super-Resolution

    Justin Johnson, Alexandre Alahi, Li Fei-Fei

    cs.CVcs.LGarXiv:1603.08155v12016