Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

18,121 to 18,180 of 18,866

  1. Surflo: Consistent 3D Surface Flow Model with Global State

    Antoine Guédon, Shu Nakamura, Nicolas Dufour +3

    cs.CVarXiv:2606.13644v12026
  2. Emerging Properties in Self-Supervised Vision Transformers

    Mathilde Caron, Hugo Touvron, Ishan Misra +4

    cs.CVarXiv:2104.14294v22021
  3. MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffold

    Yang Zhou, Ziheng Wang, Yuqin Lu +4

    cs.CVarXiv:2606.13376v22026
  4. Searching for MobileNetV3

    Andrew Howard, Mark Sandler, Grace Chu +9

    cs.CVarXiv:1905.02244v52019
  5. Training data-efficient image transformers & distillation through attention

    Hugo Touvron, Matthieu Cord, Matthijs Douze +3

    cs.CVarXiv:2012.12877v22020
  6. Hierarchical Data Selection via Manifold Coverage and Sparse Feature Coverage in LLM Post-training

    Peng Sun, Yi Yang, Antong Zhang +7

    cs.LGcs.CVarXiv:2608.16927v12026
  7. Image Super-Resolution Using Deep Convolutional Networks

    Chao Dong, Chen Change Loy, Kaiming He +1

    cs.CVcs.NEarXiv:1501.00092v32014
  8. UNet++: A Nested U-Net Architecture for Medical Image Segmentation

    Zongwei Zhou, Md Mahfuzur Rahman Siddiquee, Nima Tajbakhsh +1

    cs.CVcs.LGeess.IVarXiv:1807.10165v12018
  9. YOLO9000: Better, Faster, Stronger

    Joseph Redmon, Ali Farhadi

    cs.CVarXiv:1612.08242v12016
  10. MagViT: Interpretable Multi-Magnification Transformers with Patient-Level Model Selection for Breast Histopathology

    Nabil Ashab, Soumit Kumar Kundu, Saif Mahmud Parvez +3

    eess.IVcs.CVcs.LGarXiv:2608.16959v12026
  11. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset

    Joao Carreira, Andrew Zisserman

    cs.CVcs.LGarXiv:1705.07750v32017
  12. Deep Learning Face Attributes in the Wild

    Ziwei Liu, Ping Luo, Xiaogang Wang +1

    cs.CVarXiv:1411.7766v32014
  13. Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting

    Xingjian Shi, Zhourong Chen, Hao Wang +3

    cs.CVarXiv:1506.04214v22015
  14. VideoMDM: Towards 3D Human Motion Generation From 2D Supervision

    Amir Mann, Gal Michael Harari, Merav Keidar +1

    cs.LGcs.CVarXiv:2606.13364v12026
  15. InterleaveThinker: Reinforcing Agentic Interleaved Generation

    Dian Zheng, Harry Lee, Manyuan Zhang +4

    cs.CVarXiv:2606.13679v22026
  16. Mask R-CNN

    Kaiming He, Georgia Gkioxari, Piotr Dollár +1

    cs.CVarXiv:1703.06870v32017
  17. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

    Song Han, Huizi Mao, William J. Dally

    cs.CVcs.NEarXiv:1510.00149v52015
  18. V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation

    Fausto Milletari, Nassir Navab, Seyed-Ahmad Ahmadi

    cs.CVarXiv:1606.04797v12016
  19. Identity Mappings in Deep Residual Networks

    Kaiming He, Xiangyu Zhang, Shaoqing Ren +1

    cs.CVcs.LGarXiv:1603.05027v32016
  20. Lymphocyte Mimicry Correction via Region-Level Tissue Reasoning and Unbalanced Optimal Transport

    Xiang Li, Yuqi Wang, Casey C. Heirman +2

    cs.CVcs.LGarXiv:2608.17151v12026
  21. Rethinking Atrous Convolution for Semantic Image Segmentation

    Liang-Chieh Chen, George Papandreou, Florian Schroff +1

    cs.CVarXiv:1706.05587v32017
  22. Non-local Neural Networks

    Xiaolong Wang, Ross Girshick, Abhinav Gupta +1

    cs.CVarXiv:1711.07971v32017
  23. Visual Instruction Tuning

    Haotian Liu, Chunyuan Li, Qingyang Wu +1

    cs.CVcs.AIcs.CLarXiv:2304.08485v22023
  24. Learning Deep Features for Discriminative Localization

    Bolei Zhou, Aditya Khosla, Agata Lapedriza +2

    cs.CVarXiv:1512.04150v12015
  25. Show, Attend and Tell: Neural Image Caption Generation with Visual Attention

    Kelvin Xu, Jimmy Ba, Ryan Kiros +5

    cs.LGcs.CVarXiv:1502.03044v32015
  26. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors

    Chien-Yao Wang, Alexey Bochkovskiy, Hong-Yuan Mark Liao

    cs.CVarXiv:2207.02696v12022
  27. Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms

    Han Xiao, Kashif Rasul, Roland Vollgraf

    cs.LGcs.CVstat.MLarXiv:1708.07747v22017
  28. Track2View: 4D-Consistent Camera-Controlled Video Generation via Paired 3D Point Tracks

    Feng Qiao, Zhaochong An, Zhexiao Xiong +2

    cs.CVarXiv:2606.15534v12026
  29. Perceptual Losses for Real-Time Style Transfer and Super-Resolution

    Justin Johnson, Alexandre Alahi, Li Fei-Fei

    cs.CVcs.LGarXiv:1603.08155v12016
  30. Conditional Generative Adversarial Nets

    Mehdi Mirza, Simon Osindero

    cs.LGcs.AIcs.CVarXiv:1411.1784v12014
  31. Aggregated Residual Transformations for Deep Neural Networks

    Saining Xie, Ross Girshick, Piotr Dollár +2

    cs.CVarXiv:1611.05431v22016
  32. Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren +1

    cs.CVarXiv:1406.4729v42014
  33. Caffe: Convolutional Architecture for Fast Feature Embedding

    Yangqing Jia, Evan Shelhamer, Jeff Donahue +5

    cs.CVcs.LGcs.NEarXiv:1408.5093v12014
  34. FaceNet: A Unified Embedding for Face Recognition and Clustering

    Florian Schroff, Dmitry Kalenichenko, James Philbin

    cs.CVarXiv:1503.03832v32015
  35. Diffusion Models Beat GANs on Image Synthesis

    Prafulla Dhariwal, Alex Nichol

    cs.LGcs.AIcs.CVarXiv:2105.05233v42021
  36. SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation

    Vijay Badrinarayanan, Alex Kendall, Roberto Cipolla

    cs.CVcs.LGcs.NEarXiv:1511.00561v32015
  37. End-to-End Object Detection with Transformers

    Nicolas Carion, Francisco Massa, Gabriel Synnaeve +3

    cs.CVarXiv:2005.12872v32020
  38. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric

    Richard Zhang, Phillip Isola, Alexei A. Efros +2

    cs.CVcs.GRarXiv:1801.03924v22018
  39. Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network

    Christian Ledig, Lucas Theis, Ferenc Huszar +8

    cs.CVstat.MLarXiv:1609.04802v52016
  40. Masked Autoencoders Are Scalable Vision Learners

    Kaiming He, Xinlei Chen, Saining Xie +3

    cs.CVarXiv:2111.06377v32021
  41. A Survey on Deep Learning in Medical Image Analysis

    Geert Litjens, Thijs Kooi, Babak Ehteshami Bejnordi +6

    cs.CVarXiv:1702.05747v22017
  42. Optimize Your Sampling: Tuned Diffusion Sampling with Bayesian Optimization

    Travis Zhang, Christian Belardi, Justin Lovelace +4

    cs.LGcs.CVarXiv:2608.18040v12026
  43. Pyramid Scene Parsing Network

    Hengshuang Zhao, Jianping Shi, Xiaojuan Qi +2

    cs.CVarXiv:1612.01105v22016
  44. Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks

    Alec Radford, Luke Metz, Soumith Chintala

    cs.LGcs.CVarXiv:1511.06434v22015
  45. Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning

    Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke +1

    cs.CVarXiv:1602.07261v22016
  46. YOLOv4: Optimal Speed and Accuracy of Object Detection

    Alexey Bochkovskiy, Chien-Yao Wang, Hong-Yuan Mark Liao

    cs.CVeess.IVarXiv:2004.10934v12020
  47. Visualizing and Understanding Convolutional Networks

    Matthew D Zeiler, Rob Fergus

    cs.CVarXiv:1311.2901v32013
  48. Intriguing properties of neural networks

    Christian Szegedy, Wojciech Zaremba, Ilya Sutskever +4

    cs.CVcs.LGcs.NEarXiv:1312.6199v42013
  49. The Llama 3 Herd of Models

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri +558

    cs.AIcs.CLcs.CVarXiv:2407.21783v32024
  50. Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation

    Liang-Chieh Chen, Yukun Zhu, George Papandreou +2

    cs.CVarXiv:1802.02611v32018
  51. UAV Video Deblurring via Motion-Aware Diffusion: A Path to Robust Target Detection

    Zhiqiang Hu, Shouren Huang, Masatoshi Ishikawa

    cs.CVcs.AIcs.ROarXiv:2608.15259v12026
  52. Image-to-Image Translation with Conditional Adversarial Networks

    Phillip Isola, Jun-Yan Zhu, Tinghui Zhou +1

    cs.CVarXiv:1611.07004v32016
  53. AppendiGrade: An XAI-Enhanced Deep Learning Framework for Grading Appendicitis in Ultrasound with Gaussian Blur and Grad-CAM

    Fahad Ahammed, Omar Faruq Shikdar, Navid Zaman +3

    cs.CVcs.LGarXiv:2608.17923v12026
  54. MobileNetV2: Inverted Residuals and Linear Bottlenecks

    Mark Sandler, Andrew Howard, Menglong Zhu +2

    cs.CVarXiv:1801.04381v42018
  55. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks

    Mingxing Tan, Quoc V. Le

    cs.LGcs.CVstat.MLarXiv:1905.11946v52019
  56. YOLOv3: An Incremental Improvement

    Joseph Redmon, Ali Farhadi

    cs.CVarXiv:1804.02767v12018
  57. ENAF: A Multi-Exit Network with an Adaptive Patch Fusion for Large Image Super Resolution

    Duong M. Nguyen, Tuan Nghia Nguyen, Xuan Truong Nguyen

    cs.CVcs.AIeess.IVarXiv:2608.15349v12026
  58. Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification

    Kaiming He, Xiangyu Zhang, Shaoqing Ren +1

    cs.CVcs.AIcs.LGarXiv:1502.01852v12015
  59. Primitive Representation Learning for Unsupervised Dynamic Contrast Enhanced MRI Reconstruction

    Veronika Spieker, Wenqi Huang, Cemre Ariyurek +5

    eess.IVcs.CVcs.LGarXiv:2608.18055v22026
  60. Feature Pyramid Networks for Object Detection

    Tsung-Yi Lin, Piotr Dollár, Ross Girshick +3

    cs.CVarXiv:1612.03144v22016