Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
18,121 to 18,180 of 18,866
Surflo: Consistent 3D Surface Flow Model with Global State
Antoine Guédon, Shu Nakamura, Nicolas Dufour +3
cs.CVarXiv:2606.13644v12026Emerging Properties in Self-Supervised Vision Transformers
Mathilde Caron, Hugo Touvron, Ishan Misra +4
cs.CVarXiv:2104.14294v22021MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffold
Yang Zhou, Ziheng Wang, Yuqin Lu +4
cs.CVarXiv:2606.13376v22026Searching for MobileNetV3
Andrew Howard, Mark Sandler, Grace Chu +9
cs.CVarXiv:1905.02244v52019Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze +3
cs.CVarXiv:2012.12877v22020Hierarchical Data Selection via Manifold Coverage and Sparse Feature Coverage in LLM Post-training
Peng Sun, Yi Yang, Antong Zhang +7
cs.LGcs.CVarXiv:2608.16927v12026Image Super-Resolution Using Deep Convolutional Networks
Chao Dong, Chen Change Loy, Kaiming He +1
cs.CVcs.NEarXiv:1501.00092v32014UNet++: A Nested U-Net Architecture for Medical Image Segmentation
Zongwei Zhou, Md Mahfuzur Rahman Siddiquee, Nima Tajbakhsh +1
cs.CVcs.LGeess.IVarXiv:1807.10165v12018YOLO9000: Better, Faster, Stronger
Joseph Redmon, Ali Farhadi
cs.CVarXiv:1612.08242v12016MagViT: Interpretable Multi-Magnification Transformers with Patient-Level Model Selection for Breast Histopathology
Nabil Ashab, Soumit Kumar Kundu, Saif Mahmud Parvez +3
eess.IVcs.CVcs.LGarXiv:2608.16959v12026Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
Joao Carreira, Andrew Zisserman
cs.CVcs.LGarXiv:1705.07750v32017Deep Learning Face Attributes in the Wild
Ziwei Liu, Ping Luo, Xiaogang Wang +1
cs.CVarXiv:1411.7766v32014Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting
Xingjian Shi, Zhourong Chen, Hao Wang +3
cs.CVarXiv:1506.04214v22015VideoMDM: Towards 3D Human Motion Generation From 2D Supervision
Amir Mann, Gal Michael Harari, Merav Keidar +1
cs.LGcs.CVarXiv:2606.13364v12026InterleaveThinker: Reinforcing Agentic Interleaved Generation
Dian Zheng, Harry Lee, Manyuan Zhang +4
cs.CVarXiv:2606.13679v22026Mask R-CNN
Kaiming He, Georgia Gkioxari, Piotr Dollár +1
cs.CVarXiv:1703.06870v32017Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
Song Han, Huizi Mao, William J. Dally
cs.CVcs.NEarXiv:1510.00149v52015V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation
Fausto Milletari, Nassir Navab, Seyed-Ahmad Ahmadi
cs.CVarXiv:1606.04797v12016Identity Mappings in Deep Residual Networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren +1
cs.CVcs.LGarXiv:1603.05027v32016Lymphocyte Mimicry Correction via Region-Level Tissue Reasoning and Unbalanced Optimal Transport
Xiang Li, Yuqi Wang, Casey C. Heirman +2
cs.CVcs.LGarXiv:2608.17151v12026Rethinking Atrous Convolution for Semantic Image Segmentation
Liang-Chieh Chen, George Papandreou, Florian Schroff +1
cs.CVarXiv:1706.05587v32017Non-local Neural Networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta +1
cs.CVarXiv:1711.07971v32017Visual Instruction Tuning
Haotian Liu, Chunyuan Li, Qingyang Wu +1
cs.CVcs.AIcs.CLarXiv:2304.08485v22023Learning Deep Features for Discriminative Localization
Bolei Zhou, Aditya Khosla, Agata Lapedriza +2
cs.CVarXiv:1512.04150v12015Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
Kelvin Xu, Jimmy Ba, Ryan Kiros +5
cs.LGcs.CVarXiv:1502.03044v32015YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors
Chien-Yao Wang, Alexey Bochkovskiy, Hong-Yuan Mark Liao
cs.CVarXiv:2207.02696v12022Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
Han Xiao, Kashif Rasul, Roland Vollgraf
cs.LGcs.CVstat.MLarXiv:1708.07747v22017Track2View: 4D-Consistent Camera-Controlled Video Generation via Paired 3D Point Tracks
Feng Qiao, Zhaochong An, Zhexiao Xiong +2
cs.CVarXiv:2606.15534v12026Perceptual Losses for Real-Time Style Transfer and Super-Resolution
Justin Johnson, Alexandre Alahi, Li Fei-Fei
cs.CVcs.LGarXiv:1603.08155v12016Conditional Generative Adversarial Nets
Mehdi Mirza, Simon Osindero
cs.LGcs.AIcs.CVarXiv:1411.1784v12014Aggregated Residual Transformations for Deep Neural Networks
Saining Xie, Ross Girshick, Piotr Dollár +2
cs.CVarXiv:1611.05431v22016Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren +1
cs.CVarXiv:1406.4729v42014Caffe: Convolutional Architecture for Fast Feature Embedding
Yangqing Jia, Evan Shelhamer, Jeff Donahue +5
cs.CVcs.LGcs.NEarXiv:1408.5093v12014FaceNet: A Unified Embedding for Face Recognition and Clustering
Florian Schroff, Dmitry Kalenichenko, James Philbin
cs.CVarXiv:1503.03832v32015Diffusion Models Beat GANs on Image Synthesis
Prafulla Dhariwal, Alex Nichol
cs.LGcs.AIcs.CVarXiv:2105.05233v42021SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation
Vijay Badrinarayanan, Alex Kendall, Roberto Cipolla
cs.CVcs.LGcs.NEarXiv:1511.00561v32015End-to-End Object Detection with Transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve +3
cs.CVarXiv:2005.12872v32020The Unreasonable Effectiveness of Deep Features as a Perceptual Metric
Richard Zhang, Phillip Isola, Alexei A. Efros +2
cs.CVcs.GRarXiv:1801.03924v22018Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network
Christian Ledig, Lucas Theis, Ferenc Huszar +8
cs.CVstat.MLarXiv:1609.04802v52016Masked Autoencoders Are Scalable Vision Learners
Kaiming He, Xinlei Chen, Saining Xie +3
cs.CVarXiv:2111.06377v32021A Survey on Deep Learning in Medical Image Analysis
Geert Litjens, Thijs Kooi, Babak Ehteshami Bejnordi +6
cs.CVarXiv:1702.05747v22017Optimize Your Sampling: Tuned Diffusion Sampling with Bayesian Optimization
Travis Zhang, Christian Belardi, Justin Lovelace +4
cs.LGcs.CVarXiv:2608.18040v12026Pyramid Scene Parsing Network
Hengshuang Zhao, Jianping Shi, Xiaojuan Qi +2
cs.CVarXiv:1612.01105v22016Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
Alec Radford, Luke Metz, Soumith Chintala
cs.LGcs.CVarXiv:1511.06434v22015Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke +1
cs.CVarXiv:1602.07261v22016YOLOv4: Optimal Speed and Accuracy of Object Detection
Alexey Bochkovskiy, Chien-Yao Wang, Hong-Yuan Mark Liao
cs.CVeess.IVarXiv:2004.10934v12020Visualizing and Understanding Convolutional Networks
Matthew D Zeiler, Rob Fergus
cs.CVarXiv:1311.2901v32013Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever +4
cs.CVcs.LGcs.NEarXiv:1312.6199v42013The Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri +558
cs.AIcs.CLcs.CVarXiv:2407.21783v32024Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation
Liang-Chieh Chen, Yukun Zhu, George Papandreou +2
cs.CVarXiv:1802.02611v32018UAV Video Deblurring via Motion-Aware Diffusion: A Path to Robust Target Detection
Zhiqiang Hu, Shouren Huang, Masatoshi Ishikawa
cs.CVcs.AIcs.ROarXiv:2608.15259v12026Image-to-Image Translation with Conditional Adversarial Networks
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou +1
cs.CVarXiv:1611.07004v32016AppendiGrade: An XAI-Enhanced Deep Learning Framework for Grading Appendicitis in Ultrasound with Gaussian Blur and Grad-CAM
Fahad Ahammed, Omar Faruq Shikdar, Navid Zaman +3
cs.CVcs.LGarXiv:2608.17923v12026MobileNetV2: Inverted Residuals and Linear Bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu +2
cs.CVarXiv:1801.04381v42018EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
Mingxing Tan, Quoc V. Le
cs.LGcs.CVstat.MLarXiv:1905.11946v52019YOLOv3: An Incremental Improvement
Joseph Redmon, Ali Farhadi
cs.CVarXiv:1804.02767v12018ENAF: A Multi-Exit Network with an Adaptive Patch Fusion for Large Image Super Resolution
Duong M. Nguyen, Tuan Nghia Nguyen, Xuan Truong Nguyen
cs.CVcs.AIeess.IVarXiv:2608.15349v12026Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren +1
cs.CVcs.AIcs.LGarXiv:1502.01852v12015Primitive Representation Learning for Unsupervised Dynamic Contrast Enhanced MRI Reconstruction
Veronika Spieker, Wenqi Huang, Cemre Ariyurek +5
eess.IVcs.CVcs.LGarXiv:2608.18055v22026Feature Pyramid Networks for Object Detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick +3
cs.CVarXiv:1612.03144v22016