Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
18,061 to 18,120 of 18,837
ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks
Qilong Wang, Banggu Wu, Pengfei Zhu +3
cs.CVarXiv:1910.03151v42019Long-term Recurrent Convolutional Networks for Visual Recognition and Description
Jeff Donahue, Lisa Anne Hendricks, Marcus Rohrbach +4
cs.CVarXiv:1411.4389v42014What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?
Alex Kendall, Yarin Gal
cs.CVarXiv:1703.04977v22017ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design
Ningning Ma, Xiangyu Zhang, Hai-Tao Zheng +1
cs.CVarXiv:1807.11164v12018Show and Tell: A Neural Image Caption Generator
Oriol Vinyals, Alexander Toshev, Samy Bengio +1
cs.CVarXiv:1411.4555v22014Zero-Shot Text-to-Image Generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh +5
cs.CVcs.LGarXiv:2102.12092v22021Network In Network
Min Lin, Qiang Chen, Shuicheng Yan
cs.NEcs.CVcs.LGarXiv:1312.4400v32013Deformable Convolutional Networks
Jifeng Dai, Haozhi Qi, Yuwen Xiong +4
cs.CVarXiv:1703.06211v32017CARLA: An Open Urban Driving Simulator
Alexey Dosovitskiy, German Ros, Felipe Codevilla +2
cs.LGcs.AIcs.CVarXiv:1711.03938v12017Accurate Image Super-Resolution Using Very Deep Convolutional Networks
Jiwon Kim, Jung Kwon Lee, Kyoung Mu Lee
cs.CVcs.LGarXiv:1511.04587v22015Path Aggregation Network for Instance Segmentation
Shu Liu, Lu Qi, Haifang Qin +2
cs.CVarXiv:1803.01534v42018Analyzing and Improving the Image Quality of StyleGAN
Tero Karras, Samuli Laine, Miika Aittala +3
cs.CVcs.LGcs.NEarXiv:1912.04958v22019BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Junnan Li, Dongxu Li, Caiming Xiong +1
cs.CVarXiv:2201.12086v22022EfficientDet: Scalable and Efficient Object Detection
Mingxing Tan, Ruoming Pang, Quoc V. Le
cs.CVcs.LGeess.IVarXiv:1911.09070v72019UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Khurram Soomro, Amir Roshan Zamir, Mubarak Shah
cs.CVarXiv:1212.0402v12012Enhanced Deep Residual Networks for Single Image Super-Resolution
Bee Lim, Sanghyun Son, Heewon Kim +2
cs.CVarXiv:1707.02921v12017Adding Conditional Control to Text-to-Image Diffusion Models
Lvmin Zhang, Anyi Rao, Maneesh Agrawala
cs.CVcs.AIcs.GRarXiv:2302.05543v32023Learning both Weights and Connections for Efficient Neural Networks
Song Han, Jeff Pool, John Tran +1
cs.NEcs.CVcs.LGarXiv:1506.02626v32015Spatial Transformer Networks
Max Jaderberg, Karen Simonyan, Andrew Zisserman +1
cs.CVarXiv:1506.02025v32015Deformable DETR: Deformable Transformers for End-to-End Object Detection
Xizhou Zhu, Weijie Su, Lewei Lu +3
cs.CVarXiv:2010.04159v420203D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation
Özgün Çiçek, Ahmed Abdulkadir, Soeren S. Lienkamp +2
cs.CVarXiv:1606.06650v12016Two-Stream Convolutional Networks for Action Recognition in Videos
Karen Simonyan, Andrew Zisserman
cs.CVarXiv:1406.2199v22014Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
Karen Simonyan, Andrea Vedaldi, Andrew Zisserman
cs.CVarXiv:1312.6034v22013SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5MB model size
Forrest N. Iandola, Song Han, Matthew W. Moskewicz +3
cs.CVcs.AIarXiv:1602.07360v42016ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices
Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin +1
cs.CVarXiv:1707.01083v22017BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Junnan Li, Dongxu Li, Silvio Savarese +1
cs.CVarXiv:2301.12597v32023Wide Residual Networks
Sergey Zagoruyko, Nikos Komodakis
cs.CVcs.LGcs.NEarXiv:1605.07146v42016A ConvNet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu +3
cs.CVarXiv:2201.03545v22022SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers
Enze Xie, Wenhai Wang, Zhiding Yu +3
cs.CVcs.LGarXiv:2105.15203v32021Hierarchical Text-Conditional Image Generation with CLIP Latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol +2
cs.CVarXiv:2204.06125v12022iMaC: Translating Actions into Motion and Contact Images for Embodied World Models
Zhenyu Wu, Xiuwei Xu, Yukun Zhou +8
cs.ROcs.CVarXiv:2606.09813v12026Surflo: Consistent 3D Surface Flow Model with Global State
Antoine Guédon, Shu Nakamura, Nicolas Dufour +3
cs.CVarXiv:2606.13644v12026Emerging Properties in Self-Supervised Vision Transformers
Mathilde Caron, Hugo Touvron, Ishan Misra +4
cs.CVarXiv:2104.14294v22021MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffold
Yang Zhou, Ziheng Wang, Yuqin Lu +4
cs.CVarXiv:2606.13376v22026Searching for MobileNetV3
Andrew Howard, Mark Sandler, Grace Chu +9
cs.CVarXiv:1905.02244v52019Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze +3
cs.CVarXiv:2012.12877v22020Hierarchical Data Selection via Manifold Coverage and Sparse Feature Coverage in LLM Post-training
Peng Sun, Yi Yang, Antong Zhang +7
cs.LGcs.CVarXiv:2608.16927v12026Image Super-Resolution Using Deep Convolutional Networks
Chao Dong, Chen Change Loy, Kaiming He +1
cs.CVcs.NEarXiv:1501.00092v32014UNet++: A Nested U-Net Architecture for Medical Image Segmentation
Zongwei Zhou, Md Mahfuzur Rahman Siddiquee, Nima Tajbakhsh +1
cs.CVcs.LGeess.IVarXiv:1807.10165v12018YOLO9000: Better, Faster, Stronger
Joseph Redmon, Ali Farhadi
cs.CVarXiv:1612.08242v12016MagViT: Interpretable Multi-Magnification Transformers with Patient-Level Model Selection for Breast Histopathology
Nabil Ashab, Soumit Kumar Kundu, Saif Mahmud Parvez +3
eess.IVcs.CVcs.LGarXiv:2608.16959v12026Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
Joao Carreira, Andrew Zisserman
cs.CVcs.LGarXiv:1705.07750v32017Deep Learning Face Attributes in the Wild
Ziwei Liu, Ping Luo, Xiaogang Wang +1
cs.CVarXiv:1411.7766v32014Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting
Xingjian Shi, Zhourong Chen, Hao Wang +3
cs.CVarXiv:1506.04214v22015VideoMDM: Towards 3D Human Motion Generation From 2D Supervision
Amir Mann, Gal Michael Harari, Merav Keidar +1
cs.LGcs.CVarXiv:2606.13364v12026InterleaveThinker: Reinforcing Agentic Interleaved Generation
Dian Zheng, Harry Lee, Manyuan Zhang +4
cs.CVarXiv:2606.13679v22026Mask R-CNN
Kaiming He, Georgia Gkioxari, Piotr Dollár +1
cs.CVarXiv:1703.06870v32017Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
Song Han, Huizi Mao, William J. Dally
cs.CVcs.NEarXiv:1510.00149v52015V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation
Fausto Milletari, Nassir Navab, Seyed-Ahmad Ahmadi
cs.CVarXiv:1606.04797v12016Identity Mappings in Deep Residual Networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren +1
cs.CVcs.LGarXiv:1603.05027v32016Lymphocyte Mimicry Correction via Region-Level Tissue Reasoning and Unbalanced Optimal Transport
Xiang Li, Yuqi Wang, Casey C. Heirman +2
cs.CVcs.LGarXiv:2608.17151v12026Rethinking Atrous Convolution for Semantic Image Segmentation
Liang-Chieh Chen, George Papandreou, Florian Schroff +1
cs.CVarXiv:1706.05587v32017Non-local Neural Networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta +1
cs.CVarXiv:1711.07971v32017Visual Instruction Tuning
Haotian Liu, Chunyuan Li, Qingyang Wu +1
cs.CVcs.AIcs.CLarXiv:2304.08485v22023Learning Deep Features for Discriminative Localization
Bolei Zhou, Aditya Khosla, Agata Lapedriza +2
cs.CVarXiv:1512.04150v12015Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
Kelvin Xu, Jimmy Ba, Ryan Kiros +5
cs.LGcs.CVarXiv:1502.03044v32015YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors
Chien-Yao Wang, Alexey Bochkovskiy, Hong-Yuan Mark Liao
cs.CVarXiv:2207.02696v12022Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
Han Xiao, Kashif Rasul, Roland Vollgraf
cs.LGcs.CVstat.MLarXiv:1708.07747v22017Track2View: 4D-Consistent Camera-Controlled Video Generation via Paired 3D Point Tracks
Feng Qiao, Zhaochong An, Zhexiao Xiong +2
cs.CVarXiv:2606.15534v12026Perceptual Losses for Real-Time Style Transfer and Super-Resolution
Justin Johnson, Alexandre Alahi, Li Fei-Fei
cs.CVcs.LGarXiv:1603.08155v12016