Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
18,001 to 18,060 of 18,817
Kwai Keye-VL-2.0 Technical Report
Kwai Keye Team, Bin Wen, Changyi Liu +50
cs.CVarXiv:2606.10651v12026Spectral Normalization for Generative Adversarial Networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama +1
cs.LGcs.CVstat.MLarXiv:1802.05957v12018FedCoRe: Target-Adaptive Completion for Missing Modalities in Healthcare Federated Learning
Holger R. Roth, Ziyue Xu, Peter Cnudde
cs.CVcs.AIcs.LGarXiv:2608.18311v12026Coordinate Attention for Efficient Mobile Network Design
Qibin Hou, Daquan Zhou, Jiashi Feng
cs.CVarXiv:2103.02907v12021TTSD-FAR: Test-Time Self-Distillation with Fisher-Anchored Restoration for Missing-Modality Emotion Recognition in LVLMs
Muhammad Haseeb Aslam, Alessandro Koerich, Marco Pedersoli +2
cs.CVcs.AIarXiv:2608.18386v12026Exploring Simple Siamese Representation Learning
Xinlei Chen, Kaiming He
cs.CVcs.LGarXiv:2011.10566v12020Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models
Cheng-Yu Yang, Shao-Yuan Lo, Yu-Lun Liu
cs.CVcs.AIarXiv:2606.12412v12026Deep High-Resolution Representation Learning for Human Pose Estimation
Ke Sun, Bin Xiao, Dong Liu +1
cs.CVarXiv:1902.09212v12019Dynamic Routing Between Capsules
Sara Sabour, Nicholas Frosst, Geoffrey E Hinton
cs.CVarXiv:1710.09829v22017SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Dustin Podell, Zion English, Kyle Lacey +5
cs.CVcs.AIarXiv:2307.01952v12023Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected CRFs
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos +2
cs.CVcs.LGcs.NEarXiv:1412.7062v42014Visual-Prompt Guided Wildlife Instance-Level Recognition
Mufhumudzi Muthivhi, Jiahao Huo, Terence van Zyl +1
cs.CVcs.AIcs.LGarXiv:2608.18246v12026YOLOv10: Real-Time End-to-End Object Detection
Ao Wang, Hui Chen, Lihao Liu +4
cs.CVarXiv:2405.14458v22024LAION-5B: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu +13
cs.CVcs.AIcs.LGarXiv:2210.08402v12022Image Super-Resolution Using Very Deep Residual Channel Attention Networks
Yulun Zhang, Kunpeng Li, Kai Li +3
cs.CVarXiv:1807.02758v22018Improved Baselines with Visual Instruction Tuning
Haotian Liu, Chunyuan Li, Yuheng Li +1
cs.CVcs.AIcs.CLarXiv:2310.03744v22023Adversarial Discriminative Domain Adaptation
Eric Tzeng, Judy Hoffman, Kate Saenko +1
cs.CVarXiv:1702.05464v12017Bound-Aware Per-Organ Recall Risk Control for Multi-Organ CT Segmentation under Clinical Domain Shift
Souraj Adhikary, Negar Chabi, Andre Mastmeyer
cs.CVcs.AIcs.LGarXiv:2608.18193v12026DeepFool: a simple and accurate method to fool deep neural networks
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Pascal Frossard
cs.LGcs.CVarXiv:1511.04599v32015Stacked Hourglass Networks for Human Pose Estimation
Alejandro Newell, Kaiyu Yang, Jia Deng
cs.CVarXiv:1603.06937v22016Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition
Sijie Yan, Yuanjun Xiong, Dahua Lin
cs.CVarXiv:1801.07455v22018Arbitrary Style Transfer in Real-time with Adaptive Instance Normalization
Xun Huang, Serge Belongie
cs.CVarXiv:1703.06868v22017Spectral Networks and Locally Connected Networks on Graphs
Joan Bruna, Wojciech Zaremba, Arthur Szlam +1
cs.LGcs.CVcs.NEarXiv:1312.6203v32013Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks
Jun-Yan Zhu, Taesung Park, Phillip Isola +1
cs.CVarXiv:1703.10593v72017Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Chao Jia, Yinfei Yang, Ye Xia +7
cs.CVcs.CLcs.LGarXiv:2102.05918v22021Recent Advances in Convolutional Neural Networks
Jiuxiang Gu, Zhenhua Wang, Jason Kuen +9
cs.CVcs.LGcs.NEarXiv:1512.07108v62015Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network
Wenzhe Shi, Jose Caballero, Ferenc Huszár +5
cs.CVstat.MLarXiv:1609.05158v22016Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, Hervé Jégou
cs.CVcs.DBcs.DSarXiv:1702.08734v12017ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
Angela Dai, Angel X. Chang, Manolis Savva +3
cs.CVarXiv:1702.04405v22017Learning without Forgetting
Zhizhong Li, Derek Hoiem
cs.CVcs.LGstat.MLarXiv:1606.09282v32016Generalized Intersection over Union: A Metric and A Loss for Bounding Box Regression
Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak +3
cs.CVcs.AIcs.LGarXiv:1902.09630v22019Deep Visual-Semantic Alignments for Generating Image Descriptions
Andrej Karpathy, Li Fei-Fei
cs.CVarXiv:1412.2306v22014YOLOX: Exceeding YOLO Series in 2021
Zheng Ge, Songtao Liu, Feng Wang +2
cs.CVarXiv:2107.08430v22021CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features
Sangdoo Yun, Dongyoon Han, Seong Joon Oh +3
cs.CVcs.LGarXiv:1905.04899v22019TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation
Jieneng Chen, Yongyi Lu, Qihang Yu +6
cs.CVarXiv:2102.04306v12021ShapeNet: An Information-Rich 3D Model Repository
Angel X. Chang, Thomas Funkhouser, Leonidas Guibas +10
cs.GRcs.AIcs.CGarXiv:1512.03012v12015Scalable Diffusion Models with Transformers
William Peebles, Saining Xie
cs.CVcs.LGarXiv:2212.09748v22022Learning Transferable Architectures for Scalable Image Recognition
Barret Zoph, Vijay Vasudevan, Jonathon Shlens +1
cs.CVcs.LGstat.MLarXiv:1707.07012v42017Cascade R-CNN: Delving into High Quality Object Detection
Zhaowei Cai, Nuno Vasconcelos
cs.CVarXiv:1712.00726v12017Flamingo: a Visual Language Model for Few-Shot Learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc +24
cs.CVcs.AIcs.LGarXiv:2204.14198v22022ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks
Qilong Wang, Banggu Wu, Pengfei Zhu +3
cs.CVarXiv:1910.03151v42019Long-term Recurrent Convolutional Networks for Visual Recognition and Description
Jeff Donahue, Lisa Anne Hendricks, Marcus Rohrbach +4
cs.CVarXiv:1411.4389v42014What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?
Alex Kendall, Yarin Gal
cs.CVarXiv:1703.04977v22017ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design
Ningning Ma, Xiangyu Zhang, Hai-Tao Zheng +1
cs.CVarXiv:1807.11164v12018Show and Tell: A Neural Image Caption Generator
Oriol Vinyals, Alexander Toshev, Samy Bengio +1
cs.CVarXiv:1411.4555v22014Zero-Shot Text-to-Image Generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh +5
cs.CVcs.LGarXiv:2102.12092v22021Network In Network
Min Lin, Qiang Chen, Shuicheng Yan
cs.NEcs.CVcs.LGarXiv:1312.4400v32013Deformable Convolutional Networks
Jifeng Dai, Haozhi Qi, Yuwen Xiong +4
cs.CVarXiv:1703.06211v32017CARLA: An Open Urban Driving Simulator
Alexey Dosovitskiy, German Ros, Felipe Codevilla +2
cs.LGcs.AIcs.CVarXiv:1711.03938v12017Accurate Image Super-Resolution Using Very Deep Convolutional Networks
Jiwon Kim, Jung Kwon Lee, Kyoung Mu Lee
cs.CVcs.LGarXiv:1511.04587v22015Path Aggregation Network for Instance Segmentation
Shu Liu, Lu Qi, Haifang Qin +2
cs.CVarXiv:1803.01534v42018Analyzing and Improving the Image Quality of StyleGAN
Tero Karras, Samuli Laine, Miika Aittala +3
cs.CVcs.LGcs.NEarXiv:1912.04958v22019BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Junnan Li, Dongxu Li, Caiming Xiong +1
cs.CVarXiv:2201.12086v22022EfficientDet: Scalable and Efficient Object Detection
Mingxing Tan, Ruoming Pang, Quoc V. Le
cs.CVcs.LGeess.IVarXiv:1911.09070v72019UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Khurram Soomro, Amir Roshan Zamir, Mubarak Shah
cs.CVarXiv:1212.0402v12012Enhanced Deep Residual Networks for Single Image Super-Resolution
Bee Lim, Sanghyun Son, Heewon Kim +2
cs.CVarXiv:1707.02921v12017Adding Conditional Control to Text-to-Image Diffusion Models
Lvmin Zhang, Anyi Rao, Maneesh Agrawala
cs.CVcs.AIcs.GRarXiv:2302.05543v32023Learning both Weights and Connections for Efficient Neural Networks
Song Han, Jeff Pool, John Tran +1
cs.NEcs.CVcs.LGarXiv:1506.02626v32015Spatial Transformer Networks
Max Jaderberg, Karen Simonyan, Andrew Zisserman +1
cs.CVarXiv:1506.02025v32015Deformable DETR: Deformable Transformers for End-to-End Object Detection
Xizhou Zhu, Weijie Su, Lewei Lu +3
cs.CVarXiv:2010.04159v42020