Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
1,141 to 1,200 of 18,830
Face Morphing Attack Generation & Detection: A Comprehensive Survey
Sushma Venkatesh, Raghavendra Ramachandra, Kiran Raja +1
cs.CVcs.CRcs.CYarXiv:2011.02045v12020SAM on Medical Images: A Comprehensive Study on Three Prompt Modes
Dongjie Cheng, Ziyuan Qin, Zekun Jiang +3
cs.CVcs.AIarXiv:2305.00035v12023Deep supervision with additional labels for retinal vessel segmentation task
Yishuo Zhang, Albert C. S. Chung
cs.CVarXiv:1806.02132v32018Classification of EEG-Based Brain Connectivity Networks in Schizophrenia Using a Multi-Domain Connectome Convolutional Neural Network
Chun-Ren Phang, Chee-Ming Ting, Fuad Noman +1
cs.LGcs.CVq-bio.NCarXiv:1903.08858v12019GAN Memory with No Forgetting
Yulai Cong, Miaoyun Zhao, Jianqiao Li +2
cs.CVcs.LGarXiv:2006.07543v22020DeXpression: Deep Convolutional Neural Network for Expression Recognition
Peter Burkert, Felix Trier, Muhammad Zeshan Afzal +2
cs.CVcs.LGarXiv:1509.05371v22015ProtoPShare: Prototype Sharing for Interpretable Image Classification and Similarity Discovery
Dawid Rymarczyk, Łukasz Struski, Jacek Tabor +1
cs.CVcs.AIcs.LGarXiv:2011.14340v12020Beyond accuracy: quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistency
Robert Geirhos, Kristof Meding, Felix A. Wichmann
cs.CVcs.LGq-bio.NCarXiv:2006.16736v32020CALM: Conditional Adversarial Latent Models for Directable Virtual Characters
Chen Tessler, Yoni Kasten, Yunrong Guo +3
cs.CVcs.AIcs.ROarXiv:2305.02195v12023Hybrid Convolutional and Attention Network for Hyperspectral Image Denoising
Shuai Hu, Feng Gao, Xiaowei Zhou +2
eess.IVcs.CVarXiv:2403.10067v12024Connecting Look and Feel: Associating the visual and tactile properties of physical materials
Wenzhen Yuan, Shaoxiong Wang, Siyuan Dong +1
cs.CVarXiv:1704.03822v12017Interactive Sketch & Fill: Multiclass Sketch-to-Image Translation
Arnab Ghosh, Richard Zhang, Puneet K. Dokania +4
cs.CVcs.LGeess.IVarXiv:1909.11081v22019SIZER: A Dataset and Model for Parsing 3D Clothing and Learning Size Sensitive 3D Clothing
Garvita Tiwari, Bharat Lal Bhatnagar, Tony Tung +1
cs.CVarXiv:2007.11610v12020DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion
Wenqiang Sun, Shuo Chen, Fangfu Liu +4
cs.CVcs.AIcs.GRarXiv:2411.04928v12024A Comprehensive Analysis of Deep Learning Based Representation for Face Recognition
Mostafa Mehdipour Ghazi, Hazim Kemal Ekenel
cs.CVarXiv:1606.02894v12016Ego-Pose Estimation and Forecasting as Real-Time PD Control
Ye Yuan, Kris Kitani
cs.CVcs.AIcs.LGarXiv:1906.03173v22019UniTAB: Unifying Text and Box Outputs for Grounded Vision-Language Modeling
Zhengyuan Yang, Zhe Gan, Jianfeng Wang +5
cs.CVarXiv:2111.12085v22021Multimodal Masked Autoencoders Learn Transferable Representations
Xinyang Geng, Hao Liu, Lisa Lee +3
cs.CVarXiv:2205.14204v32022FreeSeg: Unified, Universal and Open-Vocabulary Image Segmentation
Jie Qin, Jie Wu, Pengxiang Yan +8
cs.CVarXiv:2303.17225v12023To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Junke Wang, Lingchen Meng, Zejia Weng +3
cs.CVarXiv:2311.07574v22023SwinLSTM:Improving Spatiotemporal Prediction Accuracy using Swin Transformer and LSTM
Song Tang, Chuang Li, Pu Zhang +1
cs.CVcs.AIarXiv:2308.09891v22023SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Bohao Li, Yuying Ge, Yi Chen +3
cs.CVarXiv:2404.16790v12024USB: A Unified Semi-supervised Learning Benchmark for Classification
Yidong Wang, Hao Chen, Yue Fan +19
cs.LGcs.AIcs.CVarXiv:2208.07204v22022A Survey of Automatic Facial Micro-expression Analysis: Databases, Methods and Challenges
Yee-Hui Oh, John See, Anh Cat Le Ngo +2
cs.CVcs.MMarXiv:1806.05781v12018Imagination improves Multimodal Translation
Desmond Elliott, Ákos Kádár
cs.CLcs.CVarXiv:1705.04350v22017Cooperative Training of Descriptor and Generator Networks
Jianwen Xie, Yang Lu, Ruiqi Gao +2
stat.MLcs.CVarXiv:1609.09408v32016RiVaT-Fuse: Reliability-Calibrated Variational Tensor Fusion for Multimodal Prediction under Modality Uncertainty
Yingfan Xu, Tieming Liu, Ye Liang +1
cs.LGcs.CVarXiv:2609.10798v12026Unbiased Teacher v2: Semi-supervised Object Detection for Anchor-free and Anchor-based Detectors
Yen-Cheng Liu, Chih-Yao Ma, Zsolt Kira
cs.CVcs.LGarXiv:2206.09500v12022MultiSports: A Multi-Person Video Dataset of Spatio-Temporally Localized Sports Actions
Yixuan Li, Lei Chen, Runyu He +3
cs.CVarXiv:2105.07404v22021Multi-Cue Zero-Shot Learning with Strong Supervision
Zeynep Akata, Mateusz Malinowski, Mario Fritz +1
cs.CVarXiv:1603.08754v12016Language Embedded Radiance Fields for Zero-Shot Task-Oriented Grasping
Adam Rashid, Satvik Sharma, Chung Min Kim +4
cs.ROcs.CVarXiv:2309.07970v22023PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI
Yandan Yang, Baoxiong Jia, Peiyuan Zhi +1
cs.CVcs.AIcs.LGarXiv:2404.09465v22024Deep Generative Learning via Schrödinger Bridge
Gefei Wang, Yuling Jiao, Qian Xu +2
cs.LGcs.CVarXiv:2106.10410v22021Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising
Fu-Yun Wang, Wenshuo Chen, Guanglu Song +3
cs.CVarXiv:2305.18264v12023GenArtist: Multimodal LLM as an Agent for Unified Image Generation and Editing
Zhenyu Wang, Aoxue Li, Zhenguo Li +1
cs.CVarXiv:2407.05600v22024Real-time Semantic Segmentation with Fast Attention
Ping Hu, Federico Perazzi, Fabian Caba Heilbron +4
cs.CVcs.MMcs.ROarXiv:2007.03815v22020Mind Reader: Reconstructing complex images from brain activities
Sikun Lin, Thomas Sprague, Ambuj K Singh
q-bio.NCcs.CVcs.HCarXiv:2210.01769v12022Multimodal Memory Modelling for Video Captioning
Junbo Wang, Wei Wang, Yan Huang +2
cs.CVarXiv:1611.05592v12016PCPL: Predicate-Correlation Perception Learning for Unbiased Scene Graph Generation
Shaotian Yan, Chen Shen, Zhongming Jin +4
cs.CVarXiv:2009.00893v12020IDM: An Intermediate Domain Module for Domain Adaptive Person Re-ID
Yongxing Dai, Jun Liu, Yifan Sun +3
cs.CVarXiv:2108.02413v12021MindTopo: Can Foundation Models Reason in Topological Space?
Yunfei Ge, Anbang Liu, Qineng Wang +9
cs.AIcs.CLcs.CVarXiv:2609.11900v12026ModaNet: A Large-Scale Street Fashion Dataset with Polygon Annotations
Shuai Zheng, Fan Yang, M. Hadi Kiapour +1
cs.CVarXiv:1807.01394v42018Towards the Generalization of Contrastive Self-Supervised Learning
Weiran Huang, Mingyang Yi, Xuyang Zhao +1
cs.LGcs.AIcs.CVarXiv:2111.00743v42021Video as Conditional Graph Hierarchy for Multi-Granular Question Answering
Junbin Xiao, Angela Yao, Zhiyuan Liu +3
cs.CVcs.AIcs.MMarXiv:2112.06197v22021GeoDA: a geometric framework for black-box adversarial attacks
Ali Rahmati, Seyed-Mohsen Moosavi-Dezfooli, Pascal Frossard +1
cs.CVcs.CRcs.LGarXiv:2003.06468v12020Context Decoupling Augmentation for Weakly Supervised Semantic Segmentation
Yukun Su, Ruizhou Sun, Guosheng Lin +1
cs.CVarXiv:2103.01795v22021TALL: Thumbnail Layout for Deepfake Video Detection
Yuting Xu, Jian Liang, Gengyun Jia +3
cs.CVarXiv:2307.07494v32023Connecting Gaze, Scene, and Attention: Generalized Attention Estimation via Joint Modeling of Gaze and Scene Saliency
Eunji Chong, Nataniel Ruiz, Yongxin Wang +3
cs.CVarXiv:1807.10437v12018Anti-UAV: A Large Multi-Modal Benchmark for UAV Tracking
Nan Jiang, Kuiran Wang, Xiaoke Peng +7
cs.CVarXiv:2101.08466v32021HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation
Xuan Ju, Ailing Zeng, Chenchen Zhao +3
cs.CVarXiv:2304.04269v12023Automated forest inventory: analysis of high-density airborne LiDAR point clouds with 3D deep learning
Binbin Xiang, Maciej Wielgosz, Theodora Kontogianni +4
cs.CVarXiv:2312.15084v22023Improving Multimodal Datasets with Image Captioning
Thao Nguyen, Samir Yitzhak Gadre, Gabriel Ilharco +2
cs.LGcs.CVarXiv:2307.10350v22023Hyperbolic Image Segmentation
Mina GhadimiAtigh, Julian Schoep, Erman Acar +2
cs.CVarXiv:2203.05898v12022Secure Face Matching Using Fully Homomorphic Encryption
Vishnu Naresh Boddeti
cs.CVarXiv:1805.00577v22018Prediction and Localization of Student Engagement in the Wild
Amanjot Kaur, Aamir Mustafa, Love Mehta +1
cs.CVarXiv:1804.00858v42018M3-Former: Multimodal Transformer with Mixture-of-Experts for Long-Term Vessel Trajectory Prediction
Wenzhe Jin, Haina Tang
cs.LGcs.CVarXiv:2609.10559v12026Diverse Trajectory Forecasting with Determinantal Point Processes
Ye Yuan, Kris Kitani
cs.CVcs.LGcs.ROarXiv:1907.04967v22019HOI Analysis: Integrating and Decomposing Human-Object Interaction
Yong-Lu Li, Xinpeng Liu, Xiaoqian Wu +2
cs.CVcs.LGeess.IVarXiv:2010.16219v22020AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
Wenhao Chai, Enxin Song, Yilun Du +6
cs.CVarXiv:2410.03051v42024Constrained R-CNN: A general image manipulation detection model
Chao Yang, Huizhou Li, Fangting Lin +2
cs.CVcs.MMarXiv:1911.08217v32019