Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1,141 to 1,200 of 18,830

  1. Face Morphing Attack Generation & Detection: A Comprehensive Survey

    Sushma Venkatesh, Raghavendra Ramachandra, Kiran Raja +1

    cs.CVcs.CRcs.CYarXiv:2011.02045v12020
  2. SAM on Medical Images: A Comprehensive Study on Three Prompt Modes

    Dongjie Cheng, Ziyuan Qin, Zekun Jiang +3

    cs.CVcs.AIarXiv:2305.00035v12023
  3. Deep supervision with additional labels for retinal vessel segmentation task

    Yishuo Zhang, Albert C. S. Chung

    cs.CVarXiv:1806.02132v32018
  4. Classification of EEG-Based Brain Connectivity Networks in Schizophrenia Using a Multi-Domain Connectome Convolutional Neural Network

    Chun-Ren Phang, Chee-Ming Ting, Fuad Noman +1

    cs.LGcs.CVq-bio.NCarXiv:1903.08858v12019
  5. GAN Memory with No Forgetting

    Yulai Cong, Miaoyun Zhao, Jianqiao Li +2

    cs.CVcs.LGarXiv:2006.07543v22020
  6. DeXpression: Deep Convolutional Neural Network for Expression Recognition

    Peter Burkert, Felix Trier, Muhammad Zeshan Afzal +2

    cs.CVcs.LGarXiv:1509.05371v22015
  7. ProtoPShare: Prototype Sharing for Interpretable Image Classification and Similarity Discovery

    Dawid Rymarczyk, Łukasz Struski, Jacek Tabor +1

    cs.CVcs.AIcs.LGarXiv:2011.14340v12020
  8. Beyond accuracy: quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistency

    Robert Geirhos, Kristof Meding, Felix A. Wichmann

    cs.CVcs.LGq-bio.NCarXiv:2006.16736v32020
  9. CALM: Conditional Adversarial Latent Models for Directable Virtual Characters

    Chen Tessler, Yoni Kasten, Yunrong Guo +3

    cs.CVcs.AIcs.ROarXiv:2305.02195v12023
  10. Hybrid Convolutional and Attention Network for Hyperspectral Image Denoising

    Shuai Hu, Feng Gao, Xiaowei Zhou +2

    eess.IVcs.CVarXiv:2403.10067v12024
  11. Connecting Look and Feel: Associating the visual and tactile properties of physical materials

    Wenzhen Yuan, Shaoxiong Wang, Siyuan Dong +1

    cs.CVarXiv:1704.03822v12017
  12. Interactive Sketch & Fill: Multiclass Sketch-to-Image Translation

    Arnab Ghosh, Richard Zhang, Puneet K. Dokania +4

    cs.CVcs.LGeess.IVarXiv:1909.11081v22019
  13. SIZER: A Dataset and Model for Parsing 3D Clothing and Learning Size Sensitive 3D Clothing

    Garvita Tiwari, Bharat Lal Bhatnagar, Tony Tung +1

    cs.CVarXiv:2007.11610v12020
  14. DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion

    Wenqiang Sun, Shuo Chen, Fangfu Liu +4

    cs.CVcs.AIcs.GRarXiv:2411.04928v12024
  15. A Comprehensive Analysis of Deep Learning Based Representation for Face Recognition

    Mostafa Mehdipour Ghazi, Hazim Kemal Ekenel

    cs.CVarXiv:1606.02894v12016
  16. Ego-Pose Estimation and Forecasting as Real-Time PD Control

    Ye Yuan, Kris Kitani

    cs.CVcs.AIcs.LGarXiv:1906.03173v22019
  17. UniTAB: Unifying Text and Box Outputs for Grounded Vision-Language Modeling

    Zhengyuan Yang, Zhe Gan, Jianfeng Wang +5

    cs.CVarXiv:2111.12085v22021
  18. Multimodal Masked Autoencoders Learn Transferable Representations

    Xinyang Geng, Hao Liu, Lisa Lee +3

    cs.CVarXiv:2205.14204v32022
  19. FreeSeg: Unified, Universal and Open-Vocabulary Image Segmentation

    Jie Qin, Jie Wu, Pengxiang Yan +8

    cs.CVarXiv:2303.17225v12023
  20. To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

    Junke Wang, Lingchen Meng, Zejia Weng +3

    cs.CVarXiv:2311.07574v22023
  21. SwinLSTM:Improving Spatiotemporal Prediction Accuracy using Swin Transformer and LSTM

    Song Tang, Chuang Li, Pu Zhang +1

    cs.CVcs.AIarXiv:2308.09891v22023
  22. SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

    Bohao Li, Yuying Ge, Yi Chen +3

    cs.CVarXiv:2404.16790v12024
  23. USB: A Unified Semi-supervised Learning Benchmark for Classification

    Yidong Wang, Hao Chen, Yue Fan +19

    cs.LGcs.AIcs.CVarXiv:2208.07204v22022
  24. A Survey of Automatic Facial Micro-expression Analysis: Databases, Methods and Challenges

    Yee-Hui Oh, John See, Anh Cat Le Ngo +2

    cs.CVcs.MMarXiv:1806.05781v12018
  25. Imagination improves Multimodal Translation

    Desmond Elliott, Ákos Kádár

    cs.CLcs.CVarXiv:1705.04350v22017
  26. Cooperative Training of Descriptor and Generator Networks

    Jianwen Xie, Yang Lu, Ruiqi Gao +2

    stat.MLcs.CVarXiv:1609.09408v32016
  27. RiVaT-Fuse: Reliability-Calibrated Variational Tensor Fusion for Multimodal Prediction under Modality Uncertainty

    Yingfan Xu, Tieming Liu, Ye Liang +1

    cs.LGcs.CVarXiv:2609.10798v12026
  28. Unbiased Teacher v2: Semi-supervised Object Detection for Anchor-free and Anchor-based Detectors

    Yen-Cheng Liu, Chih-Yao Ma, Zsolt Kira

    cs.CVcs.LGarXiv:2206.09500v12022
  29. MultiSports: A Multi-Person Video Dataset of Spatio-Temporally Localized Sports Actions

    Yixuan Li, Lei Chen, Runyu He +3

    cs.CVarXiv:2105.07404v22021
  30. Multi-Cue Zero-Shot Learning with Strong Supervision

    Zeynep Akata, Mateusz Malinowski, Mario Fritz +1

    cs.CVarXiv:1603.08754v12016
  31. Language Embedded Radiance Fields for Zero-Shot Task-Oriented Grasping

    Adam Rashid, Satvik Sharma, Chung Min Kim +4

    cs.ROcs.CVarXiv:2309.07970v22023
  32. PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI

    Yandan Yang, Baoxiong Jia, Peiyuan Zhi +1

    cs.CVcs.AIcs.LGarXiv:2404.09465v22024
  33. Deep Generative Learning via Schrödinger Bridge

    Gefei Wang, Yuling Jiao, Qian Xu +2

    cs.LGcs.CVarXiv:2106.10410v22021
  34. Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising

    Fu-Yun Wang, Wenshuo Chen, Guanglu Song +3

    cs.CVarXiv:2305.18264v12023
  35. GenArtist: Multimodal LLM as an Agent for Unified Image Generation and Editing

    Zhenyu Wang, Aoxue Li, Zhenguo Li +1

    cs.CVarXiv:2407.05600v22024
  36. Real-time Semantic Segmentation with Fast Attention

    Ping Hu, Federico Perazzi, Fabian Caba Heilbron +4

    cs.CVcs.MMcs.ROarXiv:2007.03815v22020
  37. Mind Reader: Reconstructing complex images from brain activities

    Sikun Lin, Thomas Sprague, Ambuj K Singh

    q-bio.NCcs.CVcs.HCarXiv:2210.01769v12022
  38. Multimodal Memory Modelling for Video Captioning

    Junbo Wang, Wei Wang, Yan Huang +2

    cs.CVarXiv:1611.05592v12016
  39. PCPL: Predicate-Correlation Perception Learning for Unbiased Scene Graph Generation

    Shaotian Yan, Chen Shen, Zhongming Jin +4

    cs.CVarXiv:2009.00893v12020
  40. IDM: An Intermediate Domain Module for Domain Adaptive Person Re-ID

    Yongxing Dai, Jun Liu, Yifan Sun +3

    cs.CVarXiv:2108.02413v12021
  41. MindTopo: Can Foundation Models Reason in Topological Space?

    Yunfei Ge, Anbang Liu, Qineng Wang +9

    cs.AIcs.CLcs.CVarXiv:2609.11900v12026
  42. ModaNet: A Large-Scale Street Fashion Dataset with Polygon Annotations

    Shuai Zheng, Fan Yang, M. Hadi Kiapour +1

    cs.CVarXiv:1807.01394v42018
  43. Towards the Generalization of Contrastive Self-Supervised Learning

    Weiran Huang, Mingyang Yi, Xuyang Zhao +1

    cs.LGcs.AIcs.CVarXiv:2111.00743v42021
  44. Video as Conditional Graph Hierarchy for Multi-Granular Question Answering

    Junbin Xiao, Angela Yao, Zhiyuan Liu +3

    cs.CVcs.AIcs.MMarXiv:2112.06197v22021
  45. GeoDA: a geometric framework for black-box adversarial attacks

    Ali Rahmati, Seyed-Mohsen Moosavi-Dezfooli, Pascal Frossard +1

    cs.CVcs.CRcs.LGarXiv:2003.06468v12020
  46. Context Decoupling Augmentation for Weakly Supervised Semantic Segmentation

    Yukun Su, Ruizhou Sun, Guosheng Lin +1

    cs.CVarXiv:2103.01795v22021
  47. TALL: Thumbnail Layout for Deepfake Video Detection

    Yuting Xu, Jian Liang, Gengyun Jia +3

    cs.CVarXiv:2307.07494v32023
  48. Connecting Gaze, Scene, and Attention: Generalized Attention Estimation via Joint Modeling of Gaze and Scene Saliency

    Eunji Chong, Nataniel Ruiz, Yongxin Wang +3

    cs.CVarXiv:1807.10437v12018
  49. Anti-UAV: A Large Multi-Modal Benchmark for UAV Tracking

    Nan Jiang, Kuiran Wang, Xiaoke Peng +7

    cs.CVarXiv:2101.08466v32021
  50. HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation

    Xuan Ju, Ailing Zeng, Chenchen Zhao +3

    cs.CVarXiv:2304.04269v12023
  51. Automated forest inventory: analysis of high-density airborne LiDAR point clouds with 3D deep learning

    Binbin Xiang, Maciej Wielgosz, Theodora Kontogianni +4

    cs.CVarXiv:2312.15084v22023
  52. Improving Multimodal Datasets with Image Captioning

    Thao Nguyen, Samir Yitzhak Gadre, Gabriel Ilharco +2

    cs.LGcs.CVarXiv:2307.10350v22023
  53. Hyperbolic Image Segmentation

    Mina GhadimiAtigh, Julian Schoep, Erman Acar +2

    cs.CVarXiv:2203.05898v12022
  54. Secure Face Matching Using Fully Homomorphic Encryption

    Vishnu Naresh Boddeti

    cs.CVarXiv:1805.00577v22018
  55. Prediction and Localization of Student Engagement in the Wild

    Amanjot Kaur, Aamir Mustafa, Love Mehta +1

    cs.CVarXiv:1804.00858v42018
  56. M3-Former: Multimodal Transformer with Mixture-of-Experts for Long-Term Vessel Trajectory Prediction

    Wenzhe Jin, Haina Tang

    cs.LGcs.CVarXiv:2609.10559v12026
  57. Diverse Trajectory Forecasting with Determinantal Point Processes

    Ye Yuan, Kris Kitani

    cs.CVcs.LGcs.ROarXiv:1907.04967v22019
  58. HOI Analysis: Integrating and Decomposing Human-Object Interaction

    Yong-Lu Li, Xinpeng Liu, Xiaoqian Wu +2

    cs.CVcs.LGeess.IVarXiv:2010.16219v22020
  59. AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

    Wenhao Chai, Enxin Song, Yilun Du +6

    cs.CVarXiv:2410.03051v42024
  60. Constrained R-CNN: A general image manipulation detection model

    Chao Yang, Huizhou Li, Fangting Lin +2

    cs.CVcs.MMarXiv:1911.08217v32019