Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

13,261 to 13,320 of 18,971

  1. Multi-Garment Net: Learning to Dress 3D People from Images

    Bharat Lal Bhatnagar, Garvita Tiwari, Christian Theobalt +1

    cs.CVarXiv:1908.06903v22019
  2. TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

    Yushi Hu, Benlin Liu, Jungo Kasai +4

    cs.CVarXiv:2303.11897v32023
  3. Skip-GANomaly: Skip Connected and Adversarially Trained Encoder-Decoder Anomaly Detection

    Samet Akçay, Amir Atapour-Abarghouei, Toby P. Breckon

    cs.CVcs.LGarXiv:1901.08954v12019
  4. VITON-HD: High-Resolution Virtual Try-On via Misalignment-Aware Normalization

    Seunghwan Choi, Sunghyun Park, Minsoo Lee +1

    cs.CVarXiv:2103.16874v22021
  5. A Study and Comparison of Human and Deep Learning Recognition Performance Under Visual Distortions

    Samuel Dodge, Lina Karam

    cs.CVarXiv:1705.02498v12017
  6. Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

    Chunting Zhou, Lili Yu, Arun Babu +7

    cs.AIcs.CVarXiv:2408.11039v12024
  7. Swin-UMamba: Mamba-based UNet with ImageNet-based pretraining

    Jiarun Liu, Hao Yang, Hong-Yu Zhou +8

    eess.IVcs.CVcs.LGarXiv:2402.03302v22024
  8. Attention-Aware Compositional Network for Person Re-identification

    Jing Xu, Rui Zhao, Feng Zhu +2

    cs.CVarXiv:1805.03344v22018
  9. Assessment of algorithms for mitosis detection in breast cancer histopathology images

    Mitko Veta, Paul J. van Diest, Stefan M. Willems +26

    cs.CVarXiv:1411.5825v12014
  10. Multi-Agent Tensor Fusion for Contextual Trajectory Prediction

    Tianyang Zhao, Yifei Xu, Mathew Monfort +5

    cs.CVcs.LGarXiv:1904.04776v22019
  11. FlowFormer: A Transformer Architecture for Optical Flow

    Zhaoyang Huang, Xiaoyu Shi, Chao Zhang +5

    cs.CVarXiv:2203.16194v42022
  12. FaceShifter: Towards High Fidelity And Occlusion Aware Face Swapping

    Lingzhi Li, Jianmin Bao, Hao Yang +2

    cs.CVarXiv:1912.13457v32019
  13. Depth Pro: Sharp Monocular Metric Depth in Less Than a Second

    Aleksei Bochkovskii, Amaël Delaunoy, Hugo Germain +4

    cs.CVcs.LGarXiv:2410.02073v22024
    Summaries:한국어
  14. Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation

    Hang Zhou, Yasheng Sun, Wayne Wu +3

    cs.CVcs.LGcs.MMarXiv:2104.11116v12021
  15. Kornia: an Open Source Differentiable Computer Vision Library for PyTorch

    Edgar Riba, Dmytro Mishkin, Daniel Ponsa +2

    cs.CVarXiv:1910.02190v22019
  16. Domain Adaptive Neural Networks for Object Recognition

    Muhammad Ghifary, W. Bastiaan Kleijn, Mengjie Zhang

    cs.CVcs.AIcs.LGarXiv:1409.6041v12014
  17. Deep Metric Learning with Hierarchical Triplet Loss

    Weifeng Ge, Weilin Huang, Dengke Dong +1

    cs.CVarXiv:1810.06951v12018
  18. Stacked Generative Adversarial Networks

    Xun Huang, Yixuan Li, Omid Poursaeed +2

    cs.CVcs.LGcs.NEarXiv:1612.04357v42016
  19. Separable Self-attention for Mobile Vision Transformers

    Sachin Mehta, Mohammad Rastegari

    cs.CVcs.AIcs.LGarXiv:2206.02680v12022
  20. Data Augmentation by Pairing Samples for Images Classification

    Hiroshi Inoue

    cs.LGcs.CVstat.MLarXiv:1801.02929v22018
  21. Deformable Part Models are Convolutional Neural Networks

    Ross Girshick, Forrest Iandola, Trevor Darrell +1

    cs.CVarXiv:1409.5403v22014
  22. TransMed: Transformers Advance Multi-modal Medical Image Classification

    Yin Dai, Yifan Gao

    cs.CVarXiv:2103.05940v12021
  23. SuperYOLO: Super Resolution Assisted Object Detection in Multimodal Remote Sensing Imagery

    Jiaqing Zhang, Jie Lei, Weiying Xie +3

    cs.CVarXiv:2209.13351v22022
  24. MMRotate: A Rotated Object Detection Benchmark using PyTorch

    Yue Zhou, Xue Yang, Gefan Zhang +9

    cs.CVcs.AIarXiv:2204.13317v42022
  25. LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

    Bin Zhu, Bin Lin, Munan Ning +11

    cs.CVcs.AIarXiv:2310.01852v72023
  26. PatchmatchNet: Learned Multi-View Patchmatch Stereo

    Fangjinhua Wang, Silvano Galliani, Christoph Vogel +2

    cs.CVarXiv:2012.01411v12020
  27. Deep MANTA: A Coarse-to-fine Many-Task Network for joint 2D and 3D vehicle analysis from monocular image

    Florian Chabot, Mohamed Chaouch, Jaonary Rabarisoa +2

    cs.CVarXiv:1703.07570v12017
  28. Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model

    Yu Du, Fangyun Wei, Zihe Zhang +3

    cs.CVarXiv:2203.14940v12022
  29. CAT3D: Create Anything in 3D with Multi-View Diffusion Models

    Ruiqi Gao, Aleksander Holynski, Philipp Henzler +5

    cs.CVarXiv:2405.10314v12024
  30. Wayformer: Motion Forecasting via Simple & Efficient Attention Networks

    Nigamaa Nayakanti, Rami Al-Rfou, Aurick Zhou +3

    cs.CVarXiv:2207.05844v12022
  31. Active Object Localization with Deep Reinforcement Learning

    Juan C. Caicedo, Svetlana Lazebnik

    cs.CVarXiv:1511.06015v12015
  32. Neighbor2Neighbor: Self-Supervised Denoising from Single Noisy Images

    Tao Huang, Songjiang Li, Xu Jia +2

    eess.IVcs.CVarXiv:2101.02824v32021
  33. ScanQA: 3D Question Answering for Spatial Scene Understanding

    Daichi Azuma, Taiki Miyanishi, Shuhei Kurita +1

    cs.CVarXiv:2112.10482v32021
  34. A Gentle Introduction to Deep Learning in Medical Image Processing

    Andreas Maier, Christopher Syben, Tobias Lasser +1

    cs.CVarXiv:1810.05401v22018
  35. Capture, Learning, and Synthesis of 3D Speaking Styles

    Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw +2

    cs.CVarXiv:1905.03079v12019
  36. Normalization Techniques in Training DNNs: Methodology, Analysis and Application

    Lei Huang, Jie Qin, Yi Zhou +3

    cs.LGcs.CVstat.MLarXiv:2009.12836v12020
  37. Visual Question Answering: A Survey of Methods and Datasets

    Qi Wu, Damien Teney, Peng Wang +3

    cs.CVarXiv:1607.05910v12016
  38. Modeling Spatial-Temporal Clues in a Hybrid Deep Learning Framework for Video Classification

    Zuxuan Wu, Xi Wang, Yu-Gang Jiang +2

    cs.CVcs.MMarXiv:1504.01561v12015
  39. Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

    Zhang Li, Biao Yang, Qiang Liu +6

    cs.CVcs.AIcs.CLarXiv:2311.06607v42023
  40. Deep Industrial Image Anomaly Detection: A Survey

    Jiaqi Liu, Guoyang Xie, Jinbao Wang +4

    cs.CVarXiv:2301.11514v52023
  41. Anchor-free Oriented Proposal Generator for Object Detection

    Gong Cheng, Jiabao Wang, Ke Li +4

    cs.CVarXiv:2110.01931v22021
  42. Large Pose 3D Face Reconstruction from a Single Image via Direct Volumetric CNN Regression

    Aaron S. Jackson, Adrian Bulat, Vasileios Argyriou +1

    cs.CVarXiv:1703.07834v22017
  43. ActionVLAD: Learning spatio-temporal aggregation for action classification

    Rohit Girdhar, Deva Ramanan, Abhinav Gupta +2

    cs.CVarXiv:1704.02895v12017
  44. Occ3D: A Large-Scale 3D Occupancy Prediction Benchmark for Autonomous Driving

    Xiaoyu Tian, Tao Jiang, Longfei Yun +5

    cs.CVarXiv:2304.14365v32023
  45. Analysis of Explainers of Black Box Deep Neural Networks for Computer Vision: A Survey

    Vanessa Buhrmester, David Münch, Michael Arens

    cs.AIcs.CVarXiv:1911.12116v12019
  46. Textual Explanations for Self-Driving Vehicles

    Jinkyu Kim, Anna Rohrbach, Trevor Darrell +2

    cs.CVarXiv:1807.11546v12018
  47. Knowledge Distillation Driven Semantic NOMA with GAN Refinement for 6G Robotic Vehicle Networks

    Qifei Wang, Zhen Gao, Li Qiao +4

    cs.ITcs.CVeess.IVarXiv:2608.27198v12026
  48. D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry

    Nan Yang, Lukas von Stumberg, Rui Wang +1

    cs.CVcs.AIarXiv:2003.01060v22020
  49. Heavy Rain Image Restoration: Integrating Physics Model and Conditional Adversarial Learning

    Ruotent Li, Loong Fah Cheong, Robby T. Tan

    cs.CVarXiv:1904.05050v12019
  50. University-1652: A Multi-view Multi-source Benchmark for Drone-based Geo-localization

    Zhedong Zheng, Yunchao Wei, Yi Yang

    cs.CVarXiv:2002.12186v22020
  51. ActBERT: Learning Global-Local Video-Text Representations

    Linchao Zhu, Yi Yang

    cs.CVarXiv:2011.07231v12020
  52. Noise or Signal: The Role of Image Backgrounds in Object Recognition

    Kai Xiao, Logan Engstrom, Andrew Ilyas +1

    cs.CVcs.LGarXiv:2006.09994v12020
  53. V2V-PoseNet: Voxel-to-Voxel Prediction Network for Accurate 3D Hand and Human Pose Estimation from a Single Depth Map

    Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee

    cs.CVarXiv:1711.07399v32017
  54. YOLOP: You Only Look Once for Panoptic Driving Perception

    Dong Wu, Manwen Liao, Weitian Zhang +4

    cs.CVarXiv:2108.11250v72021
  55. Robust Classification with Convolutional Prototype Learning

    Hong-Ming Yang, Xu-Yao Zhang, Fei Yin +1

    cs.CVarXiv:1805.03438v12018
  56. Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation

    Bingxin Ke, Anton Obukhov, Shengyu Huang +3

    cs.CVarXiv:2312.02145v22023
  57. Learning Pyramid-Context Encoder Network for High-Quality Image Inpainting

    Yanhong Zeng, Jianlong Fu, Hongyang Chao +1

    cs.CVarXiv:1904.07475v42019
  58. Robust Compressed Sensing MRI with Deep Generative Priors

    Ajil Jalal, Marius Arvinte, Giannis Daras +3

    cs.LGcs.CVcs.ITarXiv:2108.01368v22021
  59. An Empirical Study of Training End-to-End Vision-and-Language Transformers

    Zi-Yi Dou, Yichong Xu, Zhe Gan +9

    cs.CVcs.CLcs.LGarXiv:2111.02387v32021
  60. Proxy Anchor Loss for Deep Metric Learning

    Sungyeon Kim, Dongwon Kim, Minsu Cho +1

    cs.CVcs.LGarXiv:2003.13911v12020