Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
13,261 to 13,320 of 18,971
Multi-Garment Net: Learning to Dress 3D People from Images
Bharat Lal Bhatnagar, Garvita Tiwari, Christian Theobalt +1
cs.CVarXiv:1908.06903v22019TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering
Yushi Hu, Benlin Liu, Jungo Kasai +4
cs.CVarXiv:2303.11897v32023Skip-GANomaly: Skip Connected and Adversarially Trained Encoder-Decoder Anomaly Detection
Samet Akçay, Amir Atapour-Abarghouei, Toby P. Breckon
cs.CVcs.LGarXiv:1901.08954v12019VITON-HD: High-Resolution Virtual Try-On via Misalignment-Aware Normalization
Seunghwan Choi, Sunghyun Park, Minsoo Lee +1
cs.CVarXiv:2103.16874v22021A Study and Comparison of Human and Deep Learning Recognition Performance Under Visual Distortions
Samuel Dodge, Lina Karam
cs.CVarXiv:1705.02498v12017Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Chunting Zhou, Lili Yu, Arun Babu +7
cs.AIcs.CVarXiv:2408.11039v12024Swin-UMamba: Mamba-based UNet with ImageNet-based pretraining
Jiarun Liu, Hao Yang, Hong-Yu Zhou +8
eess.IVcs.CVcs.LGarXiv:2402.03302v22024Attention-Aware Compositional Network for Person Re-identification
Jing Xu, Rui Zhao, Feng Zhu +2
cs.CVarXiv:1805.03344v22018Assessment of algorithms for mitosis detection in breast cancer histopathology images
Mitko Veta, Paul J. van Diest, Stefan M. Willems +26
cs.CVarXiv:1411.5825v12014Multi-Agent Tensor Fusion for Contextual Trajectory Prediction
Tianyang Zhao, Yifei Xu, Mathew Monfort +5
cs.CVcs.LGarXiv:1904.04776v22019FlowFormer: A Transformer Architecture for Optical Flow
Zhaoyang Huang, Xiaoyu Shi, Chao Zhang +5
cs.CVarXiv:2203.16194v42022FaceShifter: Towards High Fidelity And Occlusion Aware Face Swapping
Lingzhi Li, Jianmin Bao, Hao Yang +2
cs.CVarXiv:1912.13457v32019Depth Pro: Sharp Monocular Metric Depth in Less Than a Second
Aleksei Bochkovskii, Amaël Delaunoy, Hugo Germain +4
cs.CVcs.LGarXiv:2410.02073v22024Summaries:한국어Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation
Hang Zhou, Yasheng Sun, Wayne Wu +3
cs.CVcs.LGcs.MMarXiv:2104.11116v12021Kornia: an Open Source Differentiable Computer Vision Library for PyTorch
Edgar Riba, Dmytro Mishkin, Daniel Ponsa +2
cs.CVarXiv:1910.02190v22019Domain Adaptive Neural Networks for Object Recognition
Muhammad Ghifary, W. Bastiaan Kleijn, Mengjie Zhang
cs.CVcs.AIcs.LGarXiv:1409.6041v12014Deep Metric Learning with Hierarchical Triplet Loss
Weifeng Ge, Weilin Huang, Dengke Dong +1
cs.CVarXiv:1810.06951v12018Stacked Generative Adversarial Networks
Xun Huang, Yixuan Li, Omid Poursaeed +2
cs.CVcs.LGcs.NEarXiv:1612.04357v42016Separable Self-attention for Mobile Vision Transformers
Sachin Mehta, Mohammad Rastegari
cs.CVcs.AIcs.LGarXiv:2206.02680v12022Data Augmentation by Pairing Samples for Images Classification
Hiroshi Inoue
cs.LGcs.CVstat.MLarXiv:1801.02929v22018Deformable Part Models are Convolutional Neural Networks
Ross Girshick, Forrest Iandola, Trevor Darrell +1
cs.CVarXiv:1409.5403v22014TransMed: Transformers Advance Multi-modal Medical Image Classification
Yin Dai, Yifan Gao
cs.CVarXiv:2103.05940v12021SuperYOLO: Super Resolution Assisted Object Detection in Multimodal Remote Sensing Imagery
Jiaqing Zhang, Jie Lei, Weiying Xie +3
cs.CVarXiv:2209.13351v22022MMRotate: A Rotated Object Detection Benchmark using PyTorch
Yue Zhou, Xue Yang, Gefan Zhang +9
cs.CVcs.AIarXiv:2204.13317v42022LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
Bin Zhu, Bin Lin, Munan Ning +11
cs.CVcs.AIarXiv:2310.01852v72023PatchmatchNet: Learned Multi-View Patchmatch Stereo
Fangjinhua Wang, Silvano Galliani, Christoph Vogel +2
cs.CVarXiv:2012.01411v12020Deep MANTA: A Coarse-to-fine Many-Task Network for joint 2D and 3D vehicle analysis from monocular image
Florian Chabot, Mohamed Chaouch, Jaonary Rabarisoa +2
cs.CVarXiv:1703.07570v12017Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model
Yu Du, Fangyun Wei, Zihe Zhang +3
cs.CVarXiv:2203.14940v12022CAT3D: Create Anything in 3D with Multi-View Diffusion Models
Ruiqi Gao, Aleksander Holynski, Philipp Henzler +5
cs.CVarXiv:2405.10314v12024Wayformer: Motion Forecasting via Simple & Efficient Attention Networks
Nigamaa Nayakanti, Rami Al-Rfou, Aurick Zhou +3
cs.CVarXiv:2207.05844v12022Active Object Localization with Deep Reinforcement Learning
Juan C. Caicedo, Svetlana Lazebnik
cs.CVarXiv:1511.06015v12015Neighbor2Neighbor: Self-Supervised Denoising from Single Noisy Images
Tao Huang, Songjiang Li, Xu Jia +2
eess.IVcs.CVarXiv:2101.02824v32021ScanQA: 3D Question Answering for Spatial Scene Understanding
Daichi Azuma, Taiki Miyanishi, Shuhei Kurita +1
cs.CVarXiv:2112.10482v32021A Gentle Introduction to Deep Learning in Medical Image Processing
Andreas Maier, Christopher Syben, Tobias Lasser +1
cs.CVarXiv:1810.05401v22018Capture, Learning, and Synthesis of 3D Speaking Styles
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw +2
cs.CVarXiv:1905.03079v12019Normalization Techniques in Training DNNs: Methodology, Analysis and Application
Lei Huang, Jie Qin, Yi Zhou +3
cs.LGcs.CVstat.MLarXiv:2009.12836v12020Visual Question Answering: A Survey of Methods and Datasets
Qi Wu, Damien Teney, Peng Wang +3
cs.CVarXiv:1607.05910v12016Modeling Spatial-Temporal Clues in a Hybrid Deep Learning Framework for Video Classification
Zuxuan Wu, Xi Wang, Yu-Gang Jiang +2
cs.CVcs.MMarXiv:1504.01561v12015Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
Zhang Li, Biao Yang, Qiang Liu +6
cs.CVcs.AIcs.CLarXiv:2311.06607v42023Deep Industrial Image Anomaly Detection: A Survey
Jiaqi Liu, Guoyang Xie, Jinbao Wang +4
cs.CVarXiv:2301.11514v52023Anchor-free Oriented Proposal Generator for Object Detection
Gong Cheng, Jiabao Wang, Ke Li +4
cs.CVarXiv:2110.01931v22021Large Pose 3D Face Reconstruction from a Single Image via Direct Volumetric CNN Regression
Aaron S. Jackson, Adrian Bulat, Vasileios Argyriou +1
cs.CVarXiv:1703.07834v22017ActionVLAD: Learning spatio-temporal aggregation for action classification
Rohit Girdhar, Deva Ramanan, Abhinav Gupta +2
cs.CVarXiv:1704.02895v12017Occ3D: A Large-Scale 3D Occupancy Prediction Benchmark for Autonomous Driving
Xiaoyu Tian, Tao Jiang, Longfei Yun +5
cs.CVarXiv:2304.14365v32023Analysis of Explainers of Black Box Deep Neural Networks for Computer Vision: A Survey
Vanessa Buhrmester, David Münch, Michael Arens
cs.AIcs.CVarXiv:1911.12116v12019Textual Explanations for Self-Driving Vehicles
Jinkyu Kim, Anna Rohrbach, Trevor Darrell +2
cs.CVarXiv:1807.11546v12018Knowledge Distillation Driven Semantic NOMA with GAN Refinement for 6G Robotic Vehicle Networks
Qifei Wang, Zhen Gao, Li Qiao +4
cs.ITcs.CVeess.IVarXiv:2608.27198v12026D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry
Nan Yang, Lukas von Stumberg, Rui Wang +1
cs.CVcs.AIarXiv:2003.01060v22020Heavy Rain Image Restoration: Integrating Physics Model and Conditional Adversarial Learning
Ruotent Li, Loong Fah Cheong, Robby T. Tan
cs.CVarXiv:1904.05050v12019University-1652: A Multi-view Multi-source Benchmark for Drone-based Geo-localization
Zhedong Zheng, Yunchao Wei, Yi Yang
cs.CVarXiv:2002.12186v22020ActBERT: Learning Global-Local Video-Text Representations
Linchao Zhu, Yi Yang
cs.CVarXiv:2011.07231v12020Noise or Signal: The Role of Image Backgrounds in Object Recognition
Kai Xiao, Logan Engstrom, Andrew Ilyas +1
cs.CVcs.LGarXiv:2006.09994v12020V2V-PoseNet: Voxel-to-Voxel Prediction Network for Accurate 3D Hand and Human Pose Estimation from a Single Depth Map
Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee
cs.CVarXiv:1711.07399v32017YOLOP: You Only Look Once for Panoptic Driving Perception
Dong Wu, Manwen Liao, Weitian Zhang +4
cs.CVarXiv:2108.11250v72021Robust Classification with Convolutional Prototype Learning
Hong-Ming Yang, Xu-Yao Zhang, Fei Yin +1
cs.CVarXiv:1805.03438v12018Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation
Bingxin Ke, Anton Obukhov, Shengyu Huang +3
cs.CVarXiv:2312.02145v22023Learning Pyramid-Context Encoder Network for High-Quality Image Inpainting
Yanhong Zeng, Jianlong Fu, Hongyang Chao +1
cs.CVarXiv:1904.07475v42019Robust Compressed Sensing MRI with Deep Generative Priors
Ajil Jalal, Marius Arvinte, Giannis Daras +3
cs.LGcs.CVcs.ITarXiv:2108.01368v22021An Empirical Study of Training End-to-End Vision-and-Language Transformers
Zi-Yi Dou, Yichong Xu, Zhe Gan +9
cs.CVcs.CLcs.LGarXiv:2111.02387v32021Proxy Anchor Loss for Deep Metric Learning
Sungyeon Kim, Dongwon Kim, Minsu Cho +1
cs.CVcs.LGarXiv:2003.13911v12020