Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
13,141 to 13,200 of 18,839
Depth Pro: Sharp Monocular Metric Depth in Less Than a Second
Aleksei Bochkovskii, Amaël Delaunoy, Hugo Germain +4
cs.CVcs.LGarXiv:2410.02073v22024Summaries:한국어Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation
Hang Zhou, Yasheng Sun, Wayne Wu +3
cs.CVcs.LGcs.MMarXiv:2104.11116v12021Kornia: an Open Source Differentiable Computer Vision Library for PyTorch
Edgar Riba, Dmytro Mishkin, Daniel Ponsa +2
cs.CVarXiv:1910.02190v22019Domain Adaptive Neural Networks for Object Recognition
Muhammad Ghifary, W. Bastiaan Kleijn, Mengjie Zhang
cs.CVcs.AIcs.LGarXiv:1409.6041v12014Deep Metric Learning with Hierarchical Triplet Loss
Weifeng Ge, Weilin Huang, Dengke Dong +1
cs.CVarXiv:1810.06951v12018Stacked Generative Adversarial Networks
Xun Huang, Yixuan Li, Omid Poursaeed +2
cs.CVcs.LGcs.NEarXiv:1612.04357v42016Separable Self-attention for Mobile Vision Transformers
Sachin Mehta, Mohammad Rastegari
cs.CVcs.AIcs.LGarXiv:2206.02680v12022Data Augmentation by Pairing Samples for Images Classification
Hiroshi Inoue
cs.LGcs.CVstat.MLarXiv:1801.02929v22018Deformable Part Models are Convolutional Neural Networks
Ross Girshick, Forrest Iandola, Trevor Darrell +1
cs.CVarXiv:1409.5403v22014TransMed: Transformers Advance Multi-modal Medical Image Classification
Yin Dai, Yifan Gao
cs.CVarXiv:2103.05940v12021SuperYOLO: Super Resolution Assisted Object Detection in Multimodal Remote Sensing Imagery
Jiaqing Zhang, Jie Lei, Weiying Xie +3
cs.CVarXiv:2209.13351v22022MMRotate: A Rotated Object Detection Benchmark using PyTorch
Yue Zhou, Xue Yang, Gefan Zhang +9
cs.CVcs.AIarXiv:2204.13317v42022LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
Bin Zhu, Bin Lin, Munan Ning +11
cs.CVcs.AIarXiv:2310.01852v72023PatchmatchNet: Learned Multi-View Patchmatch Stereo
Fangjinhua Wang, Silvano Galliani, Christoph Vogel +2
cs.CVarXiv:2012.01411v12020Deep MANTA: A Coarse-to-fine Many-Task Network for joint 2D and 3D vehicle analysis from monocular image
Florian Chabot, Mohamed Chaouch, Jaonary Rabarisoa +2
cs.CVarXiv:1703.07570v12017Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model
Yu Du, Fangyun Wei, Zihe Zhang +3
cs.CVarXiv:2203.14940v12022CAT3D: Create Anything in 3D with Multi-View Diffusion Models
Ruiqi Gao, Aleksander Holynski, Philipp Henzler +5
cs.CVarXiv:2405.10314v12024Wayformer: Motion Forecasting via Simple & Efficient Attention Networks
Nigamaa Nayakanti, Rami Al-Rfou, Aurick Zhou +3
cs.CVarXiv:2207.05844v12022Active Object Localization with Deep Reinforcement Learning
Juan C. Caicedo, Svetlana Lazebnik
cs.CVarXiv:1511.06015v12015Neighbor2Neighbor: Self-Supervised Denoising from Single Noisy Images
Tao Huang, Songjiang Li, Xu Jia +2
eess.IVcs.CVarXiv:2101.02824v32021ScanQA: 3D Question Answering for Spatial Scene Understanding
Daichi Azuma, Taiki Miyanishi, Shuhei Kurita +1
cs.CVarXiv:2112.10482v32021A Gentle Introduction to Deep Learning in Medical Image Processing
Andreas Maier, Christopher Syben, Tobias Lasser +1
cs.CVarXiv:1810.05401v22018Capture, Learning, and Synthesis of 3D Speaking Styles
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw +2
cs.CVarXiv:1905.03079v12019Normalization Techniques in Training DNNs: Methodology, Analysis and Application
Lei Huang, Jie Qin, Yi Zhou +3
cs.LGcs.CVstat.MLarXiv:2009.12836v12020Visual Question Answering: A Survey of Methods and Datasets
Qi Wu, Damien Teney, Peng Wang +3
cs.CVarXiv:1607.05910v12016Modeling Spatial-Temporal Clues in a Hybrid Deep Learning Framework for Video Classification
Zuxuan Wu, Xi Wang, Yu-Gang Jiang +2
cs.CVcs.MMarXiv:1504.01561v12015Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
Zhang Li, Biao Yang, Qiang Liu +6
cs.CVcs.AIcs.CLarXiv:2311.06607v42023Deep Industrial Image Anomaly Detection: A Survey
Jiaqi Liu, Guoyang Xie, Jinbao Wang +4
cs.CVarXiv:2301.11514v52023Anchor-free Oriented Proposal Generator for Object Detection
Gong Cheng, Jiabao Wang, Ke Li +4
cs.CVarXiv:2110.01931v22021Large Pose 3D Face Reconstruction from a Single Image via Direct Volumetric CNN Regression
Aaron S. Jackson, Adrian Bulat, Vasileios Argyriou +1
cs.CVarXiv:1703.07834v22017ActionVLAD: Learning spatio-temporal aggregation for action classification
Rohit Girdhar, Deva Ramanan, Abhinav Gupta +2
cs.CVarXiv:1704.02895v12017Occ3D: A Large-Scale 3D Occupancy Prediction Benchmark for Autonomous Driving
Xiaoyu Tian, Tao Jiang, Longfei Yun +5
cs.CVarXiv:2304.14365v32023Analysis of Explainers of Black Box Deep Neural Networks for Computer Vision: A Survey
Vanessa Buhrmester, David Münch, Michael Arens
cs.AIcs.CVarXiv:1911.12116v12019Textual Explanations for Self-Driving Vehicles
Jinkyu Kim, Anna Rohrbach, Trevor Darrell +2
cs.CVarXiv:1807.11546v12018Knowledge Distillation Driven Semantic NOMA with GAN Refinement for 6G Robotic Vehicle Networks
Qifei Wang, Zhen Gao, Li Qiao +4
cs.ITcs.CVeess.IVarXiv:2608.27198v12026D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry
Nan Yang, Lukas von Stumberg, Rui Wang +1
cs.CVcs.AIarXiv:2003.01060v22020Heavy Rain Image Restoration: Integrating Physics Model and Conditional Adversarial Learning
Ruotent Li, Loong Fah Cheong, Robby T. Tan
cs.CVarXiv:1904.05050v12019University-1652: A Multi-view Multi-source Benchmark for Drone-based Geo-localization
Zhedong Zheng, Yunchao Wei, Yi Yang
cs.CVarXiv:2002.12186v22020ActBERT: Learning Global-Local Video-Text Representations
Linchao Zhu, Yi Yang
cs.CVarXiv:2011.07231v12020Noise or Signal: The Role of Image Backgrounds in Object Recognition
Kai Xiao, Logan Engstrom, Andrew Ilyas +1
cs.CVcs.LGarXiv:2006.09994v12020V2V-PoseNet: Voxel-to-Voxel Prediction Network for Accurate 3D Hand and Human Pose Estimation from a Single Depth Map
Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee
cs.CVarXiv:1711.07399v32017YOLOP: You Only Look Once for Panoptic Driving Perception
Dong Wu, Manwen Liao, Weitian Zhang +4
cs.CVarXiv:2108.11250v72021Robust Classification with Convolutional Prototype Learning
Hong-Ming Yang, Xu-Yao Zhang, Fei Yin +1
cs.CVarXiv:1805.03438v12018Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation
Bingxin Ke, Anton Obukhov, Shengyu Huang +3
cs.CVarXiv:2312.02145v22023Learning Pyramid-Context Encoder Network for High-Quality Image Inpainting
Yanhong Zeng, Jianlong Fu, Hongyang Chao +1
cs.CVarXiv:1904.07475v42019Robust Compressed Sensing MRI with Deep Generative Priors
Ajil Jalal, Marius Arvinte, Giannis Daras +3
cs.LGcs.CVcs.ITarXiv:2108.01368v22021An Empirical Study of Training End-to-End Vision-and-Language Transformers
Zi-Yi Dou, Yichong Xu, Zhe Gan +9
cs.CVcs.CLcs.LGarXiv:2111.02387v32021Proxy Anchor Loss for Deep Metric Learning
Sungyeon Kim, Dongwon Kim, Minsu Cho +1
cs.CVcs.LGarXiv:2003.13911v12020Fast Fusion of Multi-Band Images Based on Solving a Sylvester Equation
Qi Wei, Nicolas Dobigeon, Jean-Yves Tourneret
cs.CVarXiv:1502.03121v12015BEVFormer v2: Adapting Modern Image Backbones to Bird's-Eye-View Recognition via Perspective Supervision
Chenyu Yang, Yuntao Chen, Hao Tian +9
cs.CVarXiv:2211.10439v12022Semi-Supervised Semantic Segmentation with High- and Low-level Consistency
Sudhanshu Mittal, Maxim Tatarchenko, Thomas Brox
cs.CVarXiv:1908.05724v12019QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection
Chenhongyi Yang, Zehao Huang, Naiyan Wang
cs.CVarXiv:2103.09136v22021Denoising Diffusion Models for Plug-and-Play Image Restoration
Yuanzhi Zhu, Kai Zhang, Jingyun Liang +4
cs.CVeess.IVarXiv:2305.08995v12023Robotic Grasp Detection using Deep Convolutional Neural Networks
Sulabh Kumra, Christopher Kanan
cs.ROcs.CVarXiv:1611.08036v42016A Neural Approach to Blind Motion Deblurring
Ayan Chakrabarti
cs.CVarXiv:1603.04771v22016Text Line Segmentation of Historical Documents: a Survey
Laurence Likforman-Sulem, Abderrazak Zahour, Bruno Taconet
cs.CVarXiv:0704.1267v12007Low-bit Quantization of Neural Networks for Efficient Inference
Yoni Choukroun, Eli Kravchik, Fan Yang +1
cs.LGcs.CVstat.MLarXiv:1902.06822v22019Multispectral Deep Neural Networks for Pedestrian Detection
Jingjing Liu, Shaoting Zhang, Shu Wang +1
cs.CVarXiv:1611.02644v12016MoMask: Generative Masked Modeling of 3D Human Motions
Chuan Guo, Yuxuan Mu, Muhammad Gohar Javed +2
cs.CVarXiv:2312.00063v12023Salient Object Detection via Integrity Learning
Mingchen Zhuge, Deng-Ping Fan, Nian Liu +3
cs.CVarXiv:2101.07663v72021