Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

13,141 to 13,200 of 18,839

  1. Depth Pro: Sharp Monocular Metric Depth in Less Than a Second

    Aleksei Bochkovskii, Amaël Delaunoy, Hugo Germain +4

    cs.CVcs.LGarXiv:2410.02073v22024
    Summaries:한국어
  2. Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation

    Hang Zhou, Yasheng Sun, Wayne Wu +3

    cs.CVcs.LGcs.MMarXiv:2104.11116v12021
  3. Kornia: an Open Source Differentiable Computer Vision Library for PyTorch

    Edgar Riba, Dmytro Mishkin, Daniel Ponsa +2

    cs.CVarXiv:1910.02190v22019
  4. Domain Adaptive Neural Networks for Object Recognition

    Muhammad Ghifary, W. Bastiaan Kleijn, Mengjie Zhang

    cs.CVcs.AIcs.LGarXiv:1409.6041v12014
  5. Deep Metric Learning with Hierarchical Triplet Loss

    Weifeng Ge, Weilin Huang, Dengke Dong +1

    cs.CVarXiv:1810.06951v12018
  6. Stacked Generative Adversarial Networks

    Xun Huang, Yixuan Li, Omid Poursaeed +2

    cs.CVcs.LGcs.NEarXiv:1612.04357v42016
  7. Separable Self-attention for Mobile Vision Transformers

    Sachin Mehta, Mohammad Rastegari

    cs.CVcs.AIcs.LGarXiv:2206.02680v12022
  8. Data Augmentation by Pairing Samples for Images Classification

    Hiroshi Inoue

    cs.LGcs.CVstat.MLarXiv:1801.02929v22018
  9. Deformable Part Models are Convolutional Neural Networks

    Ross Girshick, Forrest Iandola, Trevor Darrell +1

    cs.CVarXiv:1409.5403v22014
  10. TransMed: Transformers Advance Multi-modal Medical Image Classification

    Yin Dai, Yifan Gao

    cs.CVarXiv:2103.05940v12021
  11. SuperYOLO: Super Resolution Assisted Object Detection in Multimodal Remote Sensing Imagery

    Jiaqing Zhang, Jie Lei, Weiying Xie +3

    cs.CVarXiv:2209.13351v22022
  12. MMRotate: A Rotated Object Detection Benchmark using PyTorch

    Yue Zhou, Xue Yang, Gefan Zhang +9

    cs.CVcs.AIarXiv:2204.13317v42022
  13. LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

    Bin Zhu, Bin Lin, Munan Ning +11

    cs.CVcs.AIarXiv:2310.01852v72023
  14. PatchmatchNet: Learned Multi-View Patchmatch Stereo

    Fangjinhua Wang, Silvano Galliani, Christoph Vogel +2

    cs.CVarXiv:2012.01411v12020
  15. Deep MANTA: A Coarse-to-fine Many-Task Network for joint 2D and 3D vehicle analysis from monocular image

    Florian Chabot, Mohamed Chaouch, Jaonary Rabarisoa +2

    cs.CVarXiv:1703.07570v12017
  16. Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model

    Yu Du, Fangyun Wei, Zihe Zhang +3

    cs.CVarXiv:2203.14940v12022
  17. CAT3D: Create Anything in 3D with Multi-View Diffusion Models

    Ruiqi Gao, Aleksander Holynski, Philipp Henzler +5

    cs.CVarXiv:2405.10314v12024
  18. Wayformer: Motion Forecasting via Simple & Efficient Attention Networks

    Nigamaa Nayakanti, Rami Al-Rfou, Aurick Zhou +3

    cs.CVarXiv:2207.05844v12022
  19. Active Object Localization with Deep Reinforcement Learning

    Juan C. Caicedo, Svetlana Lazebnik

    cs.CVarXiv:1511.06015v12015
  20. Neighbor2Neighbor: Self-Supervised Denoising from Single Noisy Images

    Tao Huang, Songjiang Li, Xu Jia +2

    eess.IVcs.CVarXiv:2101.02824v32021
  21. ScanQA: 3D Question Answering for Spatial Scene Understanding

    Daichi Azuma, Taiki Miyanishi, Shuhei Kurita +1

    cs.CVarXiv:2112.10482v32021
  22. A Gentle Introduction to Deep Learning in Medical Image Processing

    Andreas Maier, Christopher Syben, Tobias Lasser +1

    cs.CVarXiv:1810.05401v22018
  23. Capture, Learning, and Synthesis of 3D Speaking Styles

    Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw +2

    cs.CVarXiv:1905.03079v12019
  24. Normalization Techniques in Training DNNs: Methodology, Analysis and Application

    Lei Huang, Jie Qin, Yi Zhou +3

    cs.LGcs.CVstat.MLarXiv:2009.12836v12020
  25. Visual Question Answering: A Survey of Methods and Datasets

    Qi Wu, Damien Teney, Peng Wang +3

    cs.CVarXiv:1607.05910v12016
  26. Modeling Spatial-Temporal Clues in a Hybrid Deep Learning Framework for Video Classification

    Zuxuan Wu, Xi Wang, Yu-Gang Jiang +2

    cs.CVcs.MMarXiv:1504.01561v12015
  27. Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

    Zhang Li, Biao Yang, Qiang Liu +6

    cs.CVcs.AIcs.CLarXiv:2311.06607v42023
  28. Deep Industrial Image Anomaly Detection: A Survey

    Jiaqi Liu, Guoyang Xie, Jinbao Wang +4

    cs.CVarXiv:2301.11514v52023
  29. Anchor-free Oriented Proposal Generator for Object Detection

    Gong Cheng, Jiabao Wang, Ke Li +4

    cs.CVarXiv:2110.01931v22021
  30. Large Pose 3D Face Reconstruction from a Single Image via Direct Volumetric CNN Regression

    Aaron S. Jackson, Adrian Bulat, Vasileios Argyriou +1

    cs.CVarXiv:1703.07834v22017
  31. ActionVLAD: Learning spatio-temporal aggregation for action classification

    Rohit Girdhar, Deva Ramanan, Abhinav Gupta +2

    cs.CVarXiv:1704.02895v12017
  32. Occ3D: A Large-Scale 3D Occupancy Prediction Benchmark for Autonomous Driving

    Xiaoyu Tian, Tao Jiang, Longfei Yun +5

    cs.CVarXiv:2304.14365v32023
  33. Analysis of Explainers of Black Box Deep Neural Networks for Computer Vision: A Survey

    Vanessa Buhrmester, David Münch, Michael Arens

    cs.AIcs.CVarXiv:1911.12116v12019
  34. Textual Explanations for Self-Driving Vehicles

    Jinkyu Kim, Anna Rohrbach, Trevor Darrell +2

    cs.CVarXiv:1807.11546v12018
  35. Knowledge Distillation Driven Semantic NOMA with GAN Refinement for 6G Robotic Vehicle Networks

    Qifei Wang, Zhen Gao, Li Qiao +4

    cs.ITcs.CVeess.IVarXiv:2608.27198v12026
  36. D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry

    Nan Yang, Lukas von Stumberg, Rui Wang +1

    cs.CVcs.AIarXiv:2003.01060v22020
  37. Heavy Rain Image Restoration: Integrating Physics Model and Conditional Adversarial Learning

    Ruotent Li, Loong Fah Cheong, Robby T. Tan

    cs.CVarXiv:1904.05050v12019
  38. University-1652: A Multi-view Multi-source Benchmark for Drone-based Geo-localization

    Zhedong Zheng, Yunchao Wei, Yi Yang

    cs.CVarXiv:2002.12186v22020
  39. ActBERT: Learning Global-Local Video-Text Representations

    Linchao Zhu, Yi Yang

    cs.CVarXiv:2011.07231v12020
  40. Noise or Signal: The Role of Image Backgrounds in Object Recognition

    Kai Xiao, Logan Engstrom, Andrew Ilyas +1

    cs.CVcs.LGarXiv:2006.09994v12020
  41. V2V-PoseNet: Voxel-to-Voxel Prediction Network for Accurate 3D Hand and Human Pose Estimation from a Single Depth Map

    Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee

    cs.CVarXiv:1711.07399v32017
  42. YOLOP: You Only Look Once for Panoptic Driving Perception

    Dong Wu, Manwen Liao, Weitian Zhang +4

    cs.CVarXiv:2108.11250v72021
  43. Robust Classification with Convolutional Prototype Learning

    Hong-Ming Yang, Xu-Yao Zhang, Fei Yin +1

    cs.CVarXiv:1805.03438v12018
  44. Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation

    Bingxin Ke, Anton Obukhov, Shengyu Huang +3

    cs.CVarXiv:2312.02145v22023
  45. Learning Pyramid-Context Encoder Network for High-Quality Image Inpainting

    Yanhong Zeng, Jianlong Fu, Hongyang Chao +1

    cs.CVarXiv:1904.07475v42019
  46. Robust Compressed Sensing MRI with Deep Generative Priors

    Ajil Jalal, Marius Arvinte, Giannis Daras +3

    cs.LGcs.CVcs.ITarXiv:2108.01368v22021
  47. An Empirical Study of Training End-to-End Vision-and-Language Transformers

    Zi-Yi Dou, Yichong Xu, Zhe Gan +9

    cs.CVcs.CLcs.LGarXiv:2111.02387v32021
  48. Proxy Anchor Loss for Deep Metric Learning

    Sungyeon Kim, Dongwon Kim, Minsu Cho +1

    cs.CVcs.LGarXiv:2003.13911v12020
  49. Fast Fusion of Multi-Band Images Based on Solving a Sylvester Equation

    Qi Wei, Nicolas Dobigeon, Jean-Yves Tourneret

    cs.CVarXiv:1502.03121v12015
  50. BEVFormer v2: Adapting Modern Image Backbones to Bird's-Eye-View Recognition via Perspective Supervision

    Chenyu Yang, Yuntao Chen, Hao Tian +9

    cs.CVarXiv:2211.10439v12022
  51. Semi-Supervised Semantic Segmentation with High- and Low-level Consistency

    Sudhanshu Mittal, Maxim Tatarchenko, Thomas Brox

    cs.CVarXiv:1908.05724v12019
  52. QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection

    Chenhongyi Yang, Zehao Huang, Naiyan Wang

    cs.CVarXiv:2103.09136v22021
  53. Denoising Diffusion Models for Plug-and-Play Image Restoration

    Yuanzhi Zhu, Kai Zhang, Jingyun Liang +4

    cs.CVeess.IVarXiv:2305.08995v12023
  54. Robotic Grasp Detection using Deep Convolutional Neural Networks

    Sulabh Kumra, Christopher Kanan

    cs.ROcs.CVarXiv:1611.08036v42016
  55. A Neural Approach to Blind Motion Deblurring

    Ayan Chakrabarti

    cs.CVarXiv:1603.04771v22016
  56. Text Line Segmentation of Historical Documents: a Survey

    Laurence Likforman-Sulem, Abderrazak Zahour, Bruno Taconet

    cs.CVarXiv:0704.1267v12007
  57. Low-bit Quantization of Neural Networks for Efficient Inference

    Yoni Choukroun, Eli Kravchik, Fan Yang +1

    cs.LGcs.CVstat.MLarXiv:1902.06822v22019
  58. Multispectral Deep Neural Networks for Pedestrian Detection

    Jingjing Liu, Shaoting Zhang, Shu Wang +1

    cs.CVarXiv:1611.02644v12016
  59. MoMask: Generative Masked Modeling of 3D Human Motions

    Chuan Guo, Yuxuan Mu, Muhammad Gohar Javed +2

    cs.CVarXiv:2312.00063v12023
  60. Salient Object Detection via Integrity Learning

    Mingchen Zhuge, Deng-Ping Fan, Nian Liu +3

    cs.CVarXiv:2101.07663v72021