Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

11,821 to 11,880 of 18,815

  1. Learning Detailed Face Reconstruction from a Single Image

    Elad Richardson, Matan Sela, Roy Or-El +1

    cs.CVarXiv:1611.05053v22016
  2. Fast Image Scanning with Deep Max-Pooling Convolutional Neural Networks

    Alessandro Giusti, Dan C. Cireşan, Jonathan Masci +2

    cs.CVcs.AIarXiv:1302.1700v12013
  3. Dual-Level Collaborative Transformer for Image Captioning

    Yunpeng Luo, Jiayi Ji, Xiaoshuai Sun +5

    cs.CVarXiv:2101.06462v22021
  4. Complex-YOLO: Real-time 3D Object Detection on Point Clouds

    Martin Simon, Stefan Milz, Karl Amende +1

    cs.CVarXiv:1803.06199v22018
  5. HR-Depth: High Resolution Self-Supervised Monocular Depth Estimation

    Xiaoyang Lyu, Liang Liu, Mengmeng Wang +5

    cs.CVcs.AIarXiv:2012.07356v12020
  6. Bottom-Up Human Pose Estimation Via Disentangled Keypoint Regression

    Zigang Geng, Ke Sun, Bin Xiao +2

    cs.CVarXiv:2104.02300v12021
  7. Counterfactual Attention Learning for Fine-Grained Visual Categorization and Re-identification

    Yongming Rao, Guangyi Chen, Jiwen Lu +1

    cs.CVcs.AIcs.LGarXiv:2108.08728v22021
  8. Curriculum Domain Adaptation for Semantic Segmentation of Urban Scenes

    Yang Zhang, Philip David, Boqing Gong

    cs.CVcs.LGarXiv:1707.09465v52017
  9. Visual Saliency Detection Based on Multiscale Deep CNN Features

    Guanbin Li, Yizhou Yu

    cs.CVarXiv:1609.02077v12016
  10. EMOCA: Emotion Driven Monocular Face Capture and Animation

    Radek Danecek, Michael J. Black, Timo Bolkart

    cs.CVarXiv:2204.11312v12022
  11. DAMO-YOLO : A Report on Real-Time Object Detection Design

    Xianzhe Xu, Yiqi Jiang, Weihua Chen +3

    cs.CVarXiv:2211.15444v42022
  12. Artificial Intelligence for Digital and Computational Pathology

    Andrew H. Song, Guillaume Jaume, Drew F. K. Williamson +4

    eess.IVcs.AIcs.CVarXiv:2401.06148v12023
  13. Zero-Shot Learning via Joint Latent Similarity Embedding

    Ziming Zhang, Venkatesh Saligrama

    cs.CVarXiv:1511.04512v32015
  14. Efficient DETR: Improving End-to-End Object Detector with Dense Prior

    Zhuyu Yao, Jiangbo Ai, Boxun Li +1

    cs.CVarXiv:2104.01318v12021
  15. Self-Supervised Visual Planning with Temporal Skip Connections

    Frederik Ebert, Chelsea Finn, Alex X. Lee +1

    cs.ROcs.AIcs.CVarXiv:1710.05268v12017
  16. FD-GAN: Pose-guided Feature Distilling GAN for Robust Person Re-identification

    Yixiao Ge, Zhuowan Li, Haiyu Zhao +4

    cs.CVarXiv:1810.02936v22018
  17. Multi-Scale Positive Sample Refinement for Few-Shot Object Detection

    Jiaxi Wu, Songtao Liu, Di Huang +1

    cs.CVarXiv:2007.09384v12020
  18. WebFace260M: A Benchmark Unveiling the Power of Million-Scale Deep Face Recognition

    Zheng Zhu, Guan Huang, Jiankang Deng +8

    cs.CVarXiv:2103.04098v12021
  19. Unsupervised Deep Tracking

    Ning Wang, Yibing Song, Chao Ma +3

    cs.CVarXiv:1904.01828v12019
  20. CLIP-Mesh: Generating textured meshes from text using pretrained image-text models

    Nasir Mohammad Khalid, Tianhao Xie, Eugene Belilovsky +1

    cs.CVcs.GRcs.LGarXiv:2203.13333v22022
  21. Monocular Total Capture: Posing Face, Body, and Hands in the Wild

    Donglai Xiang, Hanbyul Joo, Yaser Sheikh

    cs.CVcs.GRarXiv:1812.01598v12018
  22. LayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation

    Guangcong Zheng, Xianpan Zhou, Xuewei Li +3

    cs.CVarXiv:2303.17189v22023
  23. SCRDet++: Detecting Small, Cluttered and Rotated Objects via Instance-Level Feature Denoising and Rotation Loss Smoothing

    Xue Yang, Junchi Yan, Wenlong Liao +3

    cs.CVcs.LGeess.IVarXiv:2004.13316v22020
  24. General Multi-label Image Classification with Transformers

    Jack Lanchantin, Tianlu Wang, Vicente Ordonez +1

    cs.CVcs.AIcs.LGarXiv:2011.14027v12020
  25. Lung and Colon Cancer Histopathological Image Dataset (LC25000)

    Andrew A. Borkowski, Marilyn M. Bui, L. Brannon Thomas +3

    eess.IVcs.CVq-bio.QMarXiv:1912.12142v12019
  26. Tube Convolutional Neural Network (T-CNN) for Action Detection in Videos

    Rui Hou, Chen Chen, Mubarak Shah

    cs.CVarXiv:1703.10664v32017
  27. Unsupervised Domain Adaptive Re-Identification: Theory and Practice

    Liangchen Song, Cheng Wang, Lefei Zhang +4

    cs.CVarXiv:1807.11334v12018
  28. Missing Data Reconstruction in Remote Sensing image with a Unified Spatial-Temporal-Spectral Deep Convolutional Neural Network

    Qiang Zhang, Qiangqiang Yuan, Chao Zeng +2

    cs.CVarXiv:1802.08369v12018
  29. Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos

    Yue Ma, Yingqing He, Xiaodong Cun +5

    cs.CVarXiv:2304.01186v22023
  30. Are Labels Required for Improving Adversarial Robustness?

    Jonathan Uesato, Jean-Baptiste Alayrac, Po-Sen Huang +3

    cs.LGcs.CVstat.MLarXiv:1905.13725v42019
  31. Understanding and Diagnosing Visual Tracking Systems

    Naiyan Wang, Jianping Shi, Dit-Yan Yeung +1

    cs.CVarXiv:1504.06055v12015
  32. Interpreting Super-Resolution Networks with Local Attribution Maps

    Jinjin Gu, Chao Dong

    cs.CVarXiv:2011.11036v22020
  33. Loss-Sensitive Generative Adversarial Networks on Lipschitz Densities

    Guo-Jun Qi

    cs.CVarXiv:1701.06264v62017
  34. Domain Generalization Using a Mixture of Multiple Latent Domains

    Toshihiko Matsuura, Tatsuya Harada

    cs.CVcs.LGarXiv:1911.07661v12019
  35. MS-Net: Multi-Site Network for Improving Prostate Segmentation with Heterogeneous MRI Data

    Quande Liu, Qi Dou, Lequan Yu +1

    eess.IVcs.CVarXiv:2002.03366v22020
  36. MPViT: Multi-Path Vision Transformer for Dense Prediction

    Youngwan Lee, Jonghee Kim, Jeff Willette +1

    cs.CVarXiv:2112.11010v22021
  37. MEMC-Net: Motion Estimation and Motion Compensation Driven Neural Network for Video Interpolation and Enhancement

    Wenbo Bao, Wei-Sheng Lai, Xiaoyun Zhang +2

    cs.CVarXiv:1810.08768v22018
  38. 3D Point Capsule Networks

    Yongheng Zhao, Tolga Birdal, Haowen Deng +1

    cs.CVcs.LGcs.NEarXiv:1812.10775v22018
  39. Fair DARTS: Eliminating Unfair Advantages in Differentiable Architecture Search

    Xiangxiang Chu, Tianbao Zhou, Bo Zhang +1

    cs.LGcs.AIcs.CVarXiv:1911.12126v42019
  40. Learning to Detect Objects with a 1 Megapixel Event Camera

    Etienne Perot, Pierre de Tournemire, Davide Nitti +2

    cs.CVcs.LGarXiv:2009.13436v22020
  41. Loam_livox: A fast, robust, high-precision LiDAR odometry and mapping package for LiDARs of small FoV

    Jiarong Lin, Fu Zhang

    cs.ROcs.CVeess.IVarXiv:1909.06700v12019
  42. Accurate Pulmonary Nodule Detection in Computed Tomography Images Using Deep Convolutional Neural Networks

    Jia Ding, Aoxue Li, Zhiqiang Hu +1

    cs.CVarXiv:1706.04303v32017
  43. On the Compactness, Efficiency, and Representation of 3D Convolutional Networks: Brain Parcellation as a Pretext Task

    Wenqi Li, Guotai Wang, Lucas Fidon +3

    cs.CVarXiv:1707.01992v12017
  44. Generalized optimal sub-pattern assignment metric

    Abu Sajana Rahmathullah, Ángel F. García-Fernández, Lennart Svensson

    eess.SYcs.CVarXiv:1601.05585v72016
  45. Enabling Deep Spiking Neural Networks with Hybrid Conversion and Spike Timing Dependent Backpropagation

    Nitin Rathi, Gopalakrishnan Srinivasan, Priyadarshini Panda +1

    cs.LGcs.CVstat.MLarXiv:2005.01807v12020
  46. A review of machine learning in processing remote sensing data for mineral exploration

    Hojat Shirmard, Ehsan Farahbakhsh, R. Dietmar Muller +1

    cs.LGcs.CVstat.AParXiv:2103.07678v22021
  47. Automatic Polyp Segmentation via Multi-scale Subtraction Network

    Xiaoqi Zhao, Lihe Zhang, Huchuan Lu

    cs.CVarXiv:2108.05082v12021
  48. CrossPoint: Self-Supervised Cross-Modal Contrastive Learning for 3D Point Cloud Understanding

    Mohamed Afham, Isuru Dissanayake, Dinithi Dissanayake +3

    cs.CVarXiv:2203.00680v32022
  49. ETH-XGaze: A Large Scale Dataset for Gaze Estimation under Extreme Head Pose and Gaze Variation

    Xucong Zhang, Seonwook Park, Thabo Beeler +3

    cs.CVarXiv:2007.15837v12020
  50. How2Sign: A Large-scale Multimodal Dataset for Continuous American Sign Language

    Amanda Duarte, Shruti Palaskar, Lucas Ventura +5

    cs.CVarXiv:2008.08143v22020
  51. Human Preference Score: Better Aligning Text-to-Image Models with Human Preference

    Xiaoshi Wu, Keqiang Sun, Feng Zhu +2

    cs.CVcs.AIarXiv:2303.14420v22023
  52. StructureFlow: Image Inpainting via Structure-aware Appearance Flow

    Yurui Ren, Xiaoming Yu, Ruonan Zhang +3

    cs.CVarXiv:1908.03852v12019
  53. Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability

    Shenyuan Gao, Jiazhi Yang, Li Chen +5

    cs.CVcs.AIarXiv:2405.17398v52024
  54. From Goals, Waypoints & Paths To Long Term Human Trajectory Forecasting

    Karttikeya Mangalam, Yang An, Harshayu Girase +1

    cs.CVcs.AIcs.ROarXiv:2012.01526v12020
  55. Automatic Liver and Tumor Segmentation of CT and MRI Volumes using Cascaded Fully Convolutional Neural Networks

    Patrick Ferdinand Christ, Florian Ettlinger, Felix Grün +17

    cs.CVcs.AIarXiv:1702.05970v22017
  56. CLIP-Forge: Towards Zero-Shot Text-to-Shape Generation

    Aditya Sanghi, Hang Chu, Joseph G. Lambourne +4

    cs.CVcs.AIarXiv:2110.02624v22021
  57. On the Use of Deep Learning for Blind Image Quality Assessment

    Simone Bianco, Luigi Celona, Paolo Napoletano +1

    cs.CVarXiv:1602.05531v52016
  58. Perception Test: A Diagnostic Benchmark for Multimodal Video Models

    Viorica Pătrăucean, Lucas Smaira, Ankush Gupta +21

    cs.CVcs.AIcs.LGarXiv:2305.13786v22023
  59. CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

    Han Fang, Pengfei Xiong, Luhui Xu +1

    cs.CVarXiv:2106.11097v12021
  60. CLIP-Driven Universal Model for Organ Segmentation and Tumor Detection

    Jie Liu, Yixiao Zhang, Jie-Neng Chen +7

    eess.IVcs.CVcs.LGarXiv:2301.00785v52023