Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

13,501 to 13,560 of 18,849

  1. Spatial-Temporal Recurrent Neural Network for Emotion Recognition

    Tong Zhang, Wenming Zheng, Zhen Cui +2

    cs.CVarXiv:1705.04515v12017
  2. Scene Text Detection and Recognition: The Deep Learning Era

    Shangbang Long, Xin He, Cong Yao

    cs.CVarXiv:1811.04256v52018
  3. TransVG: End-to-End Visual Grounding with Transformers

    Jiajun Deng, Zhengyuan Yang, Tianlang Chen +2

    cs.CVarXiv:2104.08541v42021
  4. Illumination-aware Faster R-CNN for Robust Multispectral Pedestrian Detection

    Chengyang Li, Dan Song, Ruofeng Tong +1

    cs.CVarXiv:1803.05347v22018
  5. UniPC: A Unified Predictor-Corrector Framework for Fast Sampling of Diffusion Models

    Wenliang Zhao, Lujia Bai, Yongming Rao +2

    cs.LGcs.CVarXiv:2302.04867v42023
  6. Regressing Robust and Discriminative 3D Morphable Models with a very Deep Neural Network

    Anh Tuan Tran, Tal Hassner, Iacopo Masi +1

    cs.CVarXiv:1612.04904v12016
  7. VL-Adapter: Parameter-Efficient Transfer Learning for Vision-and-Language Tasks

    Yi-Lin Sung, Jaemin Cho, Mohit Bansal

    cs.CVcs.AIcs.CLarXiv:2112.06825v22021
  8. Pluralistic Image Completion

    Chuanxia Zheng, Tat-Jen Cham, Jianfei Cai

    cs.CVarXiv:1903.04227v22019
  9. Neural Architecture Search without Training

    Joseph Mellor, Jack Turner, Amos Storkey +1

    cs.LGcs.CVstat.MLarXiv:2006.04647v32020
  10. DeepWeeds: A Multiclass Weed Species Image Dataset for Deep Learning

    Alex Olsen, Dmitry A. Konovalov, Bronson Philippa +10

    cs.CVcs.LGstat.MLarXiv:1810.05726v32018
  11. 3D Hand Shape and Pose Estimation from a Single RGB Image

    Liuhao Ge, Zhou Ren, Yuncheng Li +4

    cs.CVarXiv:1903.00812v22019
  12. Depth Prediction Without the Sensors: Leveraging Structure for Unsupervised Learning from Monocular Videos

    Vincent Casser, Soeren Pirk, Reza Mahjourian +1

    cs.CVarXiv:1811.06152v12018
  13. OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

    Yuliang Liu, Zhang Li, Mingxin Huang +7

    cs.CVcs.CLarXiv:2305.07895v72023
  14. 3D Menagerie: Modeling the 3D shape and pose of animals

    Silvia Zuffi, Angjoo Kanazawa, David Jacobs +1

    cs.CVarXiv:1611.07700v22016
  15. An inertial forward-backward algorithm for monotone inclusions

    Dirk A. Lorenz, Thomas Pock

    cs.CVmath.NAmath.OCarXiv:1403.3522v22014
  16. Unsupervised Person Re-identification: Clustering and Fine-tuning

    Hehe Fan, Liang Zheng, Yi Yang

    cs.CVarXiv:1705.10444v22017
  17. Probabilistic Anchor Assignment with IoU Prediction for Object Detection

    Kang Kim, Hee Seok Lee

    cs.CVarXiv:2007.08103v22020
  18. Understanding and Creating Art with AI: Review and Outlook

    Eva Cetinic, James She

    cs.CVcs.AIcs.MMarXiv:2102.09109v12021
  19. YOLO-FaceV2: A Scale and Occlusion Aware Face Detector

    Ziping Yu, Hongbo Huang, Weijun Chen +3

    cs.CVarXiv:2208.02019v22022
  20. Dense semantic labeling of sub-decimeter resolution images with convolutional neural networks

    Michele Volpi, Devis Tuia

    cs.CVarXiv:1608.00775v22016
  21. SynthStrip: Skull-Stripping for Any Brain Image

    Andrew Hoopes, Jocelyn S. Mora, Adrian V. Dalca +2

    eess.IVcs.CVphysics.med-pharXiv:2203.09974v22022
  22. Deep Fruit Detection in Orchards

    Suchet Bargoti, James Underwood

    cs.ROcs.AIcs.CVarXiv:1610.03677v22016
  23. Universal Guidance for Diffusion Models

    Arpit Bansal, Hong-Min Chu, Avi Schwarzschild +4

    cs.CVcs.LGarXiv:2302.07121v12023
  24. Temporal Generative Adversarial Nets with Singular Value Clipping

    Masaki Saito, Eiichi Matsumoto, Shunta Saito

    cs.LGcs.CVarXiv:1611.06624v32016
  25. MagicVideo: Efficient Video Generation With Latent Diffusion Models

    Daquan Zhou, Weimin Wang, Hanshu Yan +3

    cs.CVarXiv:2211.11018v22022
  26. What Do Single-view 3D Reconstruction Networks Learn?

    Maxim Tatarchenko, Stephan R. Richter, René Ranftl +3

    cs.CVarXiv:1905.03678v12019
  27. 12-in-1: Multi-Task Vision and Language Representation Learning

    Jiasen Lu, Vedanuj Goswami, Marcus Rohrbach +2

    cs.CVcs.CLcs.LGarXiv:1912.02315v22019
  28. DeepSaliency: Multi-Task Deep Neural Network Model for Salient Object Detection

    Xi Li, Liming Zhao, Lina Wei +5

    cs.CVarXiv:1510.05484v22015
  29. Horizontal Pyramid Matching for Person Re-identification

    Yang Fu, Yunchao Wei, Yuqian Zhou +5

    cs.CVarXiv:1804.05275v42018
  30. Need for Speed: A Benchmark for Higher Frame Rate Object Tracking

    Hamed Kiani Galoogahi, Ashton Fagg, Chen Huang +2

    cs.CVarXiv:1703.05884v22017
  31. PDEBENCH: An Extensive Benchmark for Scientific Machine Learning

    Makoto Takamoto, Timothy Praditia, Raphael Leiteritz +4

    cs.LGcs.CVphysics.flu-dynarXiv:2210.07182v72022
  32. Deep Image Matting

    Ning Xu, Brian Price, Scott Cohen +1

    cs.CVarXiv:1703.03872v32017
  33. Survey on RGB, 3D, Thermal, and Multimodal Approaches for Facial Expression Recognition: History, Trends, and Affect-related Applications

    Ciprian Corneanu, Marc Oliu, Jeffrey F. Cohn +1

    cs.CVarXiv:1606.03237v12016
  34. Exploring Uncertainty Measures in Deep Networks for Multiple Sclerosis Lesion Detection and Segmentation

    Tanya Nair, Doina Precup, Douglas L. Arnold +1

    cs.CVarXiv:1808.01200v22018
  35. Shape-aware Semi-supervised 3D Semantic Segmentation for Medical Images

    Shuailin Li, Chuyu Zhang, Xuming He

    cs.CVarXiv:2007.10732v12020
  36. The KIT Motion-Language Dataset

    Matthias Plappert, Christian Mandery, Tamim Asfour

    cs.ROcs.CLcs.CVarXiv:1607.03827v22016
  37. DeepFix: A Fully Convolutional Neural Network for predicting Human Eye Fixations

    Srinivas S. S. Kruthiventi, Kumar Ayush, R. Venkatesh Babu

    cs.CVarXiv:1510.02927v12015
  38. YOLACT++: Better Real-time Instance Segmentation

    Daniel Bolya, Chong Zhou, Fanyi Xiao +1

    cs.CVcs.LGeess.IVarXiv:1912.06218v22019
  39. I2L-MeshNet: Image-to-Lixel Prediction Network for Accurate 3D Human Pose and Mesh Estimation from a Single RGB Image

    Gyeongsik Moon, Kyoung Mu Lee

    cs.CVarXiv:2008.03713v22020
  40. Deep Reinforcement Learning for Unsupervised Video Summarization with Diversity-Representativeness Reward

    Kaiyang Zhou, Yu Qiao, Tao Xiang

    cs.CVarXiv:1801.00054v32017
  41. Oriented RepPoints for Aerial Object Detection

    Wentong Li, Yijie Chen, Kaixuan Hu +1

    cs.CVarXiv:2105.11111v42021
  42. Sim-to-Real via Sim-to-Sim: Data-efficient Robotic Grasping via Randomized-to-Canonical Adaptation Networks

    Stephen James, Paul Wohlhart, Mrinal Kalakrishnan +6

    cs.ROcs.CVcs.LGarXiv:1812.07252v32018
  43. Grader variability and the importance of reference standards for evaluating machine learning models for diabetic retinopathy

    Jonathan Krause, Varun Gulshan, Ehsan Rahimy +5

    cs.CVarXiv:1710.01711v32017
  44. Unsupervised Image Super-Resolution using Cycle-in-Cycle Generative Adversarial Networks

    Yuan Yuan, Siyuan Liu, Jiawei Zhang +3

    cs.CVarXiv:1809.00437v12018
  45. G-TAD: Sub-Graph Localization for Temporal Action Detection

    Mengmeng Xu, Chen Zhao, David S. Rojas +2

    cs.CVarXiv:1911.11462v22019
  46. Context Autoencoder for Self-Supervised Representation Learning

    Xiaokang Chen, Mingyu Ding, Xiaodi Wang +7

    cs.CVarXiv:2202.03026v32022
  47. Long-tailed Recognition by Routing Diverse Distribution-Aware Experts

    Xudong Wang, Long Lian, Zhongqi Miao +2

    cs.CVarXiv:2010.01809v42020
  48. Look at Boundary: A Boundary-Aware Face Alignment Algorithm

    Wayne Wu, Chen Qian, Shuo Yang +3

    cs.CVarXiv:1805.10483v12018
  49. Look into Person: Self-supervised Structure-sensitive Learning and A New Benchmark for Human Parsing

    Ke Gong, Xiaodan Liang, Dongyu Zhang +2

    cs.CVcs.AIcs.LGarXiv:1703.05446v22017
  50. Towards Better Analysis of Deep Convolutional Neural Networks

    Mengchen Liu, Jiaxin Shi, Zhen Li +3

    cs.CVarXiv:1604.07043v32016
  51. Learning From Noisy Large-Scale Datasets With Minimal Supervision

    Andreas Veit, Neil Alldrin, Gal Chechik +3

    cs.CVarXiv:1701.01619v22017
  52. Tightly Coupled 3D Lidar Inertial Odometry and Mapping

    Haoyang Ye, Yuying Chen, Ming Liu

    cs.ROcs.CVarXiv:1904.06993v12019
  53. Post-Training Quantization for Vision Transformer

    Zhenhua Liu, Yunhe Wang, Kai Han +2

    cs.CVarXiv:2106.14156v12021
  54. Disentangling Monocular 3D Object Detection

    Andrea Simonelli, Samuel Rota Rota Bulò, Lorenzo Porzi +2

    cs.CVarXiv:1905.12365v12019
  55. Multi-scale self-guided attention for medical image segmentation

    Ashish Sinha, Jose Dolz

    cs.CVarXiv:1906.02849v32019
  56. VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training

    Yecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman +3

    cs.ROcs.AIcs.CVarXiv:2210.00030v22022
  57. Scene Text Visual Question Answering

    Ali Furkan Biten, Ruben Tito, Andres Mafla +5

    cs.CVarXiv:1905.13648v22019
  58. SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation

    Wenxuan Zhang, Xiaodong Cun, Xuan Wang +5

    cs.CVarXiv:2211.12194v22022
  59. MonoScene: Monocular 3D Semantic Scene Completion

    Anh-Quan Cao, Raoul de Charette

    cs.CVcs.AIcs.ROarXiv:2112.00726v22021
  60. Learning Activation Functions to Improve Deep Neural Networks

    Forest Agostinelli, Matthew Hoffman, Peter Sadowski +1

    cs.NEcs.CVcs.LGarXiv:1412.6830v32014