Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
13,501 to 13,560 of 18,849
Spatial-Temporal Recurrent Neural Network for Emotion Recognition
Tong Zhang, Wenming Zheng, Zhen Cui +2
cs.CVarXiv:1705.04515v12017Scene Text Detection and Recognition: The Deep Learning Era
Shangbang Long, Xin He, Cong Yao
cs.CVarXiv:1811.04256v52018TransVG: End-to-End Visual Grounding with Transformers
Jiajun Deng, Zhengyuan Yang, Tianlang Chen +2
cs.CVarXiv:2104.08541v42021Illumination-aware Faster R-CNN for Robust Multispectral Pedestrian Detection
Chengyang Li, Dan Song, Ruofeng Tong +1
cs.CVarXiv:1803.05347v22018UniPC: A Unified Predictor-Corrector Framework for Fast Sampling of Diffusion Models
Wenliang Zhao, Lujia Bai, Yongming Rao +2
cs.LGcs.CVarXiv:2302.04867v42023Regressing Robust and Discriminative 3D Morphable Models with a very Deep Neural Network
Anh Tuan Tran, Tal Hassner, Iacopo Masi +1
cs.CVarXiv:1612.04904v12016VL-Adapter: Parameter-Efficient Transfer Learning for Vision-and-Language Tasks
Yi-Lin Sung, Jaemin Cho, Mohit Bansal
cs.CVcs.AIcs.CLarXiv:2112.06825v22021Pluralistic Image Completion
Chuanxia Zheng, Tat-Jen Cham, Jianfei Cai
cs.CVarXiv:1903.04227v22019Neural Architecture Search without Training
Joseph Mellor, Jack Turner, Amos Storkey +1
cs.LGcs.CVstat.MLarXiv:2006.04647v32020DeepWeeds: A Multiclass Weed Species Image Dataset for Deep Learning
Alex Olsen, Dmitry A. Konovalov, Bronson Philippa +10
cs.CVcs.LGstat.MLarXiv:1810.05726v320183D Hand Shape and Pose Estimation from a Single RGB Image
Liuhao Ge, Zhou Ren, Yuncheng Li +4
cs.CVarXiv:1903.00812v22019Depth Prediction Without the Sensors: Leveraging Structure for Unsupervised Learning from Monocular Videos
Vincent Casser, Soeren Pirk, Reza Mahjourian +1
cs.CVarXiv:1811.06152v12018OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
Yuliang Liu, Zhang Li, Mingxin Huang +7
cs.CVcs.CLarXiv:2305.07895v720233D Menagerie: Modeling the 3D shape and pose of animals
Silvia Zuffi, Angjoo Kanazawa, David Jacobs +1
cs.CVarXiv:1611.07700v22016An inertial forward-backward algorithm for monotone inclusions
Dirk A. Lorenz, Thomas Pock
cs.CVmath.NAmath.OCarXiv:1403.3522v22014Unsupervised Person Re-identification: Clustering and Fine-tuning
Hehe Fan, Liang Zheng, Yi Yang
cs.CVarXiv:1705.10444v22017Probabilistic Anchor Assignment with IoU Prediction for Object Detection
Kang Kim, Hee Seok Lee
cs.CVarXiv:2007.08103v22020Understanding and Creating Art with AI: Review and Outlook
Eva Cetinic, James She
cs.CVcs.AIcs.MMarXiv:2102.09109v12021YOLO-FaceV2: A Scale and Occlusion Aware Face Detector
Ziping Yu, Hongbo Huang, Weijun Chen +3
cs.CVarXiv:2208.02019v22022Dense semantic labeling of sub-decimeter resolution images with convolutional neural networks
Michele Volpi, Devis Tuia
cs.CVarXiv:1608.00775v22016SynthStrip: Skull-Stripping for Any Brain Image
Andrew Hoopes, Jocelyn S. Mora, Adrian V. Dalca +2
eess.IVcs.CVphysics.med-pharXiv:2203.09974v22022Deep Fruit Detection in Orchards
Suchet Bargoti, James Underwood
cs.ROcs.AIcs.CVarXiv:1610.03677v22016Universal Guidance for Diffusion Models
Arpit Bansal, Hong-Min Chu, Avi Schwarzschild +4
cs.CVcs.LGarXiv:2302.07121v12023Temporal Generative Adversarial Nets with Singular Value Clipping
Masaki Saito, Eiichi Matsumoto, Shunta Saito
cs.LGcs.CVarXiv:1611.06624v32016MagicVideo: Efficient Video Generation With Latent Diffusion Models
Daquan Zhou, Weimin Wang, Hanshu Yan +3
cs.CVarXiv:2211.11018v22022What Do Single-view 3D Reconstruction Networks Learn?
Maxim Tatarchenko, Stephan R. Richter, René Ranftl +3
cs.CVarXiv:1905.03678v1201912-in-1: Multi-Task Vision and Language Representation Learning
Jiasen Lu, Vedanuj Goswami, Marcus Rohrbach +2
cs.CVcs.CLcs.LGarXiv:1912.02315v22019DeepSaliency: Multi-Task Deep Neural Network Model for Salient Object Detection
Xi Li, Liming Zhao, Lina Wei +5
cs.CVarXiv:1510.05484v22015Horizontal Pyramid Matching for Person Re-identification
Yang Fu, Yunchao Wei, Yuqian Zhou +5
cs.CVarXiv:1804.05275v42018Need for Speed: A Benchmark for Higher Frame Rate Object Tracking
Hamed Kiani Galoogahi, Ashton Fagg, Chen Huang +2
cs.CVarXiv:1703.05884v22017PDEBENCH: An Extensive Benchmark for Scientific Machine Learning
Makoto Takamoto, Timothy Praditia, Raphael Leiteritz +4
cs.LGcs.CVphysics.flu-dynarXiv:2210.07182v72022Deep Image Matting
Ning Xu, Brian Price, Scott Cohen +1
cs.CVarXiv:1703.03872v32017Survey on RGB, 3D, Thermal, and Multimodal Approaches for Facial Expression Recognition: History, Trends, and Affect-related Applications
Ciprian Corneanu, Marc Oliu, Jeffrey F. Cohn +1
cs.CVarXiv:1606.03237v12016Exploring Uncertainty Measures in Deep Networks for Multiple Sclerosis Lesion Detection and Segmentation
Tanya Nair, Doina Precup, Douglas L. Arnold +1
cs.CVarXiv:1808.01200v22018Shape-aware Semi-supervised 3D Semantic Segmentation for Medical Images
Shuailin Li, Chuyu Zhang, Xuming He
cs.CVarXiv:2007.10732v12020The KIT Motion-Language Dataset
Matthias Plappert, Christian Mandery, Tamim Asfour
cs.ROcs.CLcs.CVarXiv:1607.03827v22016DeepFix: A Fully Convolutional Neural Network for predicting Human Eye Fixations
Srinivas S. S. Kruthiventi, Kumar Ayush, R. Venkatesh Babu
cs.CVarXiv:1510.02927v12015YOLACT++: Better Real-time Instance Segmentation
Daniel Bolya, Chong Zhou, Fanyi Xiao +1
cs.CVcs.LGeess.IVarXiv:1912.06218v22019I2L-MeshNet: Image-to-Lixel Prediction Network for Accurate 3D Human Pose and Mesh Estimation from a Single RGB Image
Gyeongsik Moon, Kyoung Mu Lee
cs.CVarXiv:2008.03713v22020Deep Reinforcement Learning for Unsupervised Video Summarization with Diversity-Representativeness Reward
Kaiyang Zhou, Yu Qiao, Tao Xiang
cs.CVarXiv:1801.00054v32017Oriented RepPoints for Aerial Object Detection
Wentong Li, Yijie Chen, Kaixuan Hu +1
cs.CVarXiv:2105.11111v42021Sim-to-Real via Sim-to-Sim: Data-efficient Robotic Grasping via Randomized-to-Canonical Adaptation Networks
Stephen James, Paul Wohlhart, Mrinal Kalakrishnan +6
cs.ROcs.CVcs.LGarXiv:1812.07252v32018Grader variability and the importance of reference standards for evaluating machine learning models for diabetic retinopathy
Jonathan Krause, Varun Gulshan, Ehsan Rahimy +5
cs.CVarXiv:1710.01711v32017Unsupervised Image Super-Resolution using Cycle-in-Cycle Generative Adversarial Networks
Yuan Yuan, Siyuan Liu, Jiawei Zhang +3
cs.CVarXiv:1809.00437v12018G-TAD: Sub-Graph Localization for Temporal Action Detection
Mengmeng Xu, Chen Zhao, David S. Rojas +2
cs.CVarXiv:1911.11462v22019Context Autoencoder for Self-Supervised Representation Learning
Xiaokang Chen, Mingyu Ding, Xiaodi Wang +7
cs.CVarXiv:2202.03026v32022Long-tailed Recognition by Routing Diverse Distribution-Aware Experts
Xudong Wang, Long Lian, Zhongqi Miao +2
cs.CVarXiv:2010.01809v42020Look at Boundary: A Boundary-Aware Face Alignment Algorithm
Wayne Wu, Chen Qian, Shuo Yang +3
cs.CVarXiv:1805.10483v12018Look into Person: Self-supervised Structure-sensitive Learning and A New Benchmark for Human Parsing
Ke Gong, Xiaodan Liang, Dongyu Zhang +2
cs.CVcs.AIcs.LGarXiv:1703.05446v22017Towards Better Analysis of Deep Convolutional Neural Networks
Mengchen Liu, Jiaxin Shi, Zhen Li +3
cs.CVarXiv:1604.07043v32016Learning From Noisy Large-Scale Datasets With Minimal Supervision
Andreas Veit, Neil Alldrin, Gal Chechik +3
cs.CVarXiv:1701.01619v22017Tightly Coupled 3D Lidar Inertial Odometry and Mapping
Haoyang Ye, Yuying Chen, Ming Liu
cs.ROcs.CVarXiv:1904.06993v12019Post-Training Quantization for Vision Transformer
Zhenhua Liu, Yunhe Wang, Kai Han +2
cs.CVarXiv:2106.14156v12021Disentangling Monocular 3D Object Detection
Andrea Simonelli, Samuel Rota Rota Bulò, Lorenzo Porzi +2
cs.CVarXiv:1905.12365v12019Multi-scale self-guided attention for medical image segmentation
Ashish Sinha, Jose Dolz
cs.CVarXiv:1906.02849v32019VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training
Yecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman +3
cs.ROcs.AIcs.CVarXiv:2210.00030v22022Scene Text Visual Question Answering
Ali Furkan Biten, Ruben Tito, Andres Mafla +5
cs.CVarXiv:1905.13648v22019SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation
Wenxuan Zhang, Xiaodong Cun, Xuan Wang +5
cs.CVarXiv:2211.12194v22022MonoScene: Monocular 3D Semantic Scene Completion
Anh-Quan Cao, Raoul de Charette
cs.CVcs.AIcs.ROarXiv:2112.00726v22021Learning Activation Functions to Improve Deep Neural Networks
Forest Agostinelli, Matthew Hoffman, Peter Sadowski +1
cs.NEcs.CVcs.LGarXiv:1412.6830v32014