Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
10,561 to 10,620 of 18,867
Detecting Curve Text in the Wild: New Dataset and New Solution
Liu Yuliang, Jin Lianwen, Zhang Shuaitao +1
cs.CVarXiv:1712.02170v12017Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation
Shizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi +2
cs.CVarXiv:2202.11742v12022Attention, please! A survey of Neural Attention Models in Deep Learning
Alana de Santana Correia, Esther Luna Colombini
cs.LGcs.AIcs.CVarXiv:2103.16775v12021ASFormer: Transformer for Action Segmentation
Fangqiu Yi, Hongyu Wen, Tingting Jiang
cs.CVarXiv:2110.08568v12021Multimodal Remote Sensing Benchmark Datasets for Land Cover Classification with A Shared and Specific Feature Learning Model
Danfeng Hong, Jingliang Hu, Jing Yao +2
cs.CVarXiv:2105.10196v12021Unified Contrastive Learning in Image-Text-Label Space
Jianwei Yang, Chunyuan Li, Pengchuan Zhang +4
cs.CVcs.AIcs.LGarXiv:2204.03610v12022On Face Segmentation, Face Swapping, and Face Perception
Yuval Nirkin, Iacopo Masi, Anh Tuan Tran +2
cs.CVarXiv:1704.06729v12017Multiview Transformers for Video Recognition
Shen Yan, Xuehan Xiong, Anurag Arnab +4
cs.CVcs.LGarXiv:2201.04288v42022Simple but Effective: CLIP Embeddings for Embodied AI
Apoorv Khandelwal, Luca Weihs, Roozbeh Mottaghi +1
cs.CVcs.LGarXiv:2111.09888v22021A Deep Learning based No-reference Quality Assessment Model for UGC Videos
Wei Sun, Xiongkuo Min, Wei Lu +1
cs.CVcs.MMeess.IVarXiv:2204.14047v22022TrajectoryNet: A Dynamic Optimal Transport Network for Modeling Cellular Dynamics
Alexander Tong, Jessie Huang, Guy Wolf +2
stat.MLcs.CVcs.LGarXiv:2002.04461v22020Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model
SII-GAIR, Sand. ai, : +43
cs.CVarXiv:2603.21986v12026Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Yanwei Li, Yuechen Zhang, Chengyao Wang +5
cs.CVcs.AIcs.CLarXiv:2403.18814v12024Segment Anything in Medical Images
Jun Ma, Yuting He, Feifei Li +3
eess.IVcs.CVarXiv:2304.12306v32023Deep Learning on Image Denoising: An overview
Chunwei Tian, Lunke Fei, Wenxian Zheng +3
eess.IVcs.CVarXiv:1912.13171v42019FastSurfer -- A fast and accurate deep learning based neuroimaging pipeline
Leonie Henschel, Sailesh Conjeti, Santiago Estrada +3
eess.IVcs.CVq-bio.NCarXiv:1910.03866v42019Recent Advances in Deep Learning for Object Detection
Xiongwei Wu, Doyen Sahoo, Steven C. H. Hoi
cs.CVcs.LGcs.MMarXiv:1908.03673v12019ResUNet-a: a deep learning framework for semantic segmentation of remotely sensed data
Foivos I. Diakogiannis, François Waldner, Peter Caccetta +1
cs.CVarXiv:1904.00592v32019Impact of Fully Connected Layers on Performance of Convolutional Neural Networks for Image Classification
S. H. Shabbeer Basha, Shiv Ram Dubey, Viswanath Pulabaigari +1
cs.CVcs.LGcs.NEarXiv:1902.02771v32019An overview of deep learning in medical imaging focusing on MRI
Alexander Selvikvåg Lundervold, Arvid Lundervold
cs.CVcs.LGstat.MLarXiv:1811.10052v22018A Deep Learning Framework for Unsupervised Affine and Deformable Image Registration
Bob D. de Vos, Floris F. Berendsen, Max A. Viergever +3
cs.CVarXiv:1809.06130v22018Confounding variables can degrade generalization performance of radiological deep learning models
John R. Zech, Marcus A. Badgeley, Manway Liu +3
cs.CVcs.LGstat.MLarXiv:1807.00431v22018OFF-ApexNet on Micro-expression Recognition System
Sze-Teng Liong, Y. S. Gan, Wei-Chuen Yau +2
cs.CVcs.LGarXiv:1805.08699v12018Beyond RGB: Very High Resolution Urban Remote Sensing With Multimodal Deep Networks
Nicolas Audebert, Bertrand Le Saux, Sébastien Lefèvre
cs.NEcs.CVarXiv:1711.08681v12017Remote Sensing Image Fusion Based on Two-stream Fusion Network
Xiangyu Liu, Qingjie Liu, Yunhong Wang
cs.CVarXiv:1711.02549v32017Deep Residual Bidir-LSTM for Human Activity Recognition Using Wearable Sensors
Yu Zhao, Rennong Yang, Guillaume Chevalier +1
cs.CVcs.LGarXiv:1708.08989v22017Learning Features for Offline Handwritten Signature Verification using Deep Convolutional Neural Networks
Luiz G. Hafemann, Robert Sabourin, Luiz S. Oliveira
cs.CVarXiv:1705.05787v12017Quicksilver: Fast Predictive Image Registration - a Deep Learning Approach
Xiao Yang, Roland Kwitt, Martin Styner +1
cs.CVarXiv:1703.10908v42017Algorithms for Semantic Segmentation of Multispectral Remote Sensing Imagery using Deep Learning
Ronald Kemker, Carl Salvaggio, Christopher Kanan
cs.CVcs.AIarXiv:1703.06452v32017Deep-Learning for Classification of Colorectal Polyps on Whole-Slide Images
Bruno Korbar, Andrea M. Olofson, Allen P. Miraflor +5
cs.CVarXiv:1703.01550v220173D fully convolutional networks for subcortical segmentation in MRI: A large-scale study
J. Dolz, C. Desrosiers, I. Ben Ayed
cs.CVarXiv:1612.03925v22016Superpixels: An Evaluation of the State-of-the-Art
David Stutz, Alexander Hermans, Bastian Leibe
cs.CVarXiv:1612.01601v32016Classification With an Edge: Improving Semantic Image Segmentation with Boundary Detection
Dimitrios Marmanis, Konrad Schindler, Jan Dirk Wegner +3
cs.CVarXiv:1612.01337v22016UniMiB SHAR: a new dataset for human activity recognition using acceleration data from smartphones
Daniela Micucci, Marco Mobilio, Paolo Napoletano
cs.CVarXiv:1611.07688v52016AutoInt: Automatic Integration for Fast Neural Volume Rendering
David B. Lindell, Julien N. P. Martel, Gordon Wetzstein
cs.CVcs.GRcs.LGarXiv:2012.01714v22020Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset
Jing Lin, Ailing Zeng, Shunlin Lu +4
cs.CVarXiv:2307.00818v22023Deep-Anomaly: Fully Convolutional Neural Network for Fast Anomaly Detection in Crowded Scenes
Mohammad Sabokrou, Mohsen Fayyaz, Mahmood Fathy +2
cs.CVarXiv:1609.00866v22016Less is More: Micro-expression Recognition from Video using Apex Frame
Sze-Teng Liong, John See, KokSheik Wong +1
cs.CVarXiv:1606.01721v32016A Combined Deep-Learning and Deformable-Model Approach to Fully Automatic Segmentation of the Left Ventricle in Cardiac MRI
M. R. Avendi, A. Kheradvar, H. Jafarkhani
cs.CVarXiv:1512.07951v12015Deep Feature Learning with Relative Distance Comparison for Person Re-identification
Shengyong Ding, Liang Lin, Guangrun Wang +1
cs.CVarXiv:1512.03622v12015Brain Tumor Segmentation with Deep Neural Networks
Mohammad Havaei, Axel Davy, David Warde-Farley +6
cs.CVcs.AIarXiv:1505.03540v32015Deep Learning for Classification and Severity Estimation of Coffee Leaf Biotic Stress
J. G. M. Esgario, R. A. Krohling, J. A. Ventura
cs.CVcs.LGarXiv:1907.11561v12019Multi-Scale Structure-Aware Network for Human Pose Estimation
Lipeng Ke, Ming-Ching Chang, Honggang Qi +1
cs.CVarXiv:1803.09894v32018Adversarial Manipulation of Deep Representations
Sara Sabour, Yanshuai Cao, Fartash Faghri +1
cs.CVcs.LGcs.NEarXiv:1511.05122v92015Learning From Noisy Labels By Regularized Estimation Of Annotator Confusion
Ryutaro Tanno, Ardavan Saeedi, Swami Sankaranarayanan +2
cs.LGcs.CVstat.MLarXiv:1902.03680v32019Rotation equivariant vector field networks
Diego Marcos, Michele Volpi, Nikos Komodakis +1
cs.CVarXiv:1612.09346v32016GIFT: A Real-time and Scalable 3D Shape Search Engine
Song Bai, Xiang Bai, Zhichao Zhou +2
cs.CVarXiv:1604.01879v22016PKU-MMD: A Large Scale Benchmark for Continuous Multi-Modal Human Action Understanding
Chunhui Liu, Yueyu Hu, Yanghao Li +2
cs.CVarXiv:1703.07475v22017Text2Shape: Generating Shapes from Natural Language by Learning Joint Embeddings
Kevin Chen, Christopher B. Choy, Manolis Savva +3
cs.CVcs.AIcs.GRarXiv:1803.08495v12018Visual Commonsense R-CNN
Tan Wang, Jianqiang Huang, Hanwang Zhang +1
cs.CVarXiv:2002.12204v32020PIRenderer: Controllable Portrait Image Generation via Semantic Neural Rendering
Yurui Ren, Ge Li, Yuanqi Chen +2
cs.CVcs.AIarXiv:2109.08379v12021MeshTalk: 3D Face Animation from Speech using Cross-Modality Disentanglement
Alexander Richard, Michael Zollhoefer, Yandong Wen +2
cs.CVarXiv:2104.08223v22021Tutel: Adaptive Mixture-of-Experts at Scale
Changho Hwang, Wei Cui, Yifan Xiong +12
cs.DCcs.CLcs.CVarXiv:2206.03382v22022Synthesizing Training Images for Boosting Human 3D Pose Estimation
Wenzheng Chen, Huan Wang, Yangyan Li +6
cs.CVarXiv:1604.02703v62016DoubleFusion: Real-time Capture of Human Performances with Inner Body Shapes from a Single Depth Sensor
Tao Yu, Zerong Zheng, Kaiwen Guo +5
cs.CVarXiv:1804.06023v12018Visual Interaction Networks
Nicholas Watters, Andrea Tacchetti, Theophane Weber +3
cs.CVarXiv:1706.01433v12017Reversible Architectures for Arbitrarily Deep Residual Neural Networks
Bo Chang, Lili Meng, Eldad Haber +3
cs.CVstat.MLarXiv:1709.03698v22017Scene Transformer: A unified architecture for predicting multiple agent trajectories
Jiquan Ngiam, Benjamin Caine, Vijay Vasudevan +11
cs.CVcs.LGcs.ROarXiv:2106.08417v32021Cyclical Stochastic Gradient MCMC for Bayesian Deep Learning
Ruqi Zhang, Chunyuan Li, Jianyi Zhang +2
cs.LGcs.AIcs.CVarXiv:1902.03932v22019EC-Net: an Edge-aware Point set Consolidation Network
Lequan Yu, Xianzhi Li, Chi-Wing Fu +2
cs.CVarXiv:1807.06010v12018