Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

10,561 to 10,620 of 18,867

  1. Detecting Curve Text in the Wild: New Dataset and New Solution

    Liu Yuliang, Jin Lianwen, Zhang Shuaitao +1

    cs.CVarXiv:1712.02170v12017
  2. Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation

    Shizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi +2

    cs.CVarXiv:2202.11742v12022
  3. Attention, please! A survey of Neural Attention Models in Deep Learning

    Alana de Santana Correia, Esther Luna Colombini

    cs.LGcs.AIcs.CVarXiv:2103.16775v12021
  4. ASFormer: Transformer for Action Segmentation

    Fangqiu Yi, Hongyu Wen, Tingting Jiang

    cs.CVarXiv:2110.08568v12021
  5. Multimodal Remote Sensing Benchmark Datasets for Land Cover Classification with A Shared and Specific Feature Learning Model

    Danfeng Hong, Jingliang Hu, Jing Yao +2

    cs.CVarXiv:2105.10196v12021
  6. Unified Contrastive Learning in Image-Text-Label Space

    Jianwei Yang, Chunyuan Li, Pengchuan Zhang +4

    cs.CVcs.AIcs.LGarXiv:2204.03610v12022
  7. On Face Segmentation, Face Swapping, and Face Perception

    Yuval Nirkin, Iacopo Masi, Anh Tuan Tran +2

    cs.CVarXiv:1704.06729v12017
  8. Multiview Transformers for Video Recognition

    Shen Yan, Xuehan Xiong, Anurag Arnab +4

    cs.CVcs.LGarXiv:2201.04288v42022
  9. Simple but Effective: CLIP Embeddings for Embodied AI

    Apoorv Khandelwal, Luca Weihs, Roozbeh Mottaghi +1

    cs.CVcs.LGarXiv:2111.09888v22021
  10. A Deep Learning based No-reference Quality Assessment Model for UGC Videos

    Wei Sun, Xiongkuo Min, Wei Lu +1

    cs.CVcs.MMeess.IVarXiv:2204.14047v22022
  11. TrajectoryNet: A Dynamic Optimal Transport Network for Modeling Cellular Dynamics

    Alexander Tong, Jessie Huang, Guy Wolf +2

    stat.MLcs.CVcs.LGarXiv:2002.04461v22020
  12. Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model

    SII-GAIR, Sand. ai, : +43

    cs.CVarXiv:2603.21986v12026
  13. Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

    Yanwei Li, Yuechen Zhang, Chengyao Wang +5

    cs.CVcs.AIcs.CLarXiv:2403.18814v12024
  14. Segment Anything in Medical Images

    Jun Ma, Yuting He, Feifei Li +3

    eess.IVcs.CVarXiv:2304.12306v32023
  15. Deep Learning on Image Denoising: An overview

    Chunwei Tian, Lunke Fei, Wenxian Zheng +3

    eess.IVcs.CVarXiv:1912.13171v42019
  16. FastSurfer -- A fast and accurate deep learning based neuroimaging pipeline

    Leonie Henschel, Sailesh Conjeti, Santiago Estrada +3

    eess.IVcs.CVq-bio.NCarXiv:1910.03866v42019
  17. Recent Advances in Deep Learning for Object Detection

    Xiongwei Wu, Doyen Sahoo, Steven C. H. Hoi

    cs.CVcs.LGcs.MMarXiv:1908.03673v12019
  18. ResUNet-a: a deep learning framework for semantic segmentation of remotely sensed data

    Foivos I. Diakogiannis, François Waldner, Peter Caccetta +1

    cs.CVarXiv:1904.00592v32019
  19. Impact of Fully Connected Layers on Performance of Convolutional Neural Networks for Image Classification

    S. H. Shabbeer Basha, Shiv Ram Dubey, Viswanath Pulabaigari +1

    cs.CVcs.LGcs.NEarXiv:1902.02771v32019
  20. An overview of deep learning in medical imaging focusing on MRI

    Alexander Selvikvåg Lundervold, Arvid Lundervold

    cs.CVcs.LGstat.MLarXiv:1811.10052v22018
  21. A Deep Learning Framework for Unsupervised Affine and Deformable Image Registration

    Bob D. de Vos, Floris F. Berendsen, Max A. Viergever +3

    cs.CVarXiv:1809.06130v22018
  22. Confounding variables can degrade generalization performance of radiological deep learning models

    John R. Zech, Marcus A. Badgeley, Manway Liu +3

    cs.CVcs.LGstat.MLarXiv:1807.00431v22018
  23. OFF-ApexNet on Micro-expression Recognition System

    Sze-Teng Liong, Y. S. Gan, Wei-Chuen Yau +2

    cs.CVcs.LGarXiv:1805.08699v12018
  24. Beyond RGB: Very High Resolution Urban Remote Sensing With Multimodal Deep Networks

    Nicolas Audebert, Bertrand Le Saux, Sébastien Lefèvre

    cs.NEcs.CVarXiv:1711.08681v12017
  25. Remote Sensing Image Fusion Based on Two-stream Fusion Network

    Xiangyu Liu, Qingjie Liu, Yunhong Wang

    cs.CVarXiv:1711.02549v32017
  26. Deep Residual Bidir-LSTM for Human Activity Recognition Using Wearable Sensors

    Yu Zhao, Rennong Yang, Guillaume Chevalier +1

    cs.CVcs.LGarXiv:1708.08989v22017
  27. Learning Features for Offline Handwritten Signature Verification using Deep Convolutional Neural Networks

    Luiz G. Hafemann, Robert Sabourin, Luiz S. Oliveira

    cs.CVarXiv:1705.05787v12017
  28. Quicksilver: Fast Predictive Image Registration - a Deep Learning Approach

    Xiao Yang, Roland Kwitt, Martin Styner +1

    cs.CVarXiv:1703.10908v42017
  29. Algorithms for Semantic Segmentation of Multispectral Remote Sensing Imagery using Deep Learning

    Ronald Kemker, Carl Salvaggio, Christopher Kanan

    cs.CVcs.AIarXiv:1703.06452v32017
  30. Deep-Learning for Classification of Colorectal Polyps on Whole-Slide Images

    Bruno Korbar, Andrea M. Olofson, Allen P. Miraflor +5

    cs.CVarXiv:1703.01550v22017
  31. 3D fully convolutional networks for subcortical segmentation in MRI: A large-scale study

    J. Dolz, C. Desrosiers, I. Ben Ayed

    cs.CVarXiv:1612.03925v22016
  32. Superpixels: An Evaluation of the State-of-the-Art

    David Stutz, Alexander Hermans, Bastian Leibe

    cs.CVarXiv:1612.01601v32016
  33. Classification With an Edge: Improving Semantic Image Segmentation with Boundary Detection

    Dimitrios Marmanis, Konrad Schindler, Jan Dirk Wegner +3

    cs.CVarXiv:1612.01337v22016
  34. UniMiB SHAR: a new dataset for human activity recognition using acceleration data from smartphones

    Daniela Micucci, Marco Mobilio, Paolo Napoletano

    cs.CVarXiv:1611.07688v52016
  35. AutoInt: Automatic Integration for Fast Neural Volume Rendering

    David B. Lindell, Julien N. P. Martel, Gordon Wetzstein

    cs.CVcs.GRcs.LGarXiv:2012.01714v22020
  36. Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset

    Jing Lin, Ailing Zeng, Shunlin Lu +4

    cs.CVarXiv:2307.00818v22023
  37. Deep-Anomaly: Fully Convolutional Neural Network for Fast Anomaly Detection in Crowded Scenes

    Mohammad Sabokrou, Mohsen Fayyaz, Mahmood Fathy +2

    cs.CVarXiv:1609.00866v22016
  38. Less is More: Micro-expression Recognition from Video using Apex Frame

    Sze-Teng Liong, John See, KokSheik Wong +1

    cs.CVarXiv:1606.01721v32016
  39. A Combined Deep-Learning and Deformable-Model Approach to Fully Automatic Segmentation of the Left Ventricle in Cardiac MRI

    M. R. Avendi, A. Kheradvar, H. Jafarkhani

    cs.CVarXiv:1512.07951v12015
  40. Deep Feature Learning with Relative Distance Comparison for Person Re-identification

    Shengyong Ding, Liang Lin, Guangrun Wang +1

    cs.CVarXiv:1512.03622v12015
  41. Brain Tumor Segmentation with Deep Neural Networks

    Mohammad Havaei, Axel Davy, David Warde-Farley +6

    cs.CVcs.AIarXiv:1505.03540v32015
  42. Deep Learning for Classification and Severity Estimation of Coffee Leaf Biotic Stress

    J. G. M. Esgario, R. A. Krohling, J. A. Ventura

    cs.CVcs.LGarXiv:1907.11561v12019
  43. Multi-Scale Structure-Aware Network for Human Pose Estimation

    Lipeng Ke, Ming-Ching Chang, Honggang Qi +1

    cs.CVarXiv:1803.09894v32018
  44. Adversarial Manipulation of Deep Representations

    Sara Sabour, Yanshuai Cao, Fartash Faghri +1

    cs.CVcs.LGcs.NEarXiv:1511.05122v92015
  45. Learning From Noisy Labels By Regularized Estimation Of Annotator Confusion

    Ryutaro Tanno, Ardavan Saeedi, Swami Sankaranarayanan +2

    cs.LGcs.CVstat.MLarXiv:1902.03680v32019
  46. Rotation equivariant vector field networks

    Diego Marcos, Michele Volpi, Nikos Komodakis +1

    cs.CVarXiv:1612.09346v32016
  47. GIFT: A Real-time and Scalable 3D Shape Search Engine

    Song Bai, Xiang Bai, Zhichao Zhou +2

    cs.CVarXiv:1604.01879v22016
  48. PKU-MMD: A Large Scale Benchmark for Continuous Multi-Modal Human Action Understanding

    Chunhui Liu, Yueyu Hu, Yanghao Li +2

    cs.CVarXiv:1703.07475v22017
  49. Text2Shape: Generating Shapes from Natural Language by Learning Joint Embeddings

    Kevin Chen, Christopher B. Choy, Manolis Savva +3

    cs.CVcs.AIcs.GRarXiv:1803.08495v12018
  50. Visual Commonsense R-CNN

    Tan Wang, Jianqiang Huang, Hanwang Zhang +1

    cs.CVarXiv:2002.12204v32020
  51. PIRenderer: Controllable Portrait Image Generation via Semantic Neural Rendering

    Yurui Ren, Ge Li, Yuanqi Chen +2

    cs.CVcs.AIarXiv:2109.08379v12021
  52. MeshTalk: 3D Face Animation from Speech using Cross-Modality Disentanglement

    Alexander Richard, Michael Zollhoefer, Yandong Wen +2

    cs.CVarXiv:2104.08223v22021
  53. Tutel: Adaptive Mixture-of-Experts at Scale

    Changho Hwang, Wei Cui, Yifan Xiong +12

    cs.DCcs.CLcs.CVarXiv:2206.03382v22022
  54. Synthesizing Training Images for Boosting Human 3D Pose Estimation

    Wenzheng Chen, Huan Wang, Yangyan Li +6

    cs.CVarXiv:1604.02703v62016
  55. DoubleFusion: Real-time Capture of Human Performances with Inner Body Shapes from a Single Depth Sensor

    Tao Yu, Zerong Zheng, Kaiwen Guo +5

    cs.CVarXiv:1804.06023v12018
  56. Visual Interaction Networks

    Nicholas Watters, Andrea Tacchetti, Theophane Weber +3

    cs.CVarXiv:1706.01433v12017
  57. Reversible Architectures for Arbitrarily Deep Residual Neural Networks

    Bo Chang, Lili Meng, Eldad Haber +3

    cs.CVstat.MLarXiv:1709.03698v22017
  58. Scene Transformer: A unified architecture for predicting multiple agent trajectories

    Jiquan Ngiam, Benjamin Caine, Vijay Vasudevan +11

    cs.CVcs.LGcs.ROarXiv:2106.08417v32021
  59. Cyclical Stochastic Gradient MCMC for Bayesian Deep Learning

    Ruqi Zhang, Chunyuan Li, Jianyi Zhang +2

    cs.LGcs.AIcs.CVarXiv:1902.03932v22019
  60. EC-Net: an Edge-aware Point set Consolidation Network

    Lequan Yu, Xianzhi Li, Chi-Wing Fu +2

    cs.CVarXiv:1807.06010v12018