Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
10,861 to 10,920 of 18,849
Towards More Flexible and Accurate Object Tracking with Natural Language: Algorithms and Benchmark
Xiao Wang, Xiujun Shu, Zhipeng Zhang +4
cs.CVcs.AIarXiv:2103.16746v12021RGB-D Salient Object Detection: A Survey
Tao Zhou, Deng-Ping Fan, Ming-Ming Cheng +2
cs.CVarXiv:2008.00230v42020Evolution of Image Segmentation using Deep Convolutional Neural Network: A Survey
Farhana Sultana, Abu Sufian, Paramartha Dutta
cs.CVarXiv:2001.04074v32020Plant Diseases recognition on images using Convolutional Neural Networks: A Systematic Review
Andre S. Abade, Paulo Afonso Ferreira, Flavio de Barros Vidal
cs.CVarXiv:2009.04365v12020A Survey on Deep Learning Technique for Video Segmentation
Tianfei Zhou, Fatih Porikli, David Crandall +2
cs.CVarXiv:2107.01153v42021NeRF: Neural Radiance Field in 3D Vision: A Comprehensive Review (Updated Post-Gaussian Splatting)
Kyle Gao, Yina Gao, Hongjie He +3
cs.CVarXiv:2210.00379v82022Multi-task CNN Model for Attribute Prediction
Abrar H. Abdulnabi, Gang Wang, Jiwen Lu +1
cs.CVarXiv:1601.00400v12016Learning Texture Invariant Representation for Domain Adaptation of Semantic Segmentation
Myeongjin Kim, Hyeran Byun
cs.CVarXiv:2003.00867v22020Multi-Label Zero-Shot Learning with Structured Knowledge Graphs
Chung-Wei Lee, Wei Fang, Chih-Kuan Yeh +1
cs.CVarXiv:1711.06526v22017Open-Sora Plan: Open-Source Large Video Generation Model
Bin Lin, Yunyang Ge, Xinhua Cheng +21
cs.CVcs.AIarXiv:2412.00131v12024MISSFormer: An Effective Medical Image Segmentation Transformer
Xiaohong Huang, Zhifang Deng, Dandan Li +1
cs.CVarXiv:2109.07162v22021HarDNet: A Low Memory Traffic Network
Ping Chao, Chao-Yang Kao, Yu-Shan Ruan +2
cs.CVarXiv:1909.00948v12019Multi-class Token Transformer for Weakly Supervised Semantic Segmentation
Lian Xu, Wanli Ouyang, Mohammed Bennamoun +2
cs.CVarXiv:2203.02891v12022Feature Pyramid Transformer
Dong Zhang, Hanwang Zhang, Jinhui Tang +3
cs.CVarXiv:2007.09451v12020Coherent Online Video Style Transfer
Dongdong Chen, Jing Liao, Lu Yuan +2
cs.CVarXiv:1703.09211v22017Language Conditioned Imitation Learning over Unstructured Data
Corey Lynch, Pierre Sermanet
cs.ROcs.AIcs.CLarXiv:2005.07648v22020Evo-ViT: Slow-Fast Token Evolution for Dynamic Vision Transformer
Yifan Xu, Zhijie Zhang, Mengdan Zhang +6
cs.CVarXiv:2108.01390v52021LCR-Net++: Multi-person 2D and 3D Pose Detection in Natural Images
Gregory Rogez, Philippe Weinzaepfel, Cordelia Schmid
cs.CVarXiv:1803.00455v32018Vision Transformer for Small-Size Datasets
Seung Hoon Lee, Seunghyun Lee, Byung Cheol Song
cs.CVarXiv:2112.13492v12021Towards Faster Training of Global Covariance Pooling Networks by Iterative Matrix Square Root Normalization
Peihua Li, Jiangtao Xie, Qilong Wang +1
cs.CVarXiv:1712.01034v22017DLow: Diversifying Latent Flows for Diverse Human Motion Prediction
Ye Yuan, Kris Kitani
cs.CVcs.LGeess.IVarXiv:2003.08386v22020Semantic Segmentation using Vision Transformers: A survey
Hans Thisanke, Chamli Deshan, Kavindu Chamith +3
cs.CVcs.AIcs.LGarXiv:2305.03273v12023Fast L1-Minimization Algorithms For Robust Face Recognition
Allen Y. Yang, Zihan Zhou, Arvind Ganesh +2
cs.CVmath.NAarXiv:1007.3753v42010An Empirical Study of Remote Sensing Pretraining
Di Wang, Jing Zhang, Bo Du +2
cs.CVarXiv:2204.02825v42022Improving Semantic Segmentation via Decoupled Body and Edge Supervision
Xiangtai Li, Xia Li, Li Zhang +5
cs.CVarXiv:2007.10035v220203D Shape Induction from 2D Views of Multiple Objects
Matheus Gadelha, Subhransu Maji, Rui Wang
cs.CVarXiv:1612.05872v12016Hierarchical Clustering with Hard-batch Triplet Loss for Person Re-identification
Kaiwei Zeng
cs.CVarXiv:1910.12278v22019Learning Object-Compositional Neural Radiance Field for Editable Scene Rendering
Bangbang Yang, Yinda Zhang, Yinghao Xu +5
cs.CVarXiv:2109.01847v12021Overcoming Classifier Imbalance for Long-tail Object Detection with Balanced Group Softmax
Yu Li, Tao Wang, Bingyi Kang +4
cs.CVcs.LGstat.MLarXiv:2006.10408v12020MOTRv2: Bootstrapping End-to-End Multi-Object Tracking by Pretrained Object Detectors
Yuang Zhang, Tiancai Wang, Xiangyu Zhang
cs.CVarXiv:2211.09791v22022Decoupling Zero-Shot Semantic Segmentation
Jian Ding, Nan Xue, Gui-Song Xia +1
cs.CVarXiv:2112.07910v22021NDDR-CNN: Layerwise Feature Fusing in Multi-Task CNNs by Neural Discriminative Dimensionality Reduction
Yuan Gao, Jiayi Ma, Mingbo Zhao +2
cs.CVcs.LGarXiv:1801.08297v42018Zero-Shot Visual Recognition using Semantics-Preserving Adversarial Embedding Networks
Long Chen, Hanwang Zhang, Jun Xiao +2
cs.CVarXiv:1712.01928v22017Compressed 3D Gaussian Splatting for Accelerated Novel View Synthesis
Simon Niedermayr, Josef Stumpfegger, Rüdiger Westermann
cs.CVcs.GRarXiv:2401.02436v22023Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image Models
Lukas Höllein, Ang Cao, Andrew Owens +2
cs.CVarXiv:2303.11989v22023Hidden Two-Stream Convolutional Networks for Action Recognition
Yi Zhu, Zhenzhong Lan, Shawn Newsam +1
cs.CVcs.LGcs.MMarXiv:1704.00389v42017Weakly Supervised Cascaded Convolutional Networks
Ali Diba, Vivek Sharma, Ali Pazandeh +2
cs.CVarXiv:1611.08258v12016Lipreading using Temporal Convolutional Networks
Brais Martinez, Pingchuan Ma, Stavros Petridis +1
cs.CVcs.SDeess.ASarXiv:2001.08702v12020Outlining where humans live -- The World Settlement Footprint 2015
Mattia Marconcini, Annekatrin Metz-Marconcini, Soner Üreyen +8
eess.IVcs.CVarXiv:1910.12707v12019Video Summarization Using Deep Neural Networks: A Survey
Evlampios Apostolidis, Eleni Adamantidou, Alexandros I. Metsai +2
cs.CVcs.LGcs.MMarXiv:2101.06072v22021(AF)2-S3Net: Attentive Feature Fusion with Adaptive Feature Selection for Sparse Semantic Segmentation Network
Ran Cheng, Ryan Razani, Ehsan Taghavi +2
cs.CVcs.AIcs.ROarXiv:2102.04530v12021weedNet: Dense Semantic Weed Classification Using Multispectral Images and MAV for Smart Farming
Inkyu Sa, Zetao Chen, Marija Popovic +4
cs.CVcs.ROarXiv:1709.03329v12017MeshGPT: Generating Triangle Meshes with Decoder-Only Transformers
Yawar Siddiqui, Antonio Alliegro, Alexey Artemov +5
cs.CVcs.LGarXiv:2311.15475v12023FireCaffe: near-linear acceleration of deep neural network training on compute clusters
Forrest N. Iandola, Khalid Ashraf, Matthew W. Moskewicz +1
cs.CVarXiv:1511.00175v22015Physically Realizable Adversarial Examples for LiDAR Object Detection
James Tu, Mengye Ren, Siva Manivasagam +5
cs.CVcs.CRcs.LGarXiv:2004.00543v22020A Comprehensive Study on Colorectal Polyp Segmentation with ResUNet++, Conditional Random Field and Test-Time Augmentation
Debesh Jha, Pia H. Smedsrud, Dag Johansen +4
cs.CVarXiv:2107.12435v12021Category Anchor-Guided Unsupervised Domain Adaptation for Semantic Segmentation
Qiming Zhang, Jing Zhang, Wei Liu +1
cs.CVarXiv:1910.13049v22019Weakly-Supervised Semantic Segmentation via Sub-category Exploration
Yu-Ting Chang, Qiaosong Wang, Wei-Chih Hung +3
cs.CVcs.LGeess.IVarXiv:2008.01183v12020Unmasking the abnormal events in video
Radu Tudor Ionescu, Sorina Smeureanu, Bogdan Alexe +1
cs.CVarXiv:1705.08182v32017Each Part Matters: Local Patterns Facilitate Cross-view Geo-localization
Tingyu Wang, Zhedong Zheng, Chenggang Yan +4
cs.CVcs.LGarXiv:2008.11646v32020Recent Advances in 3D Gaussian Splatting
Tong Wu, Yu-Jie Yuan, Ling-Xiao Zhang +4
cs.CVcs.GRarXiv:2403.11134v22024Channel prior convolutional attention for medical image segmentation
Hejun Huang, Zuguo Chen, Ying Zou +2
eess.IVcs.CVarXiv:2306.05196v12023Exploring the structure of a real-time, arbitrary neural artistic stylization network
Golnaz Ghiasi, Honglak Lee, Manjunath Kudlur +2
cs.CVarXiv:1705.06830v22017Embedding Deep Metric for Person Re-identication A Study Against Large Variations
Hailin Shi, Yang Yang, Xiangyu Zhu +4
cs.CVcs.LGarXiv:1611.00137v12016Spatio-temporal video autoencoder with differentiable memory
Viorica Patraucean, Ankur Handa, Roberto Cipolla
cs.LGcs.CVarXiv:1511.06309v52015OpenVSLAM: A Versatile Visual SLAM Framework
Shinya Sumikura, Mikiya Shibuya, Ken Sakurada
cs.CVcs.ROarXiv:1910.01122v32019DF-GAN: A Simple and Effective Baseline for Text-to-Image Synthesis
Ming Tao, Hao Tang, Fei Wu +3
cs.CVarXiv:2008.05865v42020MIMIC-IT: Multi-Modal In-Context Instruction Tuning
Bo Li, Yuanhan Zhang, Liangyu Chen +5
cs.CVcs.AIcs.CLarXiv:2306.05425v12023MinerU: An Open-Source Solution for Precise Document Content Extraction
Bin Wang, Chao Xu, Xiaomeng Zhao +15
cs.CVarXiv:2409.18839v12024Tool Detection and Operative Skill Assessment in Surgical Videos Using Region-Based Convolutional Neural Networks
Amy Jin, Serena Yeung, Jeffrey Jopling +4
cs.CVarXiv:1802.08774v22018