Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

10,861 to 10,920 of 18,849

  1. Towards More Flexible and Accurate Object Tracking with Natural Language: Algorithms and Benchmark

    Xiao Wang, Xiujun Shu, Zhipeng Zhang +4

    cs.CVcs.AIarXiv:2103.16746v12021
  2. RGB-D Salient Object Detection: A Survey

    Tao Zhou, Deng-Ping Fan, Ming-Ming Cheng +2

    cs.CVarXiv:2008.00230v42020
  3. Evolution of Image Segmentation using Deep Convolutional Neural Network: A Survey

    Farhana Sultana, Abu Sufian, Paramartha Dutta

    cs.CVarXiv:2001.04074v32020
  4. Plant Diseases recognition on images using Convolutional Neural Networks: A Systematic Review

    Andre S. Abade, Paulo Afonso Ferreira, Flavio de Barros Vidal

    cs.CVarXiv:2009.04365v12020
  5. A Survey on Deep Learning Technique for Video Segmentation

    Tianfei Zhou, Fatih Porikli, David Crandall +2

    cs.CVarXiv:2107.01153v42021
  6. NeRF: Neural Radiance Field in 3D Vision: A Comprehensive Review (Updated Post-Gaussian Splatting)

    Kyle Gao, Yina Gao, Hongjie He +3

    cs.CVarXiv:2210.00379v82022
  7. Multi-task CNN Model for Attribute Prediction

    Abrar H. Abdulnabi, Gang Wang, Jiwen Lu +1

    cs.CVarXiv:1601.00400v12016
  8. Learning Texture Invariant Representation for Domain Adaptation of Semantic Segmentation

    Myeongjin Kim, Hyeran Byun

    cs.CVarXiv:2003.00867v22020
  9. Multi-Label Zero-Shot Learning with Structured Knowledge Graphs

    Chung-Wei Lee, Wei Fang, Chih-Kuan Yeh +1

    cs.CVarXiv:1711.06526v22017
  10. Open-Sora Plan: Open-Source Large Video Generation Model

    Bin Lin, Yunyang Ge, Xinhua Cheng +21

    cs.CVcs.AIarXiv:2412.00131v12024
  11. MISSFormer: An Effective Medical Image Segmentation Transformer

    Xiaohong Huang, Zhifang Deng, Dandan Li +1

    cs.CVarXiv:2109.07162v22021
  12. HarDNet: A Low Memory Traffic Network

    Ping Chao, Chao-Yang Kao, Yu-Shan Ruan +2

    cs.CVarXiv:1909.00948v12019
  13. Multi-class Token Transformer for Weakly Supervised Semantic Segmentation

    Lian Xu, Wanli Ouyang, Mohammed Bennamoun +2

    cs.CVarXiv:2203.02891v12022
  14. Feature Pyramid Transformer

    Dong Zhang, Hanwang Zhang, Jinhui Tang +3

    cs.CVarXiv:2007.09451v12020
  15. Coherent Online Video Style Transfer

    Dongdong Chen, Jing Liao, Lu Yuan +2

    cs.CVarXiv:1703.09211v22017
  16. Language Conditioned Imitation Learning over Unstructured Data

    Corey Lynch, Pierre Sermanet

    cs.ROcs.AIcs.CLarXiv:2005.07648v22020
  17. Evo-ViT: Slow-Fast Token Evolution for Dynamic Vision Transformer

    Yifan Xu, Zhijie Zhang, Mengdan Zhang +6

    cs.CVarXiv:2108.01390v52021
  18. LCR-Net++: Multi-person 2D and 3D Pose Detection in Natural Images

    Gregory Rogez, Philippe Weinzaepfel, Cordelia Schmid

    cs.CVarXiv:1803.00455v32018
  19. Vision Transformer for Small-Size Datasets

    Seung Hoon Lee, Seunghyun Lee, Byung Cheol Song

    cs.CVarXiv:2112.13492v12021
  20. Towards Faster Training of Global Covariance Pooling Networks by Iterative Matrix Square Root Normalization

    Peihua Li, Jiangtao Xie, Qilong Wang +1

    cs.CVarXiv:1712.01034v22017
  21. DLow: Diversifying Latent Flows for Diverse Human Motion Prediction

    Ye Yuan, Kris Kitani

    cs.CVcs.LGeess.IVarXiv:2003.08386v22020
  22. Semantic Segmentation using Vision Transformers: A survey

    Hans Thisanke, Chamli Deshan, Kavindu Chamith +3

    cs.CVcs.AIcs.LGarXiv:2305.03273v12023
  23. Fast L1-Minimization Algorithms For Robust Face Recognition

    Allen Y. Yang, Zihan Zhou, Arvind Ganesh +2

    cs.CVmath.NAarXiv:1007.3753v42010
  24. An Empirical Study of Remote Sensing Pretraining

    Di Wang, Jing Zhang, Bo Du +2

    cs.CVarXiv:2204.02825v42022
  25. Improving Semantic Segmentation via Decoupled Body and Edge Supervision

    Xiangtai Li, Xia Li, Li Zhang +5

    cs.CVarXiv:2007.10035v22020
  26. 3D Shape Induction from 2D Views of Multiple Objects

    Matheus Gadelha, Subhransu Maji, Rui Wang

    cs.CVarXiv:1612.05872v12016
  27. Hierarchical Clustering with Hard-batch Triplet Loss for Person Re-identification

    Kaiwei Zeng

    cs.CVarXiv:1910.12278v22019
  28. Learning Object-Compositional Neural Radiance Field for Editable Scene Rendering

    Bangbang Yang, Yinda Zhang, Yinghao Xu +5

    cs.CVarXiv:2109.01847v12021
  29. Overcoming Classifier Imbalance for Long-tail Object Detection with Balanced Group Softmax

    Yu Li, Tao Wang, Bingyi Kang +4

    cs.CVcs.LGstat.MLarXiv:2006.10408v12020
  30. MOTRv2: Bootstrapping End-to-End Multi-Object Tracking by Pretrained Object Detectors

    Yuang Zhang, Tiancai Wang, Xiangyu Zhang

    cs.CVarXiv:2211.09791v22022
  31. Decoupling Zero-Shot Semantic Segmentation

    Jian Ding, Nan Xue, Gui-Song Xia +1

    cs.CVarXiv:2112.07910v22021
  32. NDDR-CNN: Layerwise Feature Fusing in Multi-Task CNNs by Neural Discriminative Dimensionality Reduction

    Yuan Gao, Jiayi Ma, Mingbo Zhao +2

    cs.CVcs.LGarXiv:1801.08297v42018
  33. Zero-Shot Visual Recognition using Semantics-Preserving Adversarial Embedding Networks

    Long Chen, Hanwang Zhang, Jun Xiao +2

    cs.CVarXiv:1712.01928v22017
  34. Compressed 3D Gaussian Splatting for Accelerated Novel View Synthesis

    Simon Niedermayr, Josef Stumpfegger, Rüdiger Westermann

    cs.CVcs.GRarXiv:2401.02436v22023
  35. Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image Models

    Lukas Höllein, Ang Cao, Andrew Owens +2

    cs.CVarXiv:2303.11989v22023
  36. Hidden Two-Stream Convolutional Networks for Action Recognition

    Yi Zhu, Zhenzhong Lan, Shawn Newsam +1

    cs.CVcs.LGcs.MMarXiv:1704.00389v42017
  37. Weakly Supervised Cascaded Convolutional Networks

    Ali Diba, Vivek Sharma, Ali Pazandeh +2

    cs.CVarXiv:1611.08258v12016
  38. Lipreading using Temporal Convolutional Networks

    Brais Martinez, Pingchuan Ma, Stavros Petridis +1

    cs.CVcs.SDeess.ASarXiv:2001.08702v12020
  39. Outlining where humans live -- The World Settlement Footprint 2015

    Mattia Marconcini, Annekatrin Metz-Marconcini, Soner Üreyen +8

    eess.IVcs.CVarXiv:1910.12707v12019
  40. Video Summarization Using Deep Neural Networks: A Survey

    Evlampios Apostolidis, Eleni Adamantidou, Alexandros I. Metsai +2

    cs.CVcs.LGcs.MMarXiv:2101.06072v22021
  41. (AF)2-S3Net: Attentive Feature Fusion with Adaptive Feature Selection for Sparse Semantic Segmentation Network

    Ran Cheng, Ryan Razani, Ehsan Taghavi +2

    cs.CVcs.AIcs.ROarXiv:2102.04530v12021
  42. weedNet: Dense Semantic Weed Classification Using Multispectral Images and MAV for Smart Farming

    Inkyu Sa, Zetao Chen, Marija Popovic +4

    cs.CVcs.ROarXiv:1709.03329v12017
  43. MeshGPT: Generating Triangle Meshes with Decoder-Only Transformers

    Yawar Siddiqui, Antonio Alliegro, Alexey Artemov +5

    cs.CVcs.LGarXiv:2311.15475v12023
  44. FireCaffe: near-linear acceleration of deep neural network training on compute clusters

    Forrest N. Iandola, Khalid Ashraf, Matthew W. Moskewicz +1

    cs.CVarXiv:1511.00175v22015
  45. Physically Realizable Adversarial Examples for LiDAR Object Detection

    James Tu, Mengye Ren, Siva Manivasagam +5

    cs.CVcs.CRcs.LGarXiv:2004.00543v22020
  46. A Comprehensive Study on Colorectal Polyp Segmentation with ResUNet++, Conditional Random Field and Test-Time Augmentation

    Debesh Jha, Pia H. Smedsrud, Dag Johansen +4

    cs.CVarXiv:2107.12435v12021
  47. Category Anchor-Guided Unsupervised Domain Adaptation for Semantic Segmentation

    Qiming Zhang, Jing Zhang, Wei Liu +1

    cs.CVarXiv:1910.13049v22019
  48. Weakly-Supervised Semantic Segmentation via Sub-category Exploration

    Yu-Ting Chang, Qiaosong Wang, Wei-Chih Hung +3

    cs.CVcs.LGeess.IVarXiv:2008.01183v12020
  49. Unmasking the abnormal events in video

    Radu Tudor Ionescu, Sorina Smeureanu, Bogdan Alexe +1

    cs.CVarXiv:1705.08182v32017
  50. Each Part Matters: Local Patterns Facilitate Cross-view Geo-localization

    Tingyu Wang, Zhedong Zheng, Chenggang Yan +4

    cs.CVcs.LGarXiv:2008.11646v32020
  51. Recent Advances in 3D Gaussian Splatting

    Tong Wu, Yu-Jie Yuan, Ling-Xiao Zhang +4

    cs.CVcs.GRarXiv:2403.11134v22024
  52. Channel prior convolutional attention for medical image segmentation

    Hejun Huang, Zuguo Chen, Ying Zou +2

    eess.IVcs.CVarXiv:2306.05196v12023
  53. Exploring the structure of a real-time, arbitrary neural artistic stylization network

    Golnaz Ghiasi, Honglak Lee, Manjunath Kudlur +2

    cs.CVarXiv:1705.06830v22017
  54. Embedding Deep Metric for Person Re-identication A Study Against Large Variations

    Hailin Shi, Yang Yang, Xiangyu Zhu +4

    cs.CVcs.LGarXiv:1611.00137v12016
  55. Spatio-temporal video autoencoder with differentiable memory

    Viorica Patraucean, Ankur Handa, Roberto Cipolla

    cs.LGcs.CVarXiv:1511.06309v52015
  56. OpenVSLAM: A Versatile Visual SLAM Framework

    Shinya Sumikura, Mikiya Shibuya, Ken Sakurada

    cs.CVcs.ROarXiv:1910.01122v32019
  57. DF-GAN: A Simple and Effective Baseline for Text-to-Image Synthesis

    Ming Tao, Hao Tang, Fei Wu +3

    cs.CVarXiv:2008.05865v42020
  58. MIMIC-IT: Multi-Modal In-Context Instruction Tuning

    Bo Li, Yuanhan Zhang, Liangyu Chen +5

    cs.CVcs.AIcs.CLarXiv:2306.05425v12023
  59. MinerU: An Open-Source Solution for Precise Document Content Extraction

    Bin Wang, Chao Xu, Xiaomeng Zhao +15

    cs.CVarXiv:2409.18839v12024
  60. Tool Detection and Operative Skill Assessment in Surgical Videos Using Region-Based Convolutional Neural Networks

    Amy Jin, Serena Yeung, Jeffrey Jopling +4

    cs.CVarXiv:1802.08774v22018