Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

12,901 to 12,960 of 18,821

  1. Learning Uncertain Convolutional Features for Accurate Saliency Detection

    Pingping Zhang, Dong Wang, Huchuan Lu +2

    cs.CVarXiv:1708.02031v12017
  2. FastReID: A Pytorch Toolbox for General Instance Re-identification

    Lingxiao He, Xingyu Liao, Wu Liu +3

    cs.CVarXiv:2006.02631v42020
  3. Data Distillation: Towards Omni-Supervised Learning

    Ilija Radosavovic, Piotr Dollár, Ross Girshick +2

    cs.CVarXiv:1712.04440v12017
  4. Towards Discriminability and Diversity: Batch Nuclear-norm Maximization under Label Insufficient Situations

    Shuhao Cui, Shuhui Wang, Junbao Zhuo +3

    cs.CVarXiv:2003.12237v12020
  5. MultiPath++: Efficient Information Fusion and Trajectory Aggregation for Behavior Prediction

    Balakrishnan Varadarajan, Ahmed Hefny, Avikalp Srivastava +8

    cs.CVcs.AIcs.LGarXiv:2111.14973v32021
  6. Pseudo-LiDAR++: Accurate Depth for 3D Object Detection in Autonomous Driving

    Yurong You, Yan Wang, Wei-Lun Chao +5

    cs.CVarXiv:1906.06310v32019
  7. Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis

    Weixi Feng, Xuehai He, Tsu-Jui Fu +6

    cs.CVcs.CLarXiv:2212.05032v32022
  8. R2M-Bench: Evaluating Revisit Memory via Relative Consistency in Interactive Video World Models

    Qiwen Gu, Bingjie Gao, Rui Chen +7

    cs.CVarXiv:2608.27328v12026
  9. Deep Multi-View Enhancement Hashing for Image Retrieval

    Chenggang Yan, Biao Gong, Yuxuan Wei +1

    cs.CVcs.LGeess.IVarXiv:2002.00169v22020
  10. Action Genome: Actions as Composition of Spatio-temporal Scene Graphs

    Jingwei Ji, Ranjay Krishna, Li Fei-Fei +1

    cs.CVarXiv:1912.06992v12019
  11. Deep Gaze I: Boosting Saliency Prediction with Feature Maps Trained on ImageNet

    Matthias Kümmerer, Lucas Theis, Matthias Bethge

    cs.CVq-bio.NCstat.AParXiv:1411.1045v42014
  12. Amortised MAP Inference for Image Super-resolution

    Casper Kaae Sønderby, Jose Caballero, Lucas Theis +2

    cs.CVcs.LGstat.MLarXiv:1610.04490v32016
  13. Directly Fine-Tuning Diffusion Models on Differentiable Rewards

    Kevin Clark, Paul Vicol, Kevin Swersky +1

    cs.CVcs.LGarXiv:2309.17400v22023
  14. TM2T: Stochastic and Tokenized Modeling for the Reciprocal Generation of 3D Human Motions and Texts

    Chuan Guo, Xinxin Zuo, Sen Wang +1

    cs.CVarXiv:2207.01696v22022
  15. Hypercorrelation Squeeze for Few-Shot Segmentation

    Juhong Min, Dahyun Kang, Minsu Cho

    cs.CVarXiv:2104.01538v32021
  16. DeepID-Net: Deformable Deep Convolutional Neural Networks for Object Detection

    Wanli Ouyang, Xiaogang Wang, Xingyu Zeng +8

    cs.CVcs.NEarXiv:1412.5661v22014
  17. Text-to-image Diffusion Models in Generative AI: A Survey

    Chenshuang Zhang, Chaoning Zhang, Mengchun Zhang +2

    cs.CVcs.AIcs.LGarXiv:2303.07909v32023
  18. Residual and Plain Convolutional Neural Networks for 3D Brain MRI Classification

    Sergey Korolev, Amir Safiullin, Mikhail Belyaev +1

    cs.CVarXiv:1701.06643v12017
  19. Data augmentation using learned transformations for one-shot medical image segmentation

    Amy Zhao, Guha Balakrishnan, Frédo Durand +2

    cs.CVarXiv:1902.09383v22019
  20. MedNeXt: Transformer-driven Scaling of ConvNets for Medical Image Segmentation

    Saikat Roy, Gregor Koehler, Constantin Ulrich +5

    eess.IVcs.CVcs.LGarXiv:2303.09975v52023
  21. Comparison of Deep Learning Approaches for Multi-Label Chest X-Ray Classification

    Ivo M. Baltruschat, Hannes Nickisch, Michael Grass +2

    cs.CVarXiv:1803.02315v22018
  22. Self-Correction for Human Parsing

    Peike Li, Yunqiu Xu, Yunchao Wei +1

    cs.CVcs.LGeess.IVarXiv:1910.09777v12019
  23. Learning by tracking: Siamese CNN for robust target association

    Laura Leal-Taixé, Cristian Canton Ferrer, Konrad Schindler

    cs.LGcs.CVarXiv:1604.07866v32016
  24. Temporal Pyramid Network for Action Recognition

    Ceyuan Yang, Yinghao Xu, Jianping Shi +2

    cs.CVarXiv:2004.03548v22020
  25. Seven ways to improve example-based single image super resolution

    Radu Timofte, Rasmus Rothe, Luc Van Gool

    cs.CVarXiv:1511.02228v12015
  26. AraMS-28k: The Largest Publicly Released Line-Level Dataset of Historical Arabic Manuscripts with Margin and Insertion-Anchor Annotations

    Mohamed Guechaoui, Mohamed Diaa Zellagui, Souleyman Chaib +1

    cs.CVcs.CLarXiv:2608.26921v12026
  27. EditaLive! Unified Character Video Editing for Live Streaming

    Zhiyuan Li, Chi-Man Pun, Peng-Tao Jiang +2

    cs.CVarXiv:2608.27123v12026
  28. StyleRig: Rigging StyleGAN for 3D Control over Portrait Images

    Ayush Tewari, Mohamed Elgharib, Gaurav Bharaj +5

    cs.CVcs.GRarXiv:2004.00121v22020
  29. Semantic Flow for Fast and Accurate Scene Parsing

    Xiangtai Li, Ansheng You, Zhen Zhu +4

    cs.CVcs.ROarXiv:2002.10120v32020
  30. GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splatting

    Yiwen Chen, Zilong Chen, Chi Zhang +7

    cs.CVarXiv:2311.14521v42023
  31. Deep Depth Completion of a Single RGB-D Image

    Yinda Zhang, Thomas Funkhouser

    cs.CVarXiv:1803.09326v22018
  32. RoboNet: Large-Scale Multi-Robot Learning

    Sudeep Dasari, Frederik Ebert, Stephen Tian +6

    cs.ROcs.CVcs.LGarXiv:1910.11215v22019
  33. Predicting Deep Zero-Shot Convolutional Neural Networks using Textual Descriptions

    Jimmy Ba, Kevin Swersky, Sanja Fidler +1

    cs.LGcs.CVcs.NEarXiv:1506.00511v22015
  34. VOS: Learning What You Don't Know by Virtual Outlier Synthesis

    Xuefeng Du, Zhaoning Wang, Mu Cai +1

    cs.LGcs.CVarXiv:2202.01197v42022
  35. CycleISP: Real Image Restoration via Improved Data Synthesis

    Syed Waqas Zamir, Aditya Arora, Salman Khan +4

    eess.IVcs.CVarXiv:2003.07761v12020
  36. Spacetime Gaussian Feature Splatting for Real-Time Dynamic View Synthesis

    Zhan Li, Zhang Chen, Zhong Li +1

    cs.CVcs.GRarXiv:2312.16812v22023
  37. Mode Seeking Generative Adversarial Networks for Diverse Image Synthesis

    Qi Mao, Hsin-Ying Lee, Hung-Yu Tseng +2

    cs.CVarXiv:1903.05628v62019
  38. A Variational U-Net for Conditional Appearance and Shape Generation

    Patrick Esser, Ekaterina Sutter, Björn Ommer

    cs.CVarXiv:1804.04694v12018
  39. Show, Attend and Read: A Simple and Strong Baseline for Irregular Text Recognition

    Hui Li, Peng Wang, Chunhua Shen +1

    cs.CVarXiv:1811.00751v22018
  40. Use What You Have: Video Retrieval Using Representations From Collaborative Experts

    Yang Liu, Samuel Albanie, Arsha Nagrani +1

    cs.CVarXiv:1907.13487v22019
  41. FIDA: Feature Instability-Driven Attack on Self-Supervised Facial Representation

    Zhiyang Chen, Changchun Yin, Huiqin Yang +1

    cs.CVcs.CRarXiv:2608.26861v12026
  42. VPGNet: Vanishing Point Guided Network for Lane and Road Marking Detection and Recognition

    Seokju Lee, Junsik Kim, Jae Shin Yoon +7

    cs.CVarXiv:1710.06288v12017
  43. Geometry-Aware Learning of Maps for Camera Localization

    Samarth Brahmbhatt, Jinwei Gu, Kihwan Kim +2

    cs.CVarXiv:1712.03342v32017
  44. Semantic Compositional Networks for Visual Captioning

    Zhe Gan, Chuang Gan, Xiaodong He +5

    cs.CVcs.CLcs.LGarXiv:1611.08002v22016
  45. CityFlow: A City-Scale Benchmark for Multi-Target Multi-Camera Vehicle Tracking and Re-Identification

    Zheng Tang, Milind Naphade, Ming-Yu Liu +6

    cs.CVarXiv:1903.09254v42019
  46. Modeling Point Clouds with Self-Attention and Gumbel Subset Sampling

    Jiancheng Yang, Qiang Zhang, Bingbing Ni +4

    cs.CVcs.LGarXiv:1904.03375v12019
  47. Prevalence of Neural Collapse during the terminal phase of deep learning training

    Vardan Papyan, X. Y. Han, David L. Donoho

    cs.LGcs.CVstat.MLarXiv:2008.08186v22020
  48. An Investigation of Why Overparameterization Exacerbates Spurious Correlations

    Shiori Sagawa, Aditi Raghunathan, Pang Wei Koh +1

    cs.LGcs.CVstat.MLarXiv:2005.04345v32020
  49. Omnidata: A Scalable Pipeline for Making Multi-Task Mid-Level Vision Datasets from 3D Scans

    Ainaz Eftekhar, Alexander Sax, Roman Bachmann +2

    cs.CVcs.AIcs.GRarXiv:2110.04994v12021
  50. Robust Registration of Multimodal Remote Sensing Images Based on Structural Similarity

    Yuanxin Ye, Jie Shan, Lorenzo Bruzzone +1

    cs.CVarXiv:2103.16871v12021
  51. RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose

    Tao Jiang, Peng Lu, Li Zhang +5

    cs.CVarXiv:2303.07399v22023
  52. Hallucination of Multimodal Large Language Models: A Survey

    Zechen Bai, Pichao Wang, Tianjun Xiao +4

    cs.CVarXiv:2404.18930v22024
  53. Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles

    Chaitanya Ryali, Yuan-Ting Hu, Daniel Bolya +10

    cs.CVcs.LGarXiv:2306.00989v12023
  54. Measuring Neural Net Robustness with Constraints

    Osbert Bastani, Yani Ioannou, Leonidas Lampropoulos +3

    cs.LGcs.CVcs.NEarXiv:1605.07262v22016
  55. An All-In-One Convolutional Neural Network for Face Analysis

    Rajeev Ranjan, Swami Sankaranarayanan, Carlos D. Castillo +1

    cs.CVarXiv:1611.00851v12016
  56. Diverse Branch Block: Building a Convolution as an Inception-like Unit

    Xiaohan Ding, Xiangyu Zhang, Jungong Han +1

    cs.CVcs.AIcs.LGarXiv:2103.13425v22021
  57. Knowledge Distillation from A Stronger Teacher

    Tao Huang, Shan You, Fei Wang +2

    cs.CVcs.AIcs.LGarXiv:2205.10536v32022
  58. Diagnosing and Enhancing VAE Models

    Bin Dai, David Wipf

    cs.LGcs.CVstat.MLarXiv:1903.05789v22019
  59. STM: SpatioTemporal and Motion Encoding for Action Recognition

    Boyuan Jiang, Mengmeng Wang, Weihao Gan +2

    cs.CVarXiv:1908.02486v22019
  60. AdaAttN: Revisit Attention Mechanism in Arbitrary Neural Style Transfer

    Songhua Liu, Tianwei Lin, Dongliang He +6

    cs.CVarXiv:2108.03647v22021