Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

6,661 to 6,720 of 18,866

  1. PandaSet: Advanced Sensor Suite Dataset for Autonomous Driving

    Pengchuan Xiao, Zhenlei Shao, Steven Hao +9

    cs.CVcs.ROarXiv:2112.12610v12021
  2. Temporal Context Mining for Learned Video Compression

    Xihua Sheng, Jiahao Li, Bin Li +3

    cs.CVcs.LGeess.IVarXiv:2111.13850v22021
  3. From Show to Tell: A Survey on Deep Learning-based Image Captioning

    Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi +3

    cs.CVcs.CLarXiv:2107.06912v32021
  4. Deep Subdomain Adaptation Network for Image Classification

    Yongchun Zhu, Fuzhen Zhuang, Jindong Wang +5

    cs.CVcs.AIarXiv:2106.09388v12021
  5. Structured Denoising Diffusion Models in Discrete State-Spaces

    Jacob Austin, Daniel D. Johnson, Jonathan Ho +2

    cs.LGcs.AIcs.CLarXiv:2107.03006v32021
  6. Recent advances and clinical applications of deep learning in medical image analysis

    Xuxin Chen, Ximin Wang, Ke Zhang +7

    cs.CVeess.IVarXiv:2105.13381v32021
  7. InfographicVQA

    Minesh Mathew, Viraj Bagal, Rubèn Pérez Tito +3

    cs.CVcs.CLarXiv:2104.12756v22021
  8. Deep Generative Modelling: A Comparative Review of VAEs, GANs, Normalizing Flows, Energy-Based and Autoregressive Models

    Sam Bond-Taylor, Adam Leach, Yang Long +1

    cs.LGcs.CVstat.MLarXiv:2103.04922v42021
  9. Remote Sensing Image Change Detection with Transformers

    Hao Chen, Zipeng Qi, Zhenwei Shi

    cs.CVarXiv:2103.00208v32021
  10. GaitSet: Cross-view Gait Recognition through Utilizing Gait as a Deep Set

    Hanqing Chao, Kun Wang, Yiwei He +2

    cs.CVarXiv:2102.03247v12021
  11. Multi-stage Attention ResU-Net for Semantic Segmentation of Fine-Resolution Remote Sensing Images

    Rui Li, Shunyi Zheng, Chenxi Duan +2

    cs.CVarXiv:2011.14302v22020
  12. Dense Attention Fluid Network for Salient Object Detection in Optical Remote Sensing Images

    Qijian Zhang, Runmin Cong, Chongyi Li +5

    cs.CVarXiv:2011.13144v12020
  13. Channel-wise Knowledge Distillation for Dense Prediction

    Changyong Shu, Yifan Liu, Jianfei Gao +2

    cs.CVarXiv:2011.13256v42020
  14. PaDiM: a Patch Distribution Modeling Framework for Anomaly Detection and Localization

    Thomas Defard, Aleksandr Setkov, Angelique Loesch +1

    cs.CVarXiv:2011.08785v12020
  15. Rethinking the competition between detection and ReID in Multi-Object Tracking

    Chao Liang, Zhipeng Zhang, Xue Zhou +3

    cs.CVarXiv:2010.12138v32020
  16. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov +9

    cs.CVcs.AIcs.LGarXiv:2010.11929v22020
  17. Improving robustness against common corruptions by covariate shift adaptation

    Steffen Schneider, Evgenia Rusak, Luisa Eck +3

    cs.LGcs.CVstat.MLarXiv:2006.16971v22020
  18. RADIATE: A Radar Dataset for Automotive Perception in Bad Weather

    Marcel Sheeny, Emanuele De Pellegrin, Saptarshi Mukherjee +3

    cs.CVcs.ROarXiv:2010.09076v32020
  19. Denoising Diffusion Implicit Models

    Jiaming Song, Chenlin Meng, Stefano Ermon

    cs.LGcs.CVarXiv:2010.02502v42020
  20. Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity

    Youngwoo Yoon, Bok Cha, Joo-Haeng Lee +4

    cs.GRcs.CVcs.HCarXiv:2009.02119v12020
  21. Backdoor Attacks and Countermeasures on Deep Learning: A Comprehensive Review

    Yansong Gao, Bao Gia Doan, Zhi Zhang +5

    cs.CRcs.CVcs.LGarXiv:2007.10760v32020
  22. Towards Robust LiDAR-based Perception in Autonomous Driving: General Black-box Adversarial Sensor Attack and Countermeasures

    Jiachen Sun, Yulong Cao, Qi Alfred Chen +1

    cs.CRcs.CVcs.LGarXiv:2006.16974v12020
  23. TLIO: Tight Learned Inertial Odometry

    Wenxin Liu, David Caruso, Eddy Ilg +5

    cs.ROcs.CVcs.LGarXiv:2007.01867v32020
  24. Opportunities and Challenges in Explainable Artificial Intelligence (XAI): A Survey

    Arun Das, Paul Rad

    cs.CVcs.AIcs.LGarXiv:2006.11371v22020
  25. Prior-aware Neural Network for Partially-Supervised Multi-Organ Segmentation

    Yuyin Zhou, Zhe Li, Song Bai +5

    cs.CVarXiv:1904.06346v22019
  26. AffectDelta: Beyond Emotion Labels for Image Editing

    Xingzu Zhan, Lin Gu, Ruogu Fang

    cs.CVarXiv:2609.02616v12026
  27. Counterfactual VQA: A Cause-Effect Look at Language Bias

    Yulei Niu, Kaihua Tang, Hanwang Zhang +3

    cs.CVcs.CLarXiv:2006.04315v42020
  28. TEASER: Fast and Certifiable Point Cloud Registration

    Heng Yang, Jingnan Shi, Luca Carlone

    cs.ROcs.CVmath.OCarXiv:2001.07715v22020
  29. ProSR: Semantic-Prototype-Guided Discrete Modeling for Physically Consistent SAR Super-Resolution

    Byoungwoo Kim, Munchurl Kim

    cs.CVarXiv:2609.02377v12026
  30. The EPIC-KITCHENS Dataset: Collection, Challenges and Baselines

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella +8

    cs.CVarXiv:2005.00343v12020
  31. Dual-Sampling Attention Network for Diagnosis of COVID-19 from Community Acquired Pneumonia

    Xi Ouyang, Jiayu Huo, Liming Xia +15

    cs.CVeess.IVarXiv:2005.02690v22020
  32. TryOnDiffusion: A Tale of Two UNets

    Luyang Zhu, Dawei Yang, Tyler Zhu +5

    cs.CVcs.GRarXiv:2306.08276v12023
  33. Review of Artificial Intelligence Techniques in Imaging Data Acquisition, Segmentation and Diagnosis for COVID-19

    Feng Shi, Jun Wang, Jun Shi +6

    eess.IVcs.CVq-bio.QMarXiv:2004.02731v22020
  34. Coronavirus (COVID-19) Classification using CT Images by Machine Learning Methods

    Mucahid Barstugan, Umut Ozkaya, Saban Ozturk

    cs.CVcs.LGeess.IVarXiv:2003.09424v12020
  35. Hybrid Linear Modeling via Local Best-fit Flats

    Teng Zhang, Arthur Szlam, Yi Wang +1

    cs.CVstat.MLarXiv:1010.3460v22010
  36. Feature Extraction for Hyperspectral Imagery: The Evolution from Shallow to Deep (Overview and Toolbox)

    Behnood Rasti, Danfeng Hong, Renlong Hang +4

    cs.CVcs.LGeess.IVarXiv:2003.02822v42020
  37. Information Density Imbalance in Visual Object Detection

    Ziwei Zhao, Yanxi Lu, Yuwei Hu +8

    cs.CVarXiv:2609.02369v12026
  38. Algorithm Unrolling: Interpretable, Efficient Deep Learning for Signal and Image Processing

    Vishal Monga, Yuelong Li, Yonina C. Eldar

    eess.IVcs.CVcs.LGarXiv:1912.10557v32019
  39. Image Segmentation Using Deep Learning: A Survey

    Shervin Minaee, Yuri Boykov, Fatih Porikli +3

    cs.CVcs.LGarXiv:2001.05566v52020
  40. Geo-PIFu: Geometry and Pixel Aligned Implicit Functions for Single-view Human Reconstruction

    Tong He, John Collomosse, Hailin Jin +1

    cs.CVcs.GRcs.LGarXiv:2006.08072v22020
  41. Pathomic Fusion: An Integrated Framework for Fusing Histopathology and Genomic Features for Cancer Diagnosis and Prognosis

    Richard J. Chen, Ming Y. Lu, Jingwen Wang +4

    cs.CVq-bio.GNq-bio.TOarXiv:1912.08937v32019
  42. Grasping in the Wild:Learning 6DoF Closed-Loop Grasping from Low-Cost Demonstrations

    Shuran Song, Andy Zeng, Johnny Lee +1

    cs.CVcs.ROarXiv:1912.04344v22019
  43. UNet++: Redesigning Skip Connections to Exploit Multiscale Features in Image Segmentation

    Zongwei Zhou, Md Mahfuzur Rahman Siddiquee, Nima Tajbakhsh +1

    eess.IVcs.CVcs.LGarXiv:1912.05074v22019
  44. Connecting Vision and Language with Localized Narratives

    Jordi Pont-Tuset, Jasper Uijlings, Soravit Changpinyo +2

    cs.CVarXiv:1912.03098v42019
  45. A Deep Spatial Contextual Long-term Recurrent Convolutional Network for Saliency Detection

    Nian Liu, Junwei Han

    cs.CVarXiv:1610.01708v12016
  46. Image-based table recognition: data, model, and evaluation

    Xu Zhong, Elaheh ShafieiBavani, Antonio Jimeno Yepes

    cs.CVarXiv:1911.10683v52019
  47. Deep Learning for Hyperspectral Image Classification: An Overview

    Shutao Li, Weiwei Song, Leyuan Fang +3

    eess.IVcs.CVarXiv:1910.12861v12019
  48. Modified U-Net (mU-Net) with Incorporation of Object-Dependent High Level Features for Improved Liver and Liver-Tumor Segmentation in CT Images

    Hyunseok Seo, Charles Huang, Maxime Bassenne +2

    eess.IVcs.CVcs.LGarXiv:1911.00140v12019
  49. HAWQ: Hessian AWare Quantization of Neural Networks with Mixed-Precision

    Zhen Dong, Zhewei Yao, Amir Gholami +2

    cs.CVarXiv:1905.03696v12019
  50. Real-time Deep Dynamic Characters

    Marc Habermann, Lingjie Liu, Weipeng Xu +3

    cs.CVarXiv:2105.01794v22021
  51. A Survey of Deep Learning-based Object Detection

    Licheng Jiao, Fan Zhang, Fang Liu +4

    cs.CVarXiv:1907.09408v22019
  52. Blind Image Quality Assessment Using A Deep Bilinear Convolutional Neural Network

    Weixia Zhang, Kede Ma, Jia Yan +2

    eess.IVcs.CVcs.MMarXiv:1907.02665v12019
  53. Unlabeled Data Improves Adversarial Robustness

    Yair Carmon, Aditi Raghunathan, Ludwig Schmidt +2

    stat.MLcs.CVcs.LGarXiv:1905.13736v42019
  54. MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe

    Tianyu Yu, Zefan Wang, Chongyi Wang +31

    cs.LGcs.CVarXiv:2509.18154v12025
  55. Res2Net: A New Multi-scale Backbone Architecture

    Shang-Hua Gao, Ming-Ming Cheng, Kai Zhao +3

    cs.CVarXiv:1904.01169v32019
  56. HybridSN: Exploring 3D-2D CNN Feature Hierarchy for Hyperspectral Image Classification

    Swalpa Kumar Roy, Gopal Krishna, Shiv Ram Dubey +1

    cs.CVarXiv:1902.06701v32019
  57. TransMEF: A Transformer-Based Multi-Exposure Image Fusion Framework using Self-Supervised Multi-Task Learning

    Linhao Qu, Shaolei Liu, Manning Wang +1

    cs.CVarXiv:2112.01030v32021
  58. Activation Functions: Comparison of trends in Practice and Research for Deep Learning

    Chigozie Nwankpa, Winifred Ijomah, Anthony Gachagan +1

    cs.LGcs.CVarXiv:1811.03378v12018
  59. An Augmented Linear Mixing Model to Address Spectral Variability for Hyperspectral Unmixing

    Danfeng Hong, Naoto Yokoya, Jocelyn Chanussot +1

    cs.CVarXiv:1810.12000v12018
  60. A General Theory of Equivariant CNNs on Homogeneous Spaces

    Taco Cohen, Mario Geiger, Maurice Weiler

    cs.LGcs.AIcs.CGarXiv:1811.02017v22018