Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
6,661 to 6,720 of 18,866
PandaSet: Advanced Sensor Suite Dataset for Autonomous Driving
Pengchuan Xiao, Zhenlei Shao, Steven Hao +9
cs.CVcs.ROarXiv:2112.12610v12021Temporal Context Mining for Learned Video Compression
Xihua Sheng, Jiahao Li, Bin Li +3
cs.CVcs.LGeess.IVarXiv:2111.13850v22021From Show to Tell: A Survey on Deep Learning-based Image Captioning
Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi +3
cs.CVcs.CLarXiv:2107.06912v32021Deep Subdomain Adaptation Network for Image Classification
Yongchun Zhu, Fuzhen Zhuang, Jindong Wang +5
cs.CVcs.AIarXiv:2106.09388v12021Structured Denoising Diffusion Models in Discrete State-Spaces
Jacob Austin, Daniel D. Johnson, Jonathan Ho +2
cs.LGcs.AIcs.CLarXiv:2107.03006v32021Recent advances and clinical applications of deep learning in medical image analysis
Xuxin Chen, Ximin Wang, Ke Zhang +7
cs.CVeess.IVarXiv:2105.13381v32021InfographicVQA
Minesh Mathew, Viraj Bagal, Rubèn Pérez Tito +3
cs.CVcs.CLarXiv:2104.12756v22021Deep Generative Modelling: A Comparative Review of VAEs, GANs, Normalizing Flows, Energy-Based and Autoregressive Models
Sam Bond-Taylor, Adam Leach, Yang Long +1
cs.LGcs.CVstat.MLarXiv:2103.04922v42021Remote Sensing Image Change Detection with Transformers
Hao Chen, Zipeng Qi, Zhenwei Shi
cs.CVarXiv:2103.00208v32021GaitSet: Cross-view Gait Recognition through Utilizing Gait as a Deep Set
Hanqing Chao, Kun Wang, Yiwei He +2
cs.CVarXiv:2102.03247v12021Multi-stage Attention ResU-Net for Semantic Segmentation of Fine-Resolution Remote Sensing Images
Rui Li, Shunyi Zheng, Chenxi Duan +2
cs.CVarXiv:2011.14302v22020Dense Attention Fluid Network for Salient Object Detection in Optical Remote Sensing Images
Qijian Zhang, Runmin Cong, Chongyi Li +5
cs.CVarXiv:2011.13144v12020Channel-wise Knowledge Distillation for Dense Prediction
Changyong Shu, Yifan Liu, Jianfei Gao +2
cs.CVarXiv:2011.13256v42020PaDiM: a Patch Distribution Modeling Framework for Anomaly Detection and Localization
Thomas Defard, Aleksandr Setkov, Angelique Loesch +1
cs.CVarXiv:2011.08785v12020Rethinking the competition between detection and ReID in Multi-Object Tracking
Chao Liang, Zhipeng Zhang, Xue Zhou +3
cs.CVarXiv:2010.12138v32020An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov +9
cs.CVcs.AIcs.LGarXiv:2010.11929v22020Improving robustness against common corruptions by covariate shift adaptation
Steffen Schneider, Evgenia Rusak, Luisa Eck +3
cs.LGcs.CVstat.MLarXiv:2006.16971v22020RADIATE: A Radar Dataset for Automotive Perception in Bad Weather
Marcel Sheeny, Emanuele De Pellegrin, Saptarshi Mukherjee +3
cs.CVcs.ROarXiv:2010.09076v32020Denoising Diffusion Implicit Models
Jiaming Song, Chenlin Meng, Stefano Ermon
cs.LGcs.CVarXiv:2010.02502v42020Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity
Youngwoo Yoon, Bok Cha, Joo-Haeng Lee +4
cs.GRcs.CVcs.HCarXiv:2009.02119v12020Backdoor Attacks and Countermeasures on Deep Learning: A Comprehensive Review
Yansong Gao, Bao Gia Doan, Zhi Zhang +5
cs.CRcs.CVcs.LGarXiv:2007.10760v32020Towards Robust LiDAR-based Perception in Autonomous Driving: General Black-box Adversarial Sensor Attack and Countermeasures
Jiachen Sun, Yulong Cao, Qi Alfred Chen +1
cs.CRcs.CVcs.LGarXiv:2006.16974v12020TLIO: Tight Learned Inertial Odometry
Wenxin Liu, David Caruso, Eddy Ilg +5
cs.ROcs.CVcs.LGarXiv:2007.01867v32020Opportunities and Challenges in Explainable Artificial Intelligence (XAI): A Survey
Arun Das, Paul Rad
cs.CVcs.AIcs.LGarXiv:2006.11371v22020Prior-aware Neural Network for Partially-Supervised Multi-Organ Segmentation
Yuyin Zhou, Zhe Li, Song Bai +5
cs.CVarXiv:1904.06346v22019AffectDelta: Beyond Emotion Labels for Image Editing
Xingzu Zhan, Lin Gu, Ruogu Fang
cs.CVarXiv:2609.02616v12026Counterfactual VQA: A Cause-Effect Look at Language Bias
Yulei Niu, Kaihua Tang, Hanwang Zhang +3
cs.CVcs.CLarXiv:2006.04315v42020TEASER: Fast and Certifiable Point Cloud Registration
Heng Yang, Jingnan Shi, Luca Carlone
cs.ROcs.CVmath.OCarXiv:2001.07715v22020ProSR: Semantic-Prototype-Guided Discrete Modeling for Physically Consistent SAR Super-Resolution
Byoungwoo Kim, Munchurl Kim
cs.CVarXiv:2609.02377v12026The EPIC-KITCHENS Dataset: Collection, Challenges and Baselines
Dima Damen, Hazel Doughty, Giovanni Maria Farinella +8
cs.CVarXiv:2005.00343v12020Dual-Sampling Attention Network for Diagnosis of COVID-19 from Community Acquired Pneumonia
Xi Ouyang, Jiayu Huo, Liming Xia +15
cs.CVeess.IVarXiv:2005.02690v22020TryOnDiffusion: A Tale of Two UNets
Luyang Zhu, Dawei Yang, Tyler Zhu +5
cs.CVcs.GRarXiv:2306.08276v12023Review of Artificial Intelligence Techniques in Imaging Data Acquisition, Segmentation and Diagnosis for COVID-19
Feng Shi, Jun Wang, Jun Shi +6
eess.IVcs.CVq-bio.QMarXiv:2004.02731v22020Coronavirus (COVID-19) Classification using CT Images by Machine Learning Methods
Mucahid Barstugan, Umut Ozkaya, Saban Ozturk
cs.CVcs.LGeess.IVarXiv:2003.09424v12020Hybrid Linear Modeling via Local Best-fit Flats
Teng Zhang, Arthur Szlam, Yi Wang +1
cs.CVstat.MLarXiv:1010.3460v22010Feature Extraction for Hyperspectral Imagery: The Evolution from Shallow to Deep (Overview and Toolbox)
Behnood Rasti, Danfeng Hong, Renlong Hang +4
cs.CVcs.LGeess.IVarXiv:2003.02822v42020Information Density Imbalance in Visual Object Detection
Ziwei Zhao, Yanxi Lu, Yuwei Hu +8
cs.CVarXiv:2609.02369v12026Algorithm Unrolling: Interpretable, Efficient Deep Learning for Signal and Image Processing
Vishal Monga, Yuelong Li, Yonina C. Eldar
eess.IVcs.CVcs.LGarXiv:1912.10557v32019Image Segmentation Using Deep Learning: A Survey
Shervin Minaee, Yuri Boykov, Fatih Porikli +3
cs.CVcs.LGarXiv:2001.05566v52020Geo-PIFu: Geometry and Pixel Aligned Implicit Functions for Single-view Human Reconstruction
Tong He, John Collomosse, Hailin Jin +1
cs.CVcs.GRcs.LGarXiv:2006.08072v22020Pathomic Fusion: An Integrated Framework for Fusing Histopathology and Genomic Features for Cancer Diagnosis and Prognosis
Richard J. Chen, Ming Y. Lu, Jingwen Wang +4
cs.CVq-bio.GNq-bio.TOarXiv:1912.08937v32019Grasping in the Wild:Learning 6DoF Closed-Loop Grasping from Low-Cost Demonstrations
Shuran Song, Andy Zeng, Johnny Lee +1
cs.CVcs.ROarXiv:1912.04344v22019UNet++: Redesigning Skip Connections to Exploit Multiscale Features in Image Segmentation
Zongwei Zhou, Md Mahfuzur Rahman Siddiquee, Nima Tajbakhsh +1
eess.IVcs.CVcs.LGarXiv:1912.05074v22019Connecting Vision and Language with Localized Narratives
Jordi Pont-Tuset, Jasper Uijlings, Soravit Changpinyo +2
cs.CVarXiv:1912.03098v42019A Deep Spatial Contextual Long-term Recurrent Convolutional Network for Saliency Detection
Nian Liu, Junwei Han
cs.CVarXiv:1610.01708v12016Image-based table recognition: data, model, and evaluation
Xu Zhong, Elaheh ShafieiBavani, Antonio Jimeno Yepes
cs.CVarXiv:1911.10683v52019Deep Learning for Hyperspectral Image Classification: An Overview
Shutao Li, Weiwei Song, Leyuan Fang +3
eess.IVcs.CVarXiv:1910.12861v12019Modified U-Net (mU-Net) with Incorporation of Object-Dependent High Level Features for Improved Liver and Liver-Tumor Segmentation in CT Images
Hyunseok Seo, Charles Huang, Maxime Bassenne +2
eess.IVcs.CVcs.LGarXiv:1911.00140v12019HAWQ: Hessian AWare Quantization of Neural Networks with Mixed-Precision
Zhen Dong, Zhewei Yao, Amir Gholami +2
cs.CVarXiv:1905.03696v12019Real-time Deep Dynamic Characters
Marc Habermann, Lingjie Liu, Weipeng Xu +3
cs.CVarXiv:2105.01794v22021A Survey of Deep Learning-based Object Detection
Licheng Jiao, Fan Zhang, Fang Liu +4
cs.CVarXiv:1907.09408v22019Blind Image Quality Assessment Using A Deep Bilinear Convolutional Neural Network
Weixia Zhang, Kede Ma, Jia Yan +2
eess.IVcs.CVcs.MMarXiv:1907.02665v12019Unlabeled Data Improves Adversarial Robustness
Yair Carmon, Aditi Raghunathan, Ludwig Schmidt +2
stat.MLcs.CVcs.LGarXiv:1905.13736v42019MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
Tianyu Yu, Zefan Wang, Chongyi Wang +31
cs.LGcs.CVarXiv:2509.18154v12025Res2Net: A New Multi-scale Backbone Architecture
Shang-Hua Gao, Ming-Ming Cheng, Kai Zhao +3
cs.CVarXiv:1904.01169v32019HybridSN: Exploring 3D-2D CNN Feature Hierarchy for Hyperspectral Image Classification
Swalpa Kumar Roy, Gopal Krishna, Shiv Ram Dubey +1
cs.CVarXiv:1902.06701v32019TransMEF: A Transformer-Based Multi-Exposure Image Fusion Framework using Self-Supervised Multi-Task Learning
Linhao Qu, Shaolei Liu, Manning Wang +1
cs.CVarXiv:2112.01030v32021Activation Functions: Comparison of trends in Practice and Research for Deep Learning
Chigozie Nwankpa, Winifred Ijomah, Anthony Gachagan +1
cs.LGcs.CVarXiv:1811.03378v12018An Augmented Linear Mixing Model to Address Spectral Variability for Hyperspectral Unmixing
Danfeng Hong, Naoto Yokoya, Jocelyn Chanussot +1
cs.CVarXiv:1810.12000v12018A General Theory of Equivariant CNNs on Homogeneous Spaces
Taco Cohen, Mario Geiger, Maurice Weiler
cs.LGcs.AIcs.CGarXiv:1811.02017v22018