Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

7,381 to 7,440 of 18,849

  1. Graph-Based Object Classification for Neuromorphic Vision Sensing

    Yin Bi, Aaron Chadha, Alhabib Abbas +2

    cs.CVarXiv:1908.06648v12019
  2. Less Is More: Picking Informative Frames for Video Captioning

    Yangyu Chen, Shuhui Wang, Weigang Zhang +1

    cs.CVarXiv:1803.01457v12018
  3. Explainable Neural Computation via Stack Neural Module Networks

    Ronghang Hu, Jacob Andreas, Trevor Darrell +1

    cs.CVarXiv:1807.08556v32018
  4. Doppio: A Dataset for Contactless Weight Estimation of Falling Particles

    Simon Kiefhaber, Jan-Martin O. Steitz, Julia Grabinski +5

    cs.CVarXiv:2609.02528v12026
  5. Applications of Artificial Neural Networks in Microorganism Image Analysis: A Comprehensive Review from Conventional Multilayer Perceptron to Popular Convolutional Neural Network and Potential Visual Transformer

    Jinghua Zhang, Chen Li, Yimin Yin +2

    cs.CVcs.AIarXiv:2108.00358v32021
  6. Artificial Intelligence-Based Methods for Fusion of Electronic Health Records and Imaging Data

    Farida Mohsen, Hazrat Ali, Nady El Hajj +1

    cs.LGcs.AIcs.CVarXiv:2210.13462v12022
  7. Disease Burden over Skin Tone: Decomposing the Dermatology-AI Generalization Gap

    Nirajan Kunwor, Sanjaya Poudel, Quoc-Huy Trinh +2

    cs.CVcs.AIcs.LGarXiv:2609.02111v12026
  8. Mamba YOLO: A Simple Baseline for Object Detection with State Space Model

    Zeyu Wang, Chen Li, Huiying Xu +2

    cs.CVarXiv:2406.05835v22024
  9. SPA-GAN: Spatial Attention GAN for Image-to-Image Translation

    Hajar Emami, Majid Moradi Aliabadi, Ming Dong +1

    cs.CVarXiv:1908.06616v32019
  10. Deep Learning-based 3D Point Cloud Classification: A Systematic Survey and Outlook

    Huang Zhang, Changshuo Wang, Shengwei Tian +4

    cs.CVarXiv:2311.02608v12023
  11. Local contrastive loss with pseudo-label based self-training for semi-supervised medical image segmentation

    Krishna Chaitanya, Ertunc Erdil, Neerav Karani +1

    cs.CVcs.AIcs.LGarXiv:2112.09645v12021
  12. Attention-Aware Face Hallucination via Deep Reinforcement Learning

    Qingxing Cao, Liang Lin, Yukai Shi +2

    cs.CVarXiv:1708.03132v12017
  13. SFFNet: A Wavelet-Based Spatial and Frequency Domain Fusion Network for Remote Sensing Segmentation

    Yunsong Yang, Genji Yuan, Jinjiang Li

    cs.CVarXiv:2405.01992v12024
  14. LeafGAN: An Effective Data Augmentation Method for Practical Plant Disease Diagnosis

    Quan Huu Cap, Hiroyuki Uga, Satoshi Kagiwada +1

    cs.CVarXiv:2002.10100v22020
  15. An Interactively Reinforced Paradigm for Joint Infrared-Visible Image Fusion and Saliency Object Detection

    Di Wang, Jinyuan Liu, Risheng Liu +1

    cs.CVarXiv:2305.09999v12023
  16. Co-Learning Feature Fusion Maps from PET-CT Images of Lung Cancer

    Ashnil Kumar, Michael Fulham, Dagan Feng +1

    cs.CVarXiv:1810.02492v22018
  17. Detecting Object Hallucinations in Large Vision-Language Models via Cross-Modal Attention Drifts and Mask-Based Verification

    Xuanbing Wen, Boxu Chen, Le Yang +4

    cs.CVarXiv:2609.02028v12026
  18. Understanding the Failure Modes of Out-of-Distribution Generalization

    Vaishnavh Nagarajan, Anders Andreassen, Behnam Neyshabur

    cs.LGcs.CVstat.MLarXiv:2010.15775v32020
  19. NAM: Normalization-based Attention Module

    Yichao Liu, Zongru Shao, Yueyang Teng +1

    cs.CVeess.IVarXiv:2111.12419v12021
  20. Perspective-Guided Convolution Networks for Crowd Counting

    Zhaoyi Yan, Yuchen Yuan, Wangmeng Zuo +4

    cs.CVarXiv:1909.06966v12019
  21. Unified Vision and Language Prompt Learning

    Yuhang Zang, Wei Li, Kaiyang Zhou +2

    cs.CVcs.AIarXiv:2210.07225v12022
  22. Diffused Heads: Diffusion Models Beat GANs on Talking-Face Generation

    Michał Stypułkowski, Konstantinos Vougioukas, Sen He +3

    cs.CVarXiv:2301.03396v22023
  23. Semantic Segmentation with Boundary Neural Fields

    Gedas Bertasius, Jianbo Shi, Lorenzo Torresani

    cs.CVarXiv:1511.02674v22015
  24. 3D Segmentation with Exponential Logarithmic Loss for Highly Unbalanced Object Sizes

    Ken C. L. Wong, Mehdi Moradi, Hui Tang +1

    cs.CVarXiv:1809.00076v22018
  25. Genesis: A Generative Engine for Hierarchical Satellite Image Synthesis

    Subash Khanal, Yangzhi Cui, Daniel Cher +4

    cs.CVarXiv:2609.02683v12026
  26. Interpretable Survival Prediction for Colorectal Cancer using Deep Learning

    Ellery Wulczyn, David F. Steiner, Melissa Moran +20

    eess.IVcs.CVarXiv:2011.08965v12020
  27. Weather Influence and Classification with Automotive Lidar Sensors

    Robin Heinzler, Philipp Schindler, Jürgen Seekircher +2

    cs.CVarXiv:1906.07675v12019
  28. Meta-Transformer: A Unified Framework for Multimodal Learning

    Yiyuan Zhang, Kaixiong Gong, Kaipeng Zhang +4

    cs.CVcs.AIcs.CLarXiv:2307.10802v12023
  29. Learning Convolutional Transforms for Lossy Point Cloud Geometry Compression

    Maurice Quach, Giuseppe Valenzise, Frederic Dufaux

    cs.CVcs.LGeess.IVarXiv:1903.08548v22019
  30. Fully Convolutional Change Detection Framework with Generative Adversarial Network for Unsupervised, Weakly Supervised and Regional Supervised Change Detection

    Chen Wu, Bo Du, Liangpei Zhang

    cs.CVcs.AIeess.IVarXiv:2201.06030v12022
  31. Is it Time to Replace CNNs with Transformers for Medical Images?

    Christos Matsoukas, Johan Fredin Haslum, Magnus Söderberg +1

    cs.CVcs.LGarXiv:2108.09038v12021
  32. Deforming Autoencoders: Unsupervised Disentangling of Shape and Appearance

    Zhixin Shu, Mihir Sahasrabudhe, Alp Guler +3

    cs.CVarXiv:1806.06503v12018
  33. LeakageBench: Document-Level Leakage Risk for Redacting Personally Identifiable Information in Document Images

    Vishnu Prasad Vijaya Kumar, Santhosh Venkatesh, Ivan P. Yamshchikov

    cs.CVcs.CLarXiv:2609.02207v12026
  34. Image Generation From Small Datasets via Batch Statistics Adaptation

    Atsuhiro Noguchi, Tatsuya Harada

    cs.CVarXiv:1904.01774v42019
  35. Semi-Supervised Brain Lesion Segmentation with an Adapted Mean Teacher Model

    Wenhui Cui, Yanlin Liu, Yuxing Li +6

    cs.CVarXiv:1903.01248v12019
  36. Kymatio: Scattering Transforms in Python

    Mathieu Andreux, Tomás Angles, Georgios Exarchakis +15

    cs.LGcs.CVcs.SDarXiv:1812.11214v32018
  37. Class Re-Activation Maps for Weakly-Supervised Semantic Segmentation

    Zhaozheng Chen, Tan Wang, Xiongwei Wu +3

    cs.CVarXiv:2203.00962v12022
  38. TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

    Yuliang Liu, Biao Yang, Qiang Liu +4

    cs.CVcs.AIarXiv:2403.04473v22024
  39. Learned Point Cloud Geometry Compression

    Jianqiang Wang, Hao Zhu, Zhan Ma +3

    cs.CVeess.IVarXiv:1909.12037v12019
  40. Switchable Whitening for Deep Representation Learning

    Xingang Pan, Xiaohang Zhan, Jianping Shi +2

    cs.CVarXiv:1904.09739v42019
  41. Intracranial Hemorrhage Segmentation Using Deep Convolutional Model

    Murtadha D. Hssayeni, M. S., Muayad S. Croock +9

    eess.IVcs.CVarXiv:1910.08643v22019
  42. Back to Event Basics: Self-Supervised Learning of Image Reconstruction for Event Cameras via Photometric Constancy

    F. Paredes-Vallés, G. C. H. E. de Croon

    cs.CVarXiv:2009.08283v22020
  43. Learning 3D Representations from 2D Pre-trained Models via Image-to-Point Masked Autoencoders

    Renrui Zhang, Liuhui Wang, Yu Qiao +2

    cs.CVcs.AIarXiv:2212.06785v12022
  44. Enriched Long-term Recurrent Convolutional Network for Facial Micro-Expression Recognition

    Huai-Qian Khor, John See, Raphael C. W. Phan +1

    cs.CVarXiv:1805.08417v12018
  45. Satellite Image Time Series Classification with Pixel-Set Encoders and Temporal Self-Attention

    Vivien Sainte Fare Garnot, Loic Landrieu, Sebastien Giordano +1

    cs.CVarXiv:1911.07757v12019
  46. ParSeNet: A Parametric Surface Fitting Network for 3D Point Clouds

    Gopal Sharma, Difan Liu, Subhransu Maji +3

    cs.CVcs.LGarXiv:2003.12181v52020
  47. Bridging the Gap between Object and Image-level Representations for Open-Vocabulary Detection

    Hanoona Rasheed, Muhammad Maaz, Muhammad Uzair Khattak +2

    cs.CVcs.AIarXiv:2207.03482v32022
  48. Attracting and Dispersing: A Simple Approach for Source-free Domain Adaptation

    Shiqi Yang, Yaxing Wang, Kai Wang +2

    cs.CVcs.LGarXiv:2205.04183v32022
  49. ChromaGAN: Adversarial Picture Colorization with Semantic Class Distribution

    Patricia Vitoria, Lara Raad, Coloma Ballester

    cs.CVarXiv:1907.09837v22019
  50. Transferability and Hardness of Supervised Classification Tasks

    Anh T. Tran, Cuong V. Nguyen, Tal Hassner

    cs.LGcs.CVstat.MLarXiv:1908.08142v12019
  51. OctAttention: Octree-Based Large-Scale Contexts Model for Point Cloud Compression

    Chunyang Fu, Ge Li, Rui Song +2

    cs.CVarXiv:2202.06028v22022
  52. PubTables-1M: Towards comprehensive table extraction from unstructured documents

    Brandon Smock, Rohith Pesala, Robin Abraham

    cs.LGcs.CVarXiv:2110.00061v32021
  53. MAOL: Morphology-Aware Ordinal Learning for Fine-Grained Industrial Defect Severity Grading

    Zhaoyang Wang, Haiyong Chen, Binyi Su +4

    cs.CVarXiv:2609.02266v12026
  54. The importance of stain normalization in colorectal tissue classification with convolutional networks

    Francesco Ciompi, Oscar Geessink, Babak Ehteshami Bejnordi +6

    cs.CVcs.LGarXiv:1702.05931v22017
  55. FuDU: A Fuzzy Dual-dimensional Uncertainty Framework for Streaming Active Learning in Industrial Defect Detection

    Zhaoyang Wang, Haiyong Chen, Binyi Su +1

    cs.CVarXiv:2609.02212v12026
  56. Neither hype nor gloom do DNNs justice

    Felix A. Wichmann, Simon Kornblith, Robert Geirhos

    cs.LGcs.CVq-bio.NCarXiv:2312.05355v12023
  57. Attention Attention Everywhere: Monocular Depth Prediction with Skip Attention

    Ashutosh Agarwal, Chetan Arora

    cs.CVarXiv:2210.09071v12022
  58. Unsupervised Pre-training for Person Re-identification

    Dengpan Fu, Dongdong Chen, Jianmin Bao +5

    cs.CVarXiv:2012.03753v22020
  59. Efficient All-in-One Weather Restoration using Spectral Harmonization

    Paula Garrido-Mellado, Daniel Feijoo, Yuning Cui +2

    cs.CVarXiv:2609.02839v12026
  60. Stereo 4D Radar for 3D Object Detection: Integrating Geometric Alignment and Absolute Velocity Estimation

    Seung-Hyun Song, Dong-Hee Paek, Woong-Chan Byun +1

    cs.CVarXiv:2609.02560v12026