Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1,861 to 1,920 of 18,866

  1. Generalized Radiograph Representation Learning via Cross-supervision between Images and Free-text Radiology Reports

    Hong-Yu Zhou, Xiaoyu Chen, Yinghao Zhang +3

    eess.IVcs.CVcs.LGarXiv:2111.03452v22021
  2. Cross-Attention in Coupled Unmixing Nets for Unsupervised Hyperspectral Super-Resolution

    Jing Yao, Danfeng Hong, Jocelyn Chanussot +3

    eess.IVcs.CVarXiv:2007.05230v32020
  3. AdvFaces: Adversarial Face Synthesis

    Debayan Deb, Jianbang Zhang, Anil K. Jain

    cs.CVarXiv:1908.05008v12019
  4. GIAOTracker: A comprehensive framework for MCMOT with global information and optimizing strategies in VisDrone 2021

    Yunhao Du, Junfeng Wan, Yanyun Zhao +3

    cs.CVarXiv:2202.11983v12022
  5. A New Learning Paradigm for Foundation Model-based Remote Sensing Change Detection

    Kaiyu Li, Xiangyong Cao, Deyu Meng

    cs.CVarXiv:2312.01163v22023
  6. How hard can it be? Estimating the difficulty of visual search in an image

    Radu Tudor Ionescu, Bogdan Alexe, Marius Leordeanu +3

    cs.CVarXiv:1705.08280v12017
  7. ColorNet: Investigating the importance of color spaces for image classification

    Shreyank N Gowda, Chun Yuan

    cs.CVarXiv:1902.00267v12019
  8. GTC: Guided Training of CTC Towards Efficient and Accurate Scene Text Recognition

    Wenyang Hu, Xiaocong Cai, Jun Hou +2

    cs.CVcs.LGeess.IVarXiv:2002.01276v12020
  9. Generating High Fidelity Images with Subscale Pixel Networks and Multidimensional Upscaling

    Jacob Menick, Nal Kalchbrenner

    cs.CVcs.GRcs.LGarXiv:1812.01608v12018
  10. LinkNet: Relational Embedding for Scene Graph

    Sanghyun Woo, Dahun Kim, Donghyeon Cho +1

    cs.CVarXiv:1811.06410v12018
  11. DriveGAN: Towards a Controllable High-Quality Neural Simulation

    Seung Wook Kim, Jonah Philion, Antonio Torralba +1

    cs.CVcs.ROarXiv:2104.15060v12021
  12. 2D3D-MatchNet: Learning to Match Keypoints Across 2D Image and 3D Point Cloud

    Mengdan Feng, Sixing Hu, Marcelo Ang +1

    cs.CVarXiv:1904.09742v12019
  13. Classification of breast cancer histology images using transfer learning

    Sulaiman Vesal, Nishant Ravikumar, AmirAbbas Davari +2

    cs.CVarXiv:1802.09424v12018
  14. SLSDeep: Skin Lesion Segmentation Based on Dilated Residual and Pyramid Pooling Networks

    Md. Mostafa Kamal Sarker, Hatem A. Rashwan, Farhan Akram +8

    cs.CVarXiv:1805.10241v22018
  15. Skin Lesion Classification Using CNNs with Patch-Based Attention and Diagnosis-Guided Loss Weighting

    Nils Gessert, Thilo Sentker, Frederic Madesta +5

    cs.CVarXiv:1905.02793v22019
  16. Semantic Graph Based Place Recognition for 3D Point Clouds

    Xin Kong, Xuemeng Yang, Guangyao Zhai +6

    cs.CVcs.ROarXiv:2008.11459v12020
  17. Regularizing Activation Distribution for Training Binarized Deep Networks

    Ruizhou Ding, Ting-Wu Chin, Zeye Liu +1

    cs.CVarXiv:1904.02823v12019
  18. Crossing Nets: Combining GANs and VAEs with a Shared Latent Space for Hand Pose Estimation

    Chengde Wan, Thomas Probst, Luc Van Gool +1

    cs.CVarXiv:1702.03431v22017
  19. Attention-based Extraction of Structured Information from Street View Imagery

    Zbigniew Wojna, Alex Gorban, Dar-Shyang Lee +4

    cs.CVarXiv:1704.03549v42017
  20. Untargeted Backdoor Watermark: Towards Harmless and Stealthy Dataset Copyright Protection

    Yiming Li, Yang Bai, Yong Jiang +3

    cs.CRcs.AIcs.CVarXiv:2210.00875v32022
  21. Diagnostic Classification Of Lung Nodules Using 3D Neural Networks

    Raunak Dey, Zhongjie Lu, Yi Hong

    cs.CVcs.LGstat.MLarXiv:1803.07192v12018
  22. MDMMT: Multidomain Multimodal Transformer for Video Retrieval

    Maksim Dzabraev, Maksim Kalashnikov, Stepan Komkov +1

    cs.CVarXiv:2103.10699v12021
  23. ReduNet: A White-box Deep Network from the Principle of Maximizing Rate Reduction

    Kwan Ho Ryan Chan, Yaodong Yu, Chong You +3

    cs.LGcs.CVcs.ITarXiv:2105.10446v32021
  24. Neuron Shapley: Discovering the Responsible Neurons

    Amirata Ghorbani, James Zou

    stat.MLcs.CVcs.LGarXiv:2002.09815v32020
  25. Instruction-driven history-aware policies for robotic manipulations

    Pierre-Louis Guhur, Shizhe Chen, Ricardo Garcia +3

    cs.ROcs.AIcs.CLarXiv:2209.04899v32022
  26. Semigroup-JEPA: Latent Dynamics Consistency for Zero-Shot Physics Generalization

    Andy Zeyi Liu, Haoran Sun, Lucas Baker +2

    cs.LGcs.AIcs.CVarXiv:2609.10464v12026
  27. CR-GAN: Learning Complete Representations for Multi-view Generation

    Yu Tian, Xi Peng, Long Zhao +2

    cs.CVarXiv:1806.11191v12018
  28. Learning Fast and Robust Target Models for Video Object Segmentation

    Andreas Robinson, Felix Järemo Lawin, Martin Danelljan +2

    cs.CVarXiv:2003.00908v22020
  29. When Person Re-identification Meets Changing Clothes

    Fangbin Wan, Yang Wu, Xuelin Qian +2

    cs.CVarXiv:2003.04070v32020
  30. Coronavirus (COVID-19) Classification using Deep Features Fusion and Ranking Technique

    Umut Ozkaya, Saban Ozturk, Mucahid Barstugan

    eess.IVcs.CVcs.LGarXiv:2004.03698v12020
  31. Learning to Discover Multi-Class Attentional Regions for Multi-Label Image Recognition

    Bin-Bin Gao, Hong-Yu Zhou

    cs.CVarXiv:2007.01755v32020
  32. Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining

    Dongyang Liu, Shitian Zhao, Le Zhuo +7

    cs.CVarXiv:2408.02657v32024
  33. On Learning the Geodesic Path for Incremental Learning

    Christian Simon, Piotr Koniusz, Mehrtash Harandi

    cs.LGcs.CVarXiv:2104.08572v12021
  34. PartSLIP: Low-Shot Part Segmentation for 3D Point Clouds via Pretrained Image-Language Models

    Minghua Liu, Yinhao Zhu, Hong Cai +4

    cs.CVcs.ROarXiv:2212.01558v22022
  35. Editing Implicit Assumptions in Text-to-Image Diffusion Models

    Hadas Orgad, Bahjat Kawar, Yonatan Belinkov

    cs.CVarXiv:2303.08084v22023
  36. Multi-Modal Temporal Attention Models for Crop Mapping from Satellite Time Series

    Vivien Sainte Fare Garnot, Loic Landrieu, Nesrine Chehata

    cs.CVeess.IVarXiv:2112.07558v12021
  37. Decorate the Newcomers: Visual Domain Prompt for Continual Test Time Adaptation

    Yulu Gan, Yan Bai, Yihang Lou +4

    cs.CVarXiv:2212.04145v22022
  38. Visual Room Rearrangement

    Luca Weihs, Matt Deitke, Aniruddha Kembhavi +1

    cs.CVcs.ROarXiv:2103.16544v12021
  39. MSDN: Mutually Semantic Distillation Network for Zero-Shot Learning

    Shiming Chen, Ziming Hong, Guo-Sen Xie +5

    cs.CVarXiv:2203.03137v22022
  40. Two-Stream Consensus Network for Weakly-Supervised Temporal Action Localization

    Yuanhao Zhai, Le Wang, Wei Tang +3

    cs.CVarXiv:2010.11594v12020
  41. Deep Image Translation with an Affinity-Based Change Prior for Unsupervised Multimodal Change Detection

    Luigi Tommaso Luppino, Michael Kampffmeyer, Filippo Maria Bianchi +4

    cs.LGcs.CVeess.IVarXiv:2001.04271v22020
  42. PICK: Processing Key Information Extraction from Documents using Improved Graph Learning-Convolutional Networks

    Wenwen Yu, Ning Lu, Xianbiao Qi +2

    cs.CVarXiv:2004.07464v32020
  43. Efficient Multimodal Learning from Data-centric Perspective

    Muyang He, Yexin Liu, Boya Wu +4

    cs.CVarXiv:2402.11530v32024
  44. Merlin: A Computed Tomography Vision-Language Foundation Model and Dataset

    Louis Blankemeier, Ashwin Kumar, Joseph Paul Cohen +37

    cs.CVcs.AIarXiv:2406.06512v22024
  45. SmallCap: Lightweight Image Captioning Prompted with Retrieval Augmentation

    Rita Ramos, Bruno Martins, Desmond Elliott +1

    cs.CVcs.CLarXiv:2209.15323v22022
  46. Bidirectional Projection Network for Cross Dimension Scene Understanding

    Wenbo Hu, Hengshuang Zhao, Li Jiang +2

    cs.CVarXiv:2103.14326v12021
  47. Feature Fusion Vision Transformer for Fine-Grained Visual Categorization

    Jun Wang, Xiaohan Yu, Yongsheng Gao

    cs.CVarXiv:2107.02341v32021
  48. NAS-Bench-1Shot1: Benchmarking and Dissecting One-shot Neural Architecture Search

    Arber Zela, Julien Siems, Frank Hutter

    cs.LGcs.CVcs.NEarXiv:2001.10422v22020
  49. Towards Lossless Dataset Distillation via Difficulty-Aligned Trajectory Matching

    Ziyao Guo, Kai Wang, George Cazenavette +3

    cs.CVarXiv:2310.05773v22023
  50. PanoFormer: Panorama Transformer for Indoor 360 Depth Estimation

    Zhijie Shen, Chunyu Lin, Kang Liao +3

    cs.CVarXiv:2203.09283v22022
  51. Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures

    Yuchen Duan, Weiyun Wang, Zhe Chen +7

    cs.CVarXiv:2403.02308v32024
  52. PACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving

    Lin Huang, Yujuan Tan, Weisheng Li +3

    cs.CVcs.AIcs.ROarXiv:2609.10372v12026
  53. Generative Low-bitwidth Data Free Quantization

    Shoukai Xu, Haokun Li, Bohan Zhuang +4

    cs.CVarXiv:2003.03603v32020
  54. Few-shot 3D Point Cloud Semantic Segmentation

    Na Zhao, Tat-Seng Chua, Gim Hee Lee

    cs.CVarXiv:2006.12052v22020
  55. One Loop, Two Gains: Can Active Learning win the Lottery for Free?

    Benedikt Tscheschner, Eduardo Veas, Marc Masana

    cs.LGcs.AIcs.CVarXiv:2609.10311v12026
  56. Experimental comparison of single-pixel imaging algorithms

    Liheng Bian, Jinli Suo, Qionghai Dai +1

    cs.CVphysics.opticsarXiv:1707.03164v22017
  57. Reliability of PET/CT shape and heterogeneity features in functional and morphological components of Non-Small Cell Lung Cancer tumors: a repeatability analysis in a prospective multi-center cohort

    Marie-Charlotte Desseroit, Florent Tixier, Wolfgang Weber +4

    cs.CVphysics.med-pharXiv:1610.01390v12016
  58. DXSLAM: A Robust and Efficient Visual SLAM System with Deep Features

    Dongjiang Li, Xuesong Shi, Qiwei Long +5

    cs.CVcs.ROarXiv:2008.05416v12020
  59. Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

    Killian Steunou, Yannis Tevissen, Mounîm A. El Yacoubi

    cs.CVcs.CLcs.MMarXiv:2609.10355v12026
  60. Learning to See the Invisible: End-to-End Trainable Amodal Instance Segmentation

    Patrick Follmann, Rebecca König, Philipp Härtinger +1

    cs.CVarXiv:1804.08864v12018