Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

8,401 to 8,460 of 18,867

  1. Omni-frequency Channel-selection Representations for Unsupervised Anomaly Detection

    Yufei Liang, Jiangning Zhang, Shiwei Zhao +3

    cs.CVarXiv:2203.00259v22022
  2. Handheld Multi-Frame Super-Resolution

    Bartlomiej Wronski, Ignacio Garcia-Dorado, Manfred Ernst +5

    cs.CVeess.IVarXiv:1905.03277v22019
  3. FCCDN: Feature Constraint Network for VHR Image Change Detection

    Pan Chen, Danfeng Hong, Zhengchao Chen +3

    cs.CVarXiv:2105.10860v22021
  4. 3D-MRL: Nested Multimodal 3D Representations via Matryoshka Representation Learning

    Márcus Lobo, Vitor Matias, Jeová Farias +1

    cs.CVarXiv:2608.29285v12026
  5. Deep Learning Autoencoder Approach for Handwritten Arabic Digits Recognition

    Mohamed Loey, Ahmed El-Sawy, Hazem EL-Bakry

    cs.CVcs.NEarXiv:1706.06720v12017
  6. Semi-Supervised Medical Image Segmentation via Learning Consistency under Transformations

    Gerda Bortsova, Florian Dubost, Laurens Hogeweg +2

    cs.CVcs.LGarXiv:1911.01218v12019
  7. From Machine to Machine: An OCT-trained Deep Learning Algorithm for Objective Quantification of Glaucomatous Damage in Fundus Photographs

    Felipe A. Medeiros, Alessandro A. Jammal, Atalie C. Thompson

    cs.CVcs.AIarXiv:1810.10343v12018
  8. Learning to Manipulate Deformable Objects without Demonstrations

    Yilin Wu, Wilson Yan, Thanard Kurutach +2

    cs.ROcs.CVcs.LGarXiv:1910.13439v22019
  9. GNeRF: GAN-based Neural Radiance Field without Posed Camera

    Quan Meng, Anpei Chen, Haimin Luo +5

    cs.CVarXiv:2103.15606v32021
  10. How to represent part-whole hierarchies in a neural network

    Geoffrey Hinton

    cs.CVarXiv:2102.12627v12021
  11. Next-ViT: Next Generation Vision Transformer for Efficient Deployment in Realistic Industrial Scenarios

    Jiashi Li, Xin Xia, Wei Li +6

    cs.CVarXiv:2207.05501v42022
  12. Dynamic-Robust Photometric-Semantic Reconstruction for Open-Vocabulary 3D Scene Understanding

    Boyu Cai, Li Yang, Yan Xu +6

    cs.CVarXiv:2608.29177v12026
  13. Generalized Max Pooling

    Naila Murray, Florent Perronnin

    cs.CVarXiv:1406.0312v12014
  14. Ground-to-Satellite Localization in Unconstrained Image Collections for 3D Scene Reconstruction

    Angel Daruna, Ben Southall, Niluthpol Chowdhury Mithun +6

    cs.CVarXiv:2608.29211v12026
  15. PERSIST: Persistent-State Discrimination for Shot Boundary Detection

    Tingyu Lin, Christian Stippel, Armin Dadras +4

    cs.CVarXiv:2608.29287v12026
  16. Demographic Inference and Representative Population Estimates from Multilingual Social Media Data

    Zijian Wang, Scott A. Hale, David Adelani +4

    cs.CYcs.CLcs.CVarXiv:1905.05961v12019
  17. Camouflaged Object Detection via Context-aware Cross-level Fusion

    Geng Chen, Si-Jie Liu, Yu-Jia Sun +3

    cs.CVarXiv:2207.13362v12022
  18. Density-aware Chamfer Distance as a Comprehensive Metric for Point Cloud Completion

    Tong Wu, Liang Pan, Junzhe Zhang +3

    cs.CVarXiv:2111.12702v12021
  19. Occlusion Robust Face Recognition Based on Mask Learning with PairwiseDifferential Siamese Network

    Lingxue Song, Dihong Gong, Zhifeng Li +2

    cs.CVarXiv:1908.06290v12019
  20. Domain Adaptation without Source Data

    Youngeun Kim, Donghyeon Cho, Kyeongtak Han +2

    cs.CVcs.LGeess.IVarXiv:2007.01524v42020
  21. Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

    Suzhen Wang, Lincheng Li, Yu Ding +2

    cs.CVcs.CLarXiv:2107.09293v12021
  22. Optimal Feature Transport for Cross-View Image Geo-Localization

    Yujiao Shi, Xin Yu, Liu Liu +2

    cs.CVarXiv:1907.05021v32019
  23. DeepCore: A Comprehensive Library for Coreset Selection in Deep Learning

    Chengcheng Guo, Bo Zhao, Yanbing Bai

    cs.LGcs.CVarXiv:2204.08499v32022
  24. Generalizable Data-free Objective for Crafting Universal Adversarial Perturbations

    Konda Reddy Mopuri, Aditya Ganeshan, R. Venkatesh Babu

    cs.CVcs.AIcs.LGarXiv:1801.08092v32018
  25. Joint Enhancement and Denoising Method via Sequential Decomposition

    Xutong Ren, Mading Li, Wen-Huang Cheng +1

    cs.CVarXiv:1804.08468v32018
  26. Provable Defense against Privacy Leakage in Federated Learning from Representation Perspective

    Jingwei Sun, Ang Li, Binghui Wang +3

    cs.LGcs.AIcs.CVarXiv:2012.06043v12020
  27. Learning to Predict Layout-to-image Conditional Convolutions for Semantic Image Synthesis

    Xihui Liu, Guojun Yin, Jing Shao +2

    cs.CVcs.LGeess.IVarXiv:1910.06809v32019
  28. BIVA: A Very Deep Hierarchy of Latent Variables for Generative Modeling

    Lars Maaløe, Marco Fraccaro, Valentin Liévin +1

    stat.MLcs.CVcs.LGarXiv:1902.02102v32019
  29. Summary Transfer: Exemplar-based Subset Selection for Video Summarization

    Ke Zhang, Wei-Lun Chao, Fei Sha +1

    cs.CVarXiv:1603.03369v32016
  30. Exploring Deep Neural Networks via Layer-Peeled Model: Minority Collapse in Imbalanced Training

    Cong Fang, Hangfeng He, Qi Long +1

    cs.LGcs.CVmath.OCarXiv:2101.12699v32021
  31. AGRICAM: A Track-Mounted Crop Pollination Monitoring Robot

    Malika Nisal Ratnayake, Adel N. Toosi, James Cook +2

    cs.CVcs.AIcs.ROarXiv:2608.29237v12026
  32. Align and Prompt: Video-and-Language Pre-training with Entity Prompts

    Dongxu Li, Junnan Li, Hongdong Li +2

    cs.CVarXiv:2112.09583v22021
  33. The Majority Can Help The Minority: Context-rich Minority Oversampling for Long-tailed Classification

    Seulki Park, Youngkyu Hong, Byeongho Heo +2

    cs.CVcs.AIarXiv:2112.00412v32021
  34. SeedFormer: Patch Seeds based Point Cloud Completion with Upsample Transformer

    Haoran Zhou, Yun Cao, Wenqing Chu +4

    cs.CVarXiv:2207.10315v12022
  35. Dual Memory Units with Uncertainty Regulation for Weakly Supervised Video Anomaly Detection

    Hang Zhou, Junqing Yu, Wei Yang

    cs.CVarXiv:2302.05160v12023
  36. VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks

    Ziyan Jiang, Rui Meng, Xinyi Yang +3

    cs.CVcs.AIcs.CLarXiv:2410.05160v32024
  37. A Survey on Deep Learning-based Architectures for Semantic Segmentation on 2D images

    Irem Ulku, Erdem Akagunduz

    cs.CVarXiv:1912.10230v52019
  38. Exploring large scale public medical image datasets

    Luke Oakden-Rayner

    eess.IVcs.CVcs.LGarXiv:1907.12720v12019
  39. DSGN: Deep Stereo Geometry Network for 3D Object Detection

    Yilun Chen, Shu Liu, Xiaoyong Shen +1

    cs.CVarXiv:2001.03398v32020
  40. Facial Emotion Recognition: State of the Art Performance on FER2013

    Yousif Khaireddin, Zhuofa Chen

    cs.CVcs.AIcs.LGarXiv:2105.03588v12021
  41. Hyper3-CLIP: Hierarchy-Conditioned Hyperbolic Vision-Language Training

    Matin Mahmood, Antonio Rueda-Toicen, Mohamed ElBassat +3

    cs.CVcs.AIarXiv:2608.29313v12026
  42. A review of deep learning-based information fusion techniques for multimodal medical image classification

    Yihao Li, Mostafa El Habib Daho, Pierre-Henri Conze +6

    cs.CVcs.AIarXiv:2404.15022v12024
  43. Elastic Token Compression for Pixel-Space Diffusion Transformers

    Eduard Zamfir, Christian Reisswig, Zongwei Wu +2

    cs.CVarXiv:2608.29281v12026
  44. Pan-Mamba: Effective pan-sharpening with State Space Model

    Xuanhua He, Ke Cao, Keyu Yan +4

    cs.CVarXiv:2402.12192v22024
  45. Training Deeper Convolutional Networks with Deep Supervision

    Liwei Wang, Chen-Yu Lee, Zhuowen Tu +1

    cs.CVarXiv:1505.02496v12015
  46. Deep Learning Fundus Image Analysis for Diabetic Retinopathy and Macular Edema Grading

    Jaakko Sahlsten, Joel Jaskari, Jyri Kivinen +4

    eess.IVcs.CVcs.LGarXiv:1904.08764v12019
  47. ID-Reveal: Identity-aware DeepFake Video Detection

    Davide Cozzolino, Andreas Rössler, Justus Thies +2

    cs.CVarXiv:2012.02512v32020
  48. Voxel Set Transformer: A Set-to-Set Approach to 3D Object Detection from Point Clouds

    Chenhang He, Ruihuang Li, Shuai Li +1

    cs.CVarXiv:2203.10314v12022
  49. JourneyDB: A Benchmark for Generative Image Understanding

    Keqiang Sun, Junting Pan, Yuying Ge +11

    cs.CVarXiv:2307.00716v22023
  50. PlainMamba: Improving Non-Hierarchical Mamba in Visual Recognition

    Chenhongyi Yang, Zehui Chen, Miguel Espinosa +4

    cs.CVcs.LGarXiv:2403.17695v22024
  51. Improved Residual Networks for Image and Video Recognition

    Ionut Cosmin Duta, Li Liu, Fan Zhu +1

    cs.CVarXiv:2004.04989v12020
  52. Deep Generative Adversarial Compression Artifact Removal

    Leonardo Galteri, Lorenzo Seidenari, Marco Bertini +1

    cs.CVarXiv:1704.02518v32017
  53. From Deterministic to Generative: Multi-Modal Stochastic RNNs for Video Captioning

    Jingkuan Song, Yuyu Guo, Lianli Gao +3

    cs.CVarXiv:1708.02478v22017
  54. Evaluating Constrained Iterative Refinement for Scalable Vector Graphics Generation with Off-the-Shelf VLMs

    Matthew Perlman, James Beetham, Niels Da Vitoria Lobo +2

    cs.CVcs.GRarXiv:2608.28678v12026
  55. Higher Order Conditional Random Fields in Deep Neural Networks

    Anurag Arnab, Sadeep Jayasumana, Shuai Zheng +1

    cs.CVarXiv:1511.08119v42015
  56. A Generalized Deep Learning Framework for Whole-Slide Image Segmentation and Analysis

    Mahendra Khened, Avinash Kori, Haran Rajkumar +2

    eess.IVcs.CVcs.LGarXiv:2001.00258v22020
  57. Coherent Multi-Sentence Video Description with Variable Level of Detail

    Anna Senina, Marcus Rohrbach, Wei Qiu +5

    cs.CVcs.CLarXiv:1403.6173v12014
  58. Multi-Sensor Mapping of Vulnerable Urban Settlements Using SAR, Multispectral, and Hyperspectral Imagery: A Case Study in Córdoba, Argentina

    Luigi Russo, Anabella Ferral, Silvia Liberata Ullo +1

    cs.CVarXiv:2608.28680v12026
  59. StableVideo: Text-driven Consistency-aware Diffusion Video Editing

    Wenhao Chai, Xun Guo, Gaoang Wang +1

    cs.CVarXiv:2308.09592v12023
  60. Automated pipeline for herbarium label digitization

    Hiba Abbad, Hanane Ariouat, Eva Perez Pimpare +6

    cs.CVarXiv:2608.28676v12026