Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

7,261 to 7,320 of 18,971

  1. Deep Gradient Projection Networks for Pan-sharpening

    Shuang Xu, Jiangshe Zhang, Zixiang Zhao +3

    cs.CVeess.IVarXiv:2103.04584v12021
  2. Deep Unsupervised Saliency Detection: A Multiple Noisy Labeling Perspective

    Jing Zhang, Tong Zhang, Yuchao Dai +2

    cs.CVarXiv:1803.10910v12018
  3. Maximum-Entropy Adversarial Data Augmentation for Improved Generalization and Robustness

    Long Zhao, Ting Liu, Xi Peng +1

    cs.LGcs.CVarXiv:2010.08001v22020
  4. Agentic Multimodal Models for Environmental Hyperspectral Unmixing

    Michał Cholewa, Luca Ciampi, Nicola Messina +2

    cs.CVarXiv:2609.01289v12026
  5. Level Playing Field for Million Scale Face Recognition

    Aaron Nech, Ira Kemelmacher-Shlizerman

    cs.CVarXiv:1705.00393v12017
  6. Fingerprint Spoof Buster

    Tarang Chugh, Kai Cao, Anil K. Jain

    cs.CVarXiv:1712.04489v12017
  7. Instant Volumetric Head Avatars

    Wojciech Zielonka, Timo Bolkart, Justus Thies

    cs.CVarXiv:2211.12499v22022
  8. Something-Else: Compositional Action Recognition with Spatial-Temporal Interaction Networks

    Joanna Materzynska, Tete Xiao, Roei Herzig +3

    cs.CVarXiv:1912.09930v32019
  9. Improved Automatic Target Recognition in Synthetic Aperture Sonar Imagery Using Large Deep Neural Networks

    C. J. Moore, Alex Hurt, Jordan Malof

    cs.CVarXiv:2609.01800v12026
  10. BEVSegFormer: Bird's Eye View Semantic Segmentation From Arbitrary Camera Rigs

    Lang Peng, Zhirong Chen, Zhangjie Fu +2

    cs.CVarXiv:2203.04050v32022
  11. Long-Tailed Recognition via Weight Balancing

    Shaden Alshammari, Yu-Xiong Wang, Deva Ramanan +1

    cs.CVarXiv:2203.14197v12022
  12. A Survey on Long-Tailed Visual Recognition

    Lu Yang, He Jiang, Qing Song +1

    cs.CVarXiv:2205.13775v12022
  13. An End-to-End Transformer Model for Crowd Localization

    Dingkang Liang, Wei Xu, Xiang Bai

    cs.CVarXiv:2202.13065v22022
  14. FeTrIL: Feature Translation for Exemplar-Free Class-Incremental Learning

    Grégoire Petit, Adrian Popescu, Hugo Schindler +2

    cs.CVcs.AIcs.LGarXiv:2211.13131v22022
  15. Remote Sensing Cross-Modal Text-Image Retrieval Based on Global and Local Information

    Zhiqiang Yuan, Wenkai Zhang, Changyuan Tian +5

    cs.CVcs.IRcs.MMarXiv:2204.09860v12022
  16. CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical Flow

    Philippe Weinzaepfel, Thomas Lucas, Vincent Leroy +7

    cs.CVarXiv:2211.10408v32022
  17. Occupancy Anticipation for Efficient Exploration and Navigation

    Santhosh K. Ramakrishnan, Ziad Al-Halah, Kristen Grauman

    cs.CVarXiv:2008.09285v22020
  18. Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental Learning

    Kai Zhu, Wei Zhai, Yang Cao +2

    cs.CVarXiv:2203.06359v22022
  19. Catching Both Gray and Black Swans: Open-set Supervised Anomaly Detection

    Choubo Ding, Guansong Pang, Chunhua Shen

    cs.CVarXiv:2203.14506v12022
  20. Fully Convolutional Networks for Continuous Sign Language Recognition

    Ka Leong Cheng, Zhaoyang Yang, Qifeng Chen +1

    cs.CVarXiv:2007.12402v12020
  21. MonoDETR: Depth-guided Transformer for Monocular 3D Object Detection

    Renrui Zhang, Han Qiu, Tai Wang +7

    cs.CVcs.AIeess.IVarXiv:2203.13310v52022
  22. Rethinking Depth Estimation for Multi-View Stereo: A Unified Representation

    Rui Peng, Rongjie Wang, Zhenyu Wang +2

    cs.CVcs.AIarXiv:2201.01501v32022
  23. RepQ-ViT: Scale Reparameterization for Post-Training Quantization of Vision Transformers

    Zhikai Li, Junrui Xiao, Lianwei Yang +1

    cs.CVcs.LGarXiv:2212.08254v22022
  24. Gaussian Shading: Provable Performance-Lossless Image Watermarking for Diffusion Models

    Zijin Yang, Kai Zeng, Kejiang Chen +3

    cs.CVcs.CRarXiv:2404.04956v32024
  25. Visual Recognition with Deep Nearest Centroids

    Wenguan Wang, Cheng Han, Tianfei Zhou +1

    cs.CVarXiv:2209.07383v22022
  26. Refining activation downsampling with SoftPool

    Alexandros Stergiou, Ronald Poppe, Grigorios Kalliatakis

    cs.CVarXiv:2101.00440v32021
  27. Spherical CNNs on Unstructured Grids

    Chiyu "Max" Jiang, Jingwei Huang, Karthik Kashinath +3

    cs.CVcs.AIcs.LGarXiv:1901.02039v12019
  28. VQFR: Blind Face Restoration with Vector-Quantized Dictionary and Parallel Decoder

    Yuchao Gu, Xintao Wang, Liangbin Xie +4

    cs.CVarXiv:2205.06803v32022
  29. CGIntrinsics: Better Intrinsic Image Decomposition through Physically-Based Rendering

    Zhengqi Li, Noah Snavely

    cs.CVarXiv:1808.08601v32018
  30. Diffusion Models for Image Restoration and Enhancement: A Comprehensive Survey

    Xin Li, Yulin Ren, Xin Jin +5

    cs.CVarXiv:2308.09388v32023
  31. Fast and Accurate Tumor Segmentation of Histology Images using Persistent Homology and Deep Convolutional Features

    Talha Qaiser, Yee-Wah Tsang, Daiki Taniyama +4

    cs.CVarXiv:1805.03699v12018
  32. A Benchmark for Vehicle Attribute Classification in Cross-Domain Surveillance Scenarios

    Sergio M. Silva, Otavio T. Remer, Gabriel E. Lima +3

    cs.CVarXiv:2609.01584v12026
  33. Optimizing the Trade-off between Single-Stage and Two-Stage Object Detectors using Image Difficulty Prediction

    Petru Soviany, Radu Tudor Ionescu

    cs.CVarXiv:1803.08707v32018
  34. InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

    Pan Zhang, Xiaoyi Dong, Yuhang Zang +24

    cs.CVcs.CLarXiv:2407.03320v12024
  35. What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models

    Sharon S. Musa, Fereshteh Forghani, Harrish Thasarathan +3

    cs.CVarXiv:2609.01551v12026
  36. Extreme View Synthesis

    Inchang Choi, Orazio Gallo, Alejandro Troccoli +2

    cs.CVarXiv:1812.04777v22018
  37. PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-time Execution on Mobile Devices

    Xiaolong Ma, Fu-Ming Guo, Wei Niu +5

    cs.LGcs.CVcs.DCarXiv:1909.05073v42019
  38. Editable Visual Design

    Junyan Ye, Wei Liu, Dongzhi Jiang +9

    cs.CVcs.CLarXiv:2609.04034v12026
  39. Introduction to the Bag of Features Paradigm for Image Classification and Retrieval

    Stephen O'Hara, Bruce A. Draper

    cs.CVcs.IRarXiv:1101.3354v12011
  40. Bag of Visual Words and Fusion Methods for Action Recognition: Comprehensive Study and Good Practice

    Xiaojiang Peng, Limin Wang, Xingxing Wang +1

    cs.CVarXiv:1405.4506v12014
  41. Beauty is in the AI of the beholder: MLLMs systematically overrate facial attractiveness

    Santiago Grandas, Juan Sebastian Cely-Acosta, Mohit Mendiratta +2

    cs.CVcs.HCarXiv:2609.02512v12026
  42. AutoCompress: An Automatic DNN Structured Pruning Framework for Ultra-High Compression Rates

    Ning Liu, Xiaolong Ma, Zhiyuan Xu +3

    cs.LGcs.AIcs.CVarXiv:1907.03141v22019
  43. Learning to Generate Images with Perceptual Similarity Metrics

    Jake Snell, Karl Ridgeway, Renjie Liao +3

    cs.LGcs.CVarXiv:1511.06409v32015
  44. Structured Prediction Helps 3D Human Motion Modelling

    Emre Aksan, Manuel Kaufmann, Otmar Hilliges

    cs.CVarXiv:1910.09070v12019
  45. CenterFormer: Center-based Transformer for 3D Object Detection

    Zixiang Zhou, Xiangchen Zhao, Yu Wang +2

    cs.CVarXiv:2209.05588v12022
  46. Adaptive Unimodal Cost Volume Filtering for Deep Stereo Matching

    Youmin Zhang, Yimin Chen, Xiao Bai +4

    cs.CVarXiv:1909.03751v22019
  47. Evaluating the Impact of Intensity Normalization on MR Image Synthesis

    Jacob C. Reinhold, Blake E. Dewey, Aaron Carass +1

    cs.CVarXiv:1812.04652v12018
  48. Real-time Driver Drowsiness Detection for Android Application Using Deep Neural Networks Techniques

    Rateb Jabbar, Khalifa Al-Khalifa, Mohamed Kharbeche +3

    cs.CVcs.HCarXiv:1811.01627v12018
  49. Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image Understanding

    Shanji Liu, Kelu Yao, Junxiao Xue +5

    cs.CVarXiv:2609.02187v12026
  50. Human-centric Indoor Scene Synthesis Using Stochastic Grammar

    Siyuan Qi, Yixin Zhu, Siyuan Huang +2

    cs.CVarXiv:1808.08473v12018
  51. KSG-Net: Key-Sparse and Global-Context Learning for Maritime 3D Ship Detection

    Zhouyuan Huai, Meiqi Wan, Yan Yang +4

    cs.CVarXiv:2609.02077v12026
  52. Looking Beyond Appearances: Synthetic Training Data for Deep CNNs in Re-identification

    Igor Barros Barbosa, Marco Cristani, Barbara Caputo +2

    cs.CVarXiv:1701.03153v22017
  53. CubeMLP: An MLP-based Model for Multimodal Sentiment Analysis and Depression Estimation

    Hao Sun, Hongyi Wang, Jiaqing Liu +2

    cs.MMcs.CLcs.CVarXiv:2207.14087v32022
  54. Low Frequency Adversarial Perturbation

    Chuan Guo, Jared S. Frank, Kilian Q. Weinberger

    cs.CVarXiv:1809.08758v22018
  55. Learning Semantic-Aware Knowledge Guidance for Low-Light Image Enhancement

    Yuhui Wu, Chen Pan, Guoqing Wang +4

    cs.CVarXiv:2304.07039v12023
  56. A Neural Temporal Model for Human Motion Prediction

    Anand Gopalakrishnan, Ankur Mali, Dan Kifer +2

    cs.CVarXiv:1809.03036v52018
  57. ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding

    Jitai Hao, Ke Yang, Qiang Huang +1

    cs.CVcs.CLarXiv:2609.02780v12026
  58. Cross-Dataset Person Re-Identification via Unsupervised Pose Disentanglement and Adaptation

    Yu-Jhe Li, Ci-Siang Lin, Yan-Bo Lin +1

    cs.CVarXiv:1909.09675v12019
  59. TACO: Trash Annotations in Context for Litter Detection

    Pedro F Proença, Pedro Simões

    cs.CVarXiv:2003.06975v22020
  60. CLIP-Driven Fine-grained Text-Image Person Re-identification

    Shuanglin Yan, Neng Dong, Liyan Zhang +1

    cs.CVarXiv:2210.10276v12022