Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1,261 to 1,320 of 18,855

  1. Multi-Branch Auxiliary Fusion YOLO with Re-parameterization Heterogeneous Convolutional for accurate object detection

    Zhiqiang Yang, Qiu Guan, Keer Zhao +4

    cs.CVcs.AIarXiv:2407.04381v12024
  2. Deep learning is a good steganalysis tool when embedding key is reused for different images, even if there is a cover source-mismatch

    Lionel Pibre, Pasquet Jérôme, Dino Ienco +1

    cs.MMcs.CVcs.LGarXiv:1511.04855v22015
  3. KVT: k-NN Attention for Boosting Vision Transformers

    Pichao Wang, Xue Wang, Fan Wang +4

    cs.CVarXiv:2106.00515v32021
  4. The Lazy Neuron Phenomenon: On Emergence of Activation Sparsity in Transformers

    Zonglin Li, Chong You, Srinadh Bhojanapalli +8

    cs.LGcs.CLcs.CVarXiv:2210.06313v22022
  5. ECG Arrhythmia Classification Using Transfer Learning from 2-Dimensional Deep CNN Features

    Milad Salem, Shayan Taheri, Jiann Shiun-Yuan

    cs.LGcs.CVstat.MLarXiv:1812.04693v12018
  6. Stereo Vision-based Semantic 3D Object and Ego-motion Tracking for Autonomous Driving

    Peiliang Li, Tong Qin, Shaojie Shen

    cs.CVarXiv:1807.02062v32018
  7. SynthRAD2023 Grand Challenge dataset: generating synthetic CT for radiotherapy

    Adrian Thummerer, Erik van der Bijl, Arthur Jr Galapon +6

    physics.med-phcs.CVarXiv:2303.16320v12023
  8. DiffiT: Diffusion Vision Transformers for Image Generation

    Ali Hatamizadeh, Jiaming Song, Guilin Liu +2

    cs.CVcs.AIcs.LGarXiv:2312.02139v32023
  9. Low Resolution Face Recognition Using a Two-Branch Deep Convolutional Neural Network Architecture

    Erfan Zangeneh, Mohammad Rahmati, Yalda Mohsenzadeh

    cs.CVarXiv:1706.06247v12017
  10. Neural Task Graphs: Generalizing to Unseen Tasks from a Single Video Demonstration

    De-An Huang, Suraj Nair, Danfei Xu +5

    cs.CVcs.AIcs.LGarXiv:1807.03480v22018
  11. Time Series Anomaly Detection Using Convolutional Neural Networks and Transfer Learning

    Tailai Wen, Roy Keyes

    cs.LGcs.CVstat.MLarXiv:1905.13628v12019
  12. 3D Human Pose Estimation using Spatio-Temporal Networks with Explicit Occlusion Training

    Yu Cheng, Bo Yang, Bo Wang +1

    cs.CVarXiv:2004.11822v12020
  13. A Deep Convolutional Neural Network for COVID-19 Detection Using Chest X-Rays

    Pedro R. A. S. Bassi, Romis Attux

    eess.IVcs.CVcs.LGarXiv:2005.01578v42020
  14. Learning rotation invariant convolutional filters for texture classification

    Diego Marcos, Michele Volpi, Devis Tuia

    cs.CVarXiv:1604.06720v22016
  15. End-to-End Saliency Mapping via Probability Distribution Prediction

    Saumya Jetley, Naila Murray, Eleonora Vig

    cs.CVcs.AIarXiv:1804.01793v12018
  16. Deep CTR Prediction in Display Advertising

    Junxuan Chen, Baigui Sun, Hao Li +2

    cs.CVcs.MMarXiv:1609.06018v12016
  17. KeepAugment: A Simple Information-Preserving Data Augmentation Approach

    Chengyue Gong, Dilin Wang, Meng Li +2

    cs.CVarXiv:2011.11778v12020
  18. Collaborative Discrepancy Optimization for Reliable Image Anomaly Localization

    Yunkang Cao, Xiaohao Xu, Zhaoge Liu +1

    cs.CVcs.AIarXiv:2302.08769v12023
  19. A Holistic Visual Place Recognition Approach using Lightweight CNNs for Significant ViewPoint and Appearance Changes

    Ahmad Khaliq, Shoaib Ehsan, Zetao Chen +2

    cs.ROcs.CVarXiv:1811.03032v42018
  20. Vision language models are blind: Failing to translate detailed visual features into words

    Pooyan Rahmanzadehgervi, Logan Bolton, Mohammad Reza Taesiri +1

    cs.AIcs.CVarXiv:2407.06581v62024
  21. Retinexmamba: Retinex-based Mamba for Low-light Image Enhancement

    Jiesong Bai, Yuhao Yin, Qiyuan He +2

    cs.CVarXiv:2405.03349v22024
  22. Deepfakes Detection with Automatic Face Weighting

    Daniel Mas Montserrat, Hanxiang Hao, S. K. Yarlagadda +8

    cs.CVeess.IVarXiv:2004.12027v22020
  23. Face Anti-Spoofing Via Disentangled Representation Learning

    Ke-Yue Zhang, Taiping Yao, Jian Zhang +6

    cs.CVarXiv:2008.08250v12020
  24. ASR is all you need: cross-modal distillation for lip reading

    Triantafyllos Afouras, Joon Son Chung, Andrew Zisserman

    cs.CVcs.SDeess.ASarXiv:1911.12747v22019
  25. Shape-Aware Organ Segmentation by Predicting Signed Distance Maps

    Yuan Xue, Hui Tang, Zhi Qiao +6

    cs.CVarXiv:1912.03849v12019
  26. Recurrent Convolutional Neural Networks for Scene Parsing

    Pedro H. O. Pinheiro, Ronan Collobert

    cs.CVarXiv:1306.2795v12013
  27. Towards neural networks that provably know when they don't know

    Alexander Meinke, Matthias Hein

    cs.LGcs.CVstat.MLarXiv:1909.12180v22019
  28. SDCNet: Video Prediction Using Spatially-Displaced Convolution

    Fitsum A. Reda, Guilin Liu, Kevin J. Shih +5

    cs.CVarXiv:1811.00684v22018
  29. High Fidelity Video Prediction with Large Stochastic Recurrent Neural Networks

    Ruben Villegas, Arkanath Pathak, Harini Kannan +3

    cs.CVarXiv:1911.01655v12019
  30. Recurrent Attention Models for Depth-Based Person Identification

    Albert Haque, Alexandre Alahi, Li Fei-Fei

    cs.CVarXiv:1611.07212v12016
  31. Weakly-supervised localization of diabetic retinopathy lesions in retinal fundus images

    Waleed M. Gondal, Jan M. Köhler, René Grzeszick +2

    cs.CVarXiv:1706.09634v12017
  32. Context-aware Captions from Context-agnostic Supervision

    Ramakrishna Vedantam, Samy Bengio, Kevin Murphy +2

    cs.CVcs.AIarXiv:1701.02870v32017
  33. FasterViT: Fast Vision Transformers with Hierarchical Attention

    Ali Hatamizadeh, Greg Heinrich, Hongxu Yin +4

    cs.CVcs.AIcs.LGarXiv:2306.06189v22023
  34. Semantics Disentangling for Generalized Zero-Shot Learning

    Zhi Chen, Yadan Luo, Ruihong Qiu +4

    cs.CVarXiv:2101.07978v52021
  35. Underwater Image Enhancement by Transformer-based Diffusion Model with Non-uniform Sampling for Skip Strategy

    Yi Tang, Takafumi Iwaguchi, Hiroshi Kawasaki

    cs.CVarXiv:2309.03445v12023
  36. Active Learning for Deep Object Detection via Probabilistic Modeling

    Jiwoong Choi, Ismail Elezi, Hyuk-Jae Lee +2

    cs.CVarXiv:2103.16130v22021
  37. Face Recognition: Too Bias, or Not Too Bias?

    Joseph P Robinson, Gennady Livitz, Yann Henon +3

    cs.CVarXiv:2002.06483v42020
  38. COAST: COntrollable Arbitrary-Sampling NeTwork for Compressive Sensing

    Di You, Jian Zhang, Jingfen Xie +2

    cs.CVeess.IVarXiv:2107.07225v12021
  39. DAG-Recurrent Neural Networks For Scene Labeling

    Bing Shuai, Zhen Zuo, Gang Wang +1

    cs.CVarXiv:1509.00552v22015
  40. VCT: A Video Compression Transformer

    Fabian Mentzer, George Toderici, David Minnen +4

    cs.CVcs.LGeess.IVarXiv:2206.07307v22022
  41. Backdoor Attack in the Physical World

    Yiming Li, Tongqing Zhai, Yong Jiang +2

    cs.CRcs.AIcs.CVarXiv:2104.02361v22021
  42. Bifurcated backbone strategy for RGB-D salient object detection

    Yingjie Zhai, Deng-Ping Fan, Jufeng Yang +4

    cs.CVarXiv:2007.02713v32020
  43. On Network Design Spaces for Visual Recognition

    Ilija Radosavovic, Justin Johnson, Saining Xie +2

    cs.CVcs.LGarXiv:1905.13214v12019
  44. Spatial Aggregation of Holistically-Nested Networks for Automated Pancreas Segmentation

    Holger R. Roth, Le Lu, Amal Farag +2

    cs.CVarXiv:1606.07830v12016
  45. Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension

    Yongdong Luo, Xiawu Zheng, Guilin Li +8

    cs.CVcs.AIarXiv:2411.13093v42024
  46. Faster R-CNN Features for Instance Search

    Amaia Salvador, Xavier Giro-i-Nieto, Ferran Marques +1

    cs.CVarXiv:1604.08893v12016
  47. Few-shot Adaptive Faster R-CNN

    Tao Wang, Xiaopeng Zhang, Li Yuan +1

    cs.CVarXiv:1903.09372v12019
  48. Rethinking Inductive Biases for Surface Normal Estimation

    Gwangbin Bae, Andrew J. Davison

    cs.CVarXiv:2403.00712v12024
  49. Improving Generalization via Scalable Neighborhood Component Analysis

    Zhirong Wu, Alexei A. Efros, Stella X. Yu

    cs.CVcs.LGarXiv:1808.04699v12018
  50. Unsupervised Domain-Specific Deblurring via Disentangled Representations

    Boyu Lu, Jun-Cheng Chen, Rama Chellappa

    cs.CVarXiv:1903.01594v22019
  51. HiFaceGAN: Face Renovation via Collaborative Suppression and Replenishment

    Lingbo Yang, Chang Liu, Pan Wang +4

    cs.CVarXiv:2005.05005v22020
  52. MultiHuSE: A Multimodal Dataset for Humour Styles and Emotions

    Mary Ogbuka Kenneth, Foaad Khosmood, Abbas Edalat

    cs.CLcs.CVcs.MMarXiv:2609.11322v12026
  53. FilterReg: Robust and Efficient Probabilistic Point-Set Registration using Gaussian Filter and Twist Parameterization

    Wei Gao, Russ Tedrake

    cs.CVarXiv:1811.10136v32018
  54. RetinaMask: Learning to predict masks improves state-of-the-art single-shot detection for free

    Cheng-Yang Fu, Mykhailo Shvets, Alexander C. Berg

    cs.CVarXiv:1901.03353v12019
  55. On Recognizing Texts of Arbitrary Shapes with 2D Self-Attention

    Junyeop Lee, Sungrae Park, Jeonghun Baek +3

    cs.CVarXiv:1910.04396v12019
  56. Robust Mean Teacher for Continual and Gradual Test-Time Adaptation

    Mario Döbler, Robert A. Marsden, Bin Yang

    cs.CVarXiv:2211.13081v22022
  57. DiT-3D: Exploring Plain Diffusion Transformers for 3D Shape Generation

    Shentong Mo, Enze Xie, Ruihang Chu +4

    cs.CVcs.AIcs.LGarXiv:2307.01831v12023
  58. Wavelet-Based Dual-Branch Network for Image Demoireing

    Lin Liu, Jianzhuang Liu, Shanxin Yuan +4

    cs.CVarXiv:2007.07173v22020
  59. Structured Bird's-Eye-View Traffic Scene Understanding from Onboard Images

    Yigit Baran Can, Alexander Liniger, Danda Pani Paudel +1

    cs.CVarXiv:2110.01997v12021
  60. 3DIoUMatch: Leveraging IoU Prediction for Semi-Supervised 3D Object Detection

    He Wang, Yezhen Cong, Or Litany +2

    cs.CVarXiv:2012.04355v32020