Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
1,201 to 1,260 of 18,795
Multi-Branch Auxiliary Fusion YOLO with Re-parameterization Heterogeneous Convolutional for accurate object detection
Zhiqiang Yang, Qiu Guan, Keer Zhao +4
cs.CVcs.AIarXiv:2407.04381v12024Deep learning is a good steganalysis tool when embedding key is reused for different images, even if there is a cover source-mismatch
Lionel Pibre, Pasquet Jérôme, Dino Ienco +1
cs.MMcs.CVcs.LGarXiv:1511.04855v22015KVT: k-NN Attention for Boosting Vision Transformers
Pichao Wang, Xue Wang, Fan Wang +4
cs.CVarXiv:2106.00515v32021The Lazy Neuron Phenomenon: On Emergence of Activation Sparsity in Transformers
Zonglin Li, Chong You, Srinadh Bhojanapalli +8
cs.LGcs.CLcs.CVarXiv:2210.06313v22022ECG Arrhythmia Classification Using Transfer Learning from 2-Dimensional Deep CNN Features
Milad Salem, Shayan Taheri, Jiann Shiun-Yuan
cs.LGcs.CVstat.MLarXiv:1812.04693v12018Stereo Vision-based Semantic 3D Object and Ego-motion Tracking for Autonomous Driving
Peiliang Li, Tong Qin, Shaojie Shen
cs.CVarXiv:1807.02062v32018SynthRAD2023 Grand Challenge dataset: generating synthetic CT for radiotherapy
Adrian Thummerer, Erik van der Bijl, Arthur Jr Galapon +6
physics.med-phcs.CVarXiv:2303.16320v12023DiffiT: Diffusion Vision Transformers for Image Generation
Ali Hatamizadeh, Jiaming Song, Guilin Liu +2
cs.CVcs.AIcs.LGarXiv:2312.02139v32023Low Resolution Face Recognition Using a Two-Branch Deep Convolutional Neural Network Architecture
Erfan Zangeneh, Mohammad Rahmati, Yalda Mohsenzadeh
cs.CVarXiv:1706.06247v12017Neural Task Graphs: Generalizing to Unseen Tasks from a Single Video Demonstration
De-An Huang, Suraj Nair, Danfei Xu +5
cs.CVcs.AIcs.LGarXiv:1807.03480v22018Time Series Anomaly Detection Using Convolutional Neural Networks and Transfer Learning
Tailai Wen, Roy Keyes
cs.LGcs.CVstat.MLarXiv:1905.13628v120193D Human Pose Estimation using Spatio-Temporal Networks with Explicit Occlusion Training
Yu Cheng, Bo Yang, Bo Wang +1
cs.CVarXiv:2004.11822v12020A Deep Convolutional Neural Network for COVID-19 Detection Using Chest X-Rays
Pedro R. A. S. Bassi, Romis Attux
eess.IVcs.CVcs.LGarXiv:2005.01578v42020Learning rotation invariant convolutional filters for texture classification
Diego Marcos, Michele Volpi, Devis Tuia
cs.CVarXiv:1604.06720v22016End-to-End Saliency Mapping via Probability Distribution Prediction
Saumya Jetley, Naila Murray, Eleonora Vig
cs.CVcs.AIarXiv:1804.01793v12018Deep CTR Prediction in Display Advertising
Junxuan Chen, Baigui Sun, Hao Li +2
cs.CVcs.MMarXiv:1609.06018v12016KeepAugment: A Simple Information-Preserving Data Augmentation Approach
Chengyue Gong, Dilin Wang, Meng Li +2
cs.CVarXiv:2011.11778v12020Collaborative Discrepancy Optimization for Reliable Image Anomaly Localization
Yunkang Cao, Xiaohao Xu, Zhaoge Liu +1
cs.CVcs.AIarXiv:2302.08769v12023A Holistic Visual Place Recognition Approach using Lightweight CNNs for Significant ViewPoint and Appearance Changes
Ahmad Khaliq, Shoaib Ehsan, Zetao Chen +2
cs.ROcs.CVarXiv:1811.03032v42018Vision language models are blind: Failing to translate detailed visual features into words
Pooyan Rahmanzadehgervi, Logan Bolton, Mohammad Reza Taesiri +1
cs.AIcs.CVarXiv:2407.06581v62024Retinexmamba: Retinex-based Mamba for Low-light Image Enhancement
Jiesong Bai, Yuhao Yin, Qiyuan He +2
cs.CVarXiv:2405.03349v22024Deepfakes Detection with Automatic Face Weighting
Daniel Mas Montserrat, Hanxiang Hao, S. K. Yarlagadda +8
cs.CVeess.IVarXiv:2004.12027v22020Face Anti-Spoofing Via Disentangled Representation Learning
Ke-Yue Zhang, Taiping Yao, Jian Zhang +6
cs.CVarXiv:2008.08250v12020ASR is all you need: cross-modal distillation for lip reading
Triantafyllos Afouras, Joon Son Chung, Andrew Zisserman
cs.CVcs.SDeess.ASarXiv:1911.12747v22019Shape-Aware Organ Segmentation by Predicting Signed Distance Maps
Yuan Xue, Hui Tang, Zhi Qiao +6
cs.CVarXiv:1912.03849v12019Recurrent Convolutional Neural Networks for Scene Parsing
Pedro H. O. Pinheiro, Ronan Collobert
cs.CVarXiv:1306.2795v12013Towards neural networks that provably know when they don't know
Alexander Meinke, Matthias Hein
cs.LGcs.CVstat.MLarXiv:1909.12180v22019SDCNet: Video Prediction Using Spatially-Displaced Convolution
Fitsum A. Reda, Guilin Liu, Kevin J. Shih +5
cs.CVarXiv:1811.00684v22018High Fidelity Video Prediction with Large Stochastic Recurrent Neural Networks
Ruben Villegas, Arkanath Pathak, Harini Kannan +3
cs.CVarXiv:1911.01655v12019Recurrent Attention Models for Depth-Based Person Identification
Albert Haque, Alexandre Alahi, Li Fei-Fei
cs.CVarXiv:1611.07212v12016Weakly-supervised localization of diabetic retinopathy lesions in retinal fundus images
Waleed M. Gondal, Jan M. Köhler, René Grzeszick +2
cs.CVarXiv:1706.09634v12017Context-aware Captions from Context-agnostic Supervision
Ramakrishna Vedantam, Samy Bengio, Kevin Murphy +2
cs.CVcs.AIarXiv:1701.02870v32017FasterViT: Fast Vision Transformers with Hierarchical Attention
Ali Hatamizadeh, Greg Heinrich, Hongxu Yin +4
cs.CVcs.AIcs.LGarXiv:2306.06189v22023Semantics Disentangling for Generalized Zero-Shot Learning
Zhi Chen, Yadan Luo, Ruihong Qiu +4
cs.CVarXiv:2101.07978v52021Underwater Image Enhancement by Transformer-based Diffusion Model with Non-uniform Sampling for Skip Strategy
Yi Tang, Takafumi Iwaguchi, Hiroshi Kawasaki
cs.CVarXiv:2309.03445v12023Active Learning for Deep Object Detection via Probabilistic Modeling
Jiwoong Choi, Ismail Elezi, Hyuk-Jae Lee +2
cs.CVarXiv:2103.16130v22021Face Recognition: Too Bias, or Not Too Bias?
Joseph P Robinson, Gennady Livitz, Yann Henon +3
cs.CVarXiv:2002.06483v42020COAST: COntrollable Arbitrary-Sampling NeTwork for Compressive Sensing
Di You, Jian Zhang, Jingfen Xie +2
cs.CVeess.IVarXiv:2107.07225v12021DAG-Recurrent Neural Networks For Scene Labeling
Bing Shuai, Zhen Zuo, Gang Wang +1
cs.CVarXiv:1509.00552v22015VCT: A Video Compression Transformer
Fabian Mentzer, George Toderici, David Minnen +4
cs.CVcs.LGeess.IVarXiv:2206.07307v22022Backdoor Attack in the Physical World
Yiming Li, Tongqing Zhai, Yong Jiang +2
cs.CRcs.AIcs.CVarXiv:2104.02361v22021Bifurcated backbone strategy for RGB-D salient object detection
Yingjie Zhai, Deng-Ping Fan, Jufeng Yang +4
cs.CVarXiv:2007.02713v32020On Network Design Spaces for Visual Recognition
Ilija Radosavovic, Justin Johnson, Saining Xie +2
cs.CVcs.LGarXiv:1905.13214v12019Spatial Aggregation of Holistically-Nested Networks for Automated Pancreas Segmentation
Holger R. Roth, Le Lu, Amal Farag +2
cs.CVarXiv:1606.07830v12016Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
Yongdong Luo, Xiawu Zheng, Guilin Li +8
cs.CVcs.AIarXiv:2411.13093v42024Faster R-CNN Features for Instance Search
Amaia Salvador, Xavier Giro-i-Nieto, Ferran Marques +1
cs.CVarXiv:1604.08893v12016Few-shot Adaptive Faster R-CNN
Tao Wang, Xiaopeng Zhang, Li Yuan +1
cs.CVarXiv:1903.09372v12019Rethinking Inductive Biases for Surface Normal Estimation
Gwangbin Bae, Andrew J. Davison
cs.CVarXiv:2403.00712v12024Improving Generalization via Scalable Neighborhood Component Analysis
Zhirong Wu, Alexei A. Efros, Stella X. Yu
cs.CVcs.LGarXiv:1808.04699v12018Unsupervised Domain-Specific Deblurring via Disentangled Representations
Boyu Lu, Jun-Cheng Chen, Rama Chellappa
cs.CVarXiv:1903.01594v22019HiFaceGAN: Face Renovation via Collaborative Suppression and Replenishment
Lingbo Yang, Chang Liu, Pan Wang +4
cs.CVarXiv:2005.05005v22020MultiHuSE: A Multimodal Dataset for Humour Styles and Emotions
Mary Ogbuka Kenneth, Foaad Khosmood, Abbas Edalat
cs.CLcs.CVcs.MMarXiv:2609.11322v12026FilterReg: Robust and Efficient Probabilistic Point-Set Registration using Gaussian Filter and Twist Parameterization
Wei Gao, Russ Tedrake
cs.CVarXiv:1811.10136v32018RetinaMask: Learning to predict masks improves state-of-the-art single-shot detection for free
Cheng-Yang Fu, Mykhailo Shvets, Alexander C. Berg
cs.CVarXiv:1901.03353v12019On Recognizing Texts of Arbitrary Shapes with 2D Self-Attention
Junyeop Lee, Sungrae Park, Jeonghun Baek +3
cs.CVarXiv:1910.04396v12019Robust Mean Teacher for Continual and Gradual Test-Time Adaptation
Mario Döbler, Robert A. Marsden, Bin Yang
cs.CVarXiv:2211.13081v22022DiT-3D: Exploring Plain Diffusion Transformers for 3D Shape Generation
Shentong Mo, Enze Xie, Ruihang Chu +4
cs.CVcs.AIcs.LGarXiv:2307.01831v12023Wavelet-Based Dual-Branch Network for Image Demoireing
Lin Liu, Jianzhuang Liu, Shanxin Yuan +4
cs.CVarXiv:2007.07173v22020Structured Bird's-Eye-View Traffic Scene Understanding from Onboard Images
Yigit Baran Can, Alexander Liniger, Danda Pani Paudel +1
cs.CVarXiv:2110.01997v120213DIoUMatch: Leveraging IoU Prediction for Semi-Supervised 3D Object Detection
He Wang, Yezhen Cong, Or Litany +2
cs.CVarXiv:2012.04355v32020