Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
2,581 to 2,640 of 18,815
SpectFormer: Frequency and Attention is what you need in a Vision Transformer
Badri N. Patro, Vinay P. Namboodiri, Vijay Srinivas Agneeswaran
cs.CVcs.AIcs.CLarXiv:2304.06446v22023A Survey of Stealth Malware: Attacks, Mitigation Measures, and Steps Toward Autonomous Open World Solutions
Ethan M. Rudd, Andras Rozsa, Manuel Günther +1
cs.CRcs.CVarXiv:1603.06028v22016Thinking Fast and Slow: Efficient Text-to-Visual Retrieval with Transformers
Antoine Miech, Jean-Baptiste Alayrac, Ivan Laptev +2
cs.CVarXiv:2103.16553v12021Feature Pyramid Network for Multi-Class Land Segmentation
Selim S. Seferbekov, Vladimir I. Iglovikov, Alexander V. Buslaev +1
cs.CVarXiv:1806.03510v22018Neural Compatibility Modeling with Attentive Knowledge Distillation
Xuemeng Song, Fuli Feng, Xianjing Han +3
cs.CVcs.MMarXiv:1805.00313v12018Imposing Hard Constraints on Deep Networks: Promises and Limitations
Pablo Márquez-Neila, Mathieu Salzmann, Pascal Fua
cs.CVarXiv:1706.02025v12017Eliciting Self-Verification in Multimodal Reasoning Agents with Reinforcement Learning
Vishwas Sathish, Viresh Ranjan, Xinliang Zhu +2
cs.AIcs.CLcs.CVarXiv:2609.08025v12026Towards the Detection of Diffusion Model Deepfakes
Jonas Ricker, Simon Damm, Thorsten Holz +1
cs.CVarXiv:2210.14571v42022ManipulaTHOR: A Framework for Visual Object Manipulation
Kiana Ehsani, Winson Han, Alvaro Herrasti +5
cs.CVcs.AIcs.LGarXiv:2104.11213v12021RevalExo: A Functional Daily-Activity Benchmark for Inertial and Visual Locomotion Mode Recognition in Older Adults and Clinical Cohorts
Diwas Lamsal, Juha Carlon, Reinhard Claeys +8
cs.AIcs.CVarXiv:2609.08090v12026Unsupervised Change Detection in Multi-temporal VHR Images Based on Deep Kernel PCA Convolutional Mapping Network
Chen Wu, Hongruixuan Chen, Bo Do +1
eess.IVcs.CVarXiv:1912.08628v12019Review of Deep Learning
Rong Zhang, Weiping Li, Tong Mo
cs.LGcs.CVcs.NEarXiv:1804.01653v22018Diffusion-SDF: Text-to-Shape via Voxelized Diffusion
Muheng Li, Yueqi Duan, Jie Zhou +1
cs.CVcs.AIcs.GRarXiv:2212.03293v22022Counterfactual Critic Multi-Agent Training for Scene Graph Generation
Long Chen, Hanwang Zhang, Jun Xiao +3
cs.CVarXiv:1812.02347v32018Interpretable and Accurate Fine-grained Recognition via Region Grouping
Zixuan Huang, Yin Li
cs.CVcs.AIcs.LGarXiv:2005.10411v12020BEVBert: Multimodal Map Pre-training for Language-guided Navigation
Dong An, Yuankai Qi, Yangguang Li +4
cs.CVcs.AIcs.CLarXiv:2212.04385v22022What is a salient object? A dataset and a baseline model for salient object detection
Ali Borji
cs.CVarXiv:1412.5027v12014Uncovering convolutional neural network decisions for diagnosing multiple sclerosis on conventional MRI using layer-wise relevance propagation
Fabian Eitel, Emily Soehler, Judith Bellmann-Strobl +10
cs.CVarXiv:1904.08771v12019OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
Raghav Kapoor, Yash Parag Butala, Melisa Russak +4
cs.AIcs.CLcs.CVarXiv:2402.17553v32024Towards Transferable Adversarial Attacks on Vision Transformers
Zhipeng Wei, Jingjing Chen, Micah Goldblum +3
cs.CVcs.AIarXiv:2109.04176v32021SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation
Soroush Mehraban, Xin Lei Lin, Vida Adeli +5
cs.CVarXiv:2609.08108v12026Incomplete Contrastive Multi-View Clustering with High-Confidence Guiding
Guoqing Chao, Yi Jiang, Dianhui Chu
cs.CVcs.LGarXiv:2312.08697v12023FastDeRain: A Novel Video Rain Streak Removal Method Using Directional Gradient Priors
Tai-Xiang Jiang, Ting-Zhu Huang, Xi-Le Zhao +2
cs.CVarXiv:1803.07487v32018Learning Multi-level Deep Representations for Image Emotion Classification
Tianrong Rao, Min Xu, Dong Xu
cs.CVarXiv:1611.07145v22016Generalized Jensen-Shannon Divergence Loss for Learning with Noisy Labels
Erik Englesson, Hossein Azizpour
cs.LGcs.CVstat.MLarXiv:2105.04522v42021Language Conditioned Spatial Relation Reasoning for 3D Object Grounding
Shizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi +2
cs.CVarXiv:2211.09646v12022LightenDiffusion: Unsupervised Low-Light Image Enhancement with Latent-Retinex Diffusion Models
Hai Jiang, Ao Luo, Xiaohong Liu +2
cs.CVarXiv:2407.08939v12024An Implementation of Faster RCNN with Study for Region Sampling
Xinlei Chen, Abhinav Gupta
cs.CVarXiv:1702.02138v22017Localization in the Crowd with Topological Constraints
Shahira Abousamra, Minh Hoai, Dimitris Samaras +1
cs.CVarXiv:2012.12482v12020Lightweight Salient Object Detection in Optical Remote-Sensing Images via Semantic Matching and Edge Alignment
Gongyang Li, Zhi Liu, Xinpeng Zhang +1
cs.CVarXiv:2301.02778v22023OpenEarthMap: A Benchmark Dataset for Global High-Resolution Land Cover Mapping
Junshi Xia, Naoto Yokoya, Bruno Adriano +1
cs.CVcs.LGarXiv:2210.10732v12022PatchNet: A Simple Face Anti-Spoofing Framework via Fine-Grained Patch Recognition
Chien-Yi Wang, Yu-Ding Lu, Shang-Ta Yang +1
cs.CVarXiv:2203.14325v12022SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
Baoxiong Jia, Yixin Chen, Huangyue Yu +5
cs.CVcs.AIcs.CLarXiv:2401.09340v32024Sparse Adversarial Perturbations for Videos
Xingxing Wei, Jun Zhu, Hang Su
cs.CVarXiv:1803.02536v12018Emu: Generative Pretraining in Multimodality
Quan Sun, Qiying Yu, Yufeng Cui +7
cs.CVarXiv:2307.05222v22023It's Moving! A Probabilistic Model for Causal Motion Segmentation in Moving Camera Videos
Pia Bideau, Erik Learned-Miller
cs.CVarXiv:1604.00136v12016EDCNN: Edge enhancement-based Densely Connected Network with Compound Loss for Low-Dose CT Denoising
Tengfei Liang, Yi Jin, Yidong Li +3
eess.IVcs.CVarXiv:2011.00139v12020Generalizing Dataset Distillation via Deep Generative Prior
George Cazenavette, Tongzhou Wang, Antonio Torralba +2
cs.CVcs.AIcs.LGarXiv:2305.01649v22023MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
Sheng-Chieh Lin, Chankyu Lee, Mohammad Shoeybi +3
cs.CLcs.AIcs.CVarXiv:2411.02571v22024Learning for Video Compression with Recurrent Auto-Encoder and Recurrent Probability Model
Ren Yang, Fabian Mentzer, Luc Van Gool +1
eess.IVcs.CVarXiv:2006.13560v42020VideoDex: Learning Dexterity from Internet Videos
Kenneth Shaw, Shikhar Bahl, Deepak Pathak
cs.ROcs.AIcs.CVarXiv:2212.04498v12022MEGANet: Multi-Scale Edge-Guided Attention Network for Weak Boundary Polyp Segmentation
Nhat-Tan Bui, Dinh-Hieu Hoang, Quang-Thuc Nguyen +2
cs.CVarXiv:2309.03329v32023Diffuse, Attend, and Segment: Unsupervised Zero-Shot Segmentation using Stable Diffusion
Junjiao Tian, Lavisha Aggarwal, Andrea Colaco +2
cs.CVarXiv:2308.12469v32023Labelling unlabelled videos from scratch with multi-modal self-supervision
Yuki M. Asano, Mandela Patrick, Christian Rupprecht +1
cs.CVcs.LGarXiv:2006.13662v32020Mean-Shifted Contrastive Loss for Anomaly Detection
Tal Reiss, Yedid Hoshen
cs.CVcs.LGarXiv:2106.03844v22021Inverse Compositional Spatial Transformer Networks
Chen-Hsuan Lin, Simon Lucey
cs.CVcs.LGarXiv:1612.03897v12016Dynamic-structured Semantic Propagation Network
Xiaodan Liang, Hongfei Zhou, Eric Xing
cs.CVarXiv:1803.06067v12018On GANs and GMMs
Eitan Richardson, Yair Weiss
cs.CVcs.LGarXiv:1805.12462v22018Multiple Video Frame Interpolation via Enhanced Deformable Separable Convolution
Xianhang Cheng, Zhenzhong Chen
cs.CVeess.IVarXiv:2006.08070v22020MobileStereoNet: Towards Lightweight Deep Networks for Stereo Matching
Faranak Shamsafar, Samuel Woerz, Rafia Rahim +1
cs.CVarXiv:2108.09770v12021FOIL it! Find One mismatch between Image and Language caption
Ravi Shekhar, Sandro Pezzelle, Yauhen Klimovich +4
cs.CVcs.CLcs.MMarXiv:1705.01359v12017MDU-Net: Multi-scale Densely Connected U-Net for biomedical image segmentation
Jiawei Zhang, Yuzhen Jin, Jilan Xu +2
cs.CVarXiv:1812.00352v32018A survey on computational spectral reconstruction methods from RGB to hyperspectral imaging
Jingang Zhang, Runmu Su, Wenqi Ren +3
eess.IVcs.CVarXiv:2106.15944v22021Active Domain Adaptation via Clustering Uncertainty-weighted Embeddings
Viraj Prabhu, Arjun Chandrasekaran, Kate Saenko +1
cs.CVcs.LGarXiv:2010.08666v32020Just Go with the Flow: Self-Supervised Scene Flow Estimation
Himangi Mittal, Brian Okorn, David Held
cs.CVcs.LGcs.ROarXiv:1912.00497v22019DVDnet: A Fast Network for Deep Video Denoising
Matias Tassano, Julie Delon, Thomas Veit
eess.IVcs.CVarXiv:1906.11890v12019Medical Image Registration Using Deep Neural Networks: A Comprehensive Review
Hamid Reza Boveiri, Raouf Khayami, Reza Javidan +1
eess.IVcs.CVcs.LGarXiv:2002.03401v12020Early-detection and classification of live bacteria using time-lapse coherent imaging and deep learning
Hongda Wang, Hatice Ceylan Koydemir, Yunzhe Qiu +8
physics.ins-detcs.CVphysics.app-pharXiv:2001.10695v12020Dense Depth Estimation in Monocular Endoscopy with Self-supervised Learning Methods
Xingtong Liu, Ayushi Sinha, Masaru Ishii +4
cs.CVstat.MLarXiv:1902.07766v22019BppAttack: Stealthy and Efficient Trojan Attacks against Deep Neural Networks via Image Quantization and Contrastive Adversarial Learning
Zhenting Wang, Juan Zhai, Shiqing Ma
cs.CVcs.CRcs.LGarXiv:2205.13383v12022