Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
12,961 to 13,020 of 18,817
Identification and Recognition of Rice Diseases and Pests Using Convolutional Neural Networks
Chowdhury Rafeed Rahman, Preetom Saha Arko, Mohammed Eunus Ali +4
cs.CVarXiv:1812.01043v32018FEELVOS: Fast End-to-End Embedding Learning for Video Object Segmentation
Paul Voigtlaender, Yuning Chai, Florian Schroff +3
cs.CVarXiv:1902.09513v22019Graph-Structured Representations for Visual Question Answering
Damien Teney, Lingqiao Liu, Anton van den Hengel
cs.CVcs.AIcs.CLarXiv:1609.05600v22016Tensor Ring Decomposition
Qibin Zhao, Guoxu Zhou, Shengli Xie +2
math.NAcs.CVcs.DSarXiv:1606.05535v12016Neural Voice Puppetry: Audio-driven Facial Reenactment
Justus Thies, Mohamed Elgharib, Ayush Tewari +2
cs.CVcs.GRarXiv:1912.05566v22019DenseNet: Implementing Efficient ConvNet Descriptor Pyramids
Forrest Iandola, Matt Moskewicz, Sergey Karayev +3
cs.CVarXiv:1404.1869v12014FickleNet: Weakly and Semi-supervised Semantic Image Segmentation using Stochastic Inference
Jungbeom Lee, Eunji Kim, Sungmin Lee +2
cs.CVarXiv:1902.10421v22019Image Segmentation for Fruit Detection and Yield Estimation in Apple Orchards
Suchet Bargoti, James Underwood
cs.ROcs.CVcs.LGarXiv:1610.08120v12016Convolutional Feature Masking for Joint Object and Stuff Segmentation
Jifeng Dai, Kaiming He, Jian Sun
cs.CVarXiv:1412.1283v42014BodyNet: Volumetric Inference of 3D Human Body Shapes
Gül Varol, Duygu Ceylan, Bryan Russell +4
cs.CVarXiv:1804.04875v32018nnU-Net for Brain Tumor Segmentation
Fabian Isensee, Paul F. Jaeger, Peter M. Full +2
eess.IVcs.CVarXiv:2011.00848v12020RUBi: Reducing Unimodal Biases in Visual Question Answering
Remi Cadene, Corentin Dancette, Hedi Ben-younes +2
cs.CVcs.CLcs.LGarXiv:1906.10169v22019CrossFuse: A Novel Cross Attention Mechanism based Infrared and Visible Image Fusion Approach
Hui Li, Xiao-Jun Wu
cs.CVarXiv:2406.10581v12024FineGym: A Hierarchical Video Dataset for Fine-grained Action Understanding
Dian Shao, Yue Zhao, Bo Dai +1
cs.CVarXiv:2004.06704v12020Learning Monocular Reactive UAV Control in Cluttered Natural Environments
Stephane Ross, Narek Melik-Barkhudarov, Kumar Shaurya Shankar +4
cs.ROcs.CVcs.LGarXiv:1211.1690v12012Multi-view Low-rank Sparse Subspace Clustering
Maria Brbic, Ivica Kopriva
cs.CVcs.LGmath.OCarXiv:1708.08732v12017Urban Change Detection for Multispectral Earth Observation Using Convolutional Neural Networks
Rodrigo Caye Daudt, Bertrand Le Saux, Alexandre Boulch +1
cs.CVcs.LGarXiv:1810.08468v12018Boosting Few-Shot Visual Learning with Self-Supervision
Spyros Gidaris, Andrei Bursuc, Nikos Komodakis +2
cs.CVcs.LGarXiv:1906.05186v12019Learning Steerable Filters for Rotation Equivariant CNNs
Maurice Weiler, Fred A. Hamprecht, Martin Storath
cs.LGcs.CVarXiv:1711.07289v32017Cold Diffusion: Inverting Arbitrary Image Transforms Without Noise
Arpit Bansal, Eitan Borgnia, Hong-Min Chu +6
cs.CVcs.LGarXiv:2208.09392v12022Summaries:한국어Rethinking and Improving Relative Position Encoding for Vision Transformer
Kan Wu, Houwen Peng, Minghao Chen +2
cs.CVarXiv:2107.14222v12021AutoZOOM: Autoencoder-based Zeroth Order Optimization Method for Attacking Black-box Neural Networks
Chun-Chen Tu, Paishun Ting, Pin-Yu Chen +5
cs.CVcs.CRstat.MLarXiv:1805.11770v52018A Fast and Accurate One-Stage Approach to Visual Grounding
Zhengyuan Yang, Boqing Gong, Liwei Wang +3
cs.CVarXiv:1908.06354v12019No-Reference Image Quality Assessment via Transformers, Relative Ranking, and Self-Consistency
S. Alireza Golestaneh, Saba Dadsetan, Kris M. Kitani
eess.IVcs.CVarXiv:2108.06858v22021SurfaceNet: An End-to-end 3D Neural Network for Multiview Stereopsis
Mengqi Ji, Juergen Gall, Haitian Zheng +2
cs.CVarXiv:1708.01749v12017Rethinking Lossy Compression: The Rate-Distortion-Perception Tradeoff
Yochai Blau, Tomer Michaeli
cs.LGcs.CVcs.ITarXiv:1901.07821v42019Importance Weighted Adversarial Nets for Partial Domain Adaptation
Jing Zhang, Zewei Ding, Wanqing Li +1
cs.CVarXiv:1803.09210v22018Fast Face-swap Using Convolutional Neural Networks
Iryna Korshunova, Wenzhe Shi, Joni Dambre +1
cs.CVarXiv:1611.09577v22016From Patches to Pictures (PaQ-2-PiQ): Mapping the Perceptual Space of Picture Quality
Zhenqiang Ying, Haoran Niu, Praful Gupta +3
cs.CVcs.MMeess.IVarXiv:1912.10088v12019Reason in the Words You Speak: Idiolectal Paraphrasing Off-Policy Traces for Reasoning Distillation in VideoLLMs
Ji Soo Lee, Jinyoung Park, Seohyun Lee +4
cs.CVarXiv:2608.26684v12026HUG-VIS: A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence
Fei Ma, Zebang Cheng, Minghui Li +11
cs.CVarXiv:2608.26517v12026A General Optimization-based Framework for Global Pose Estimation with Multiple Sensors
Tong Qin, Shaozu Cao, Jie Pan +1
cs.CVarXiv:1901.03642v12019GenImage: A Million-Scale Benchmark for Detecting AI-Generated Image
Mingjian Zhu, Hanting Chen, Qiangyu Yan +7
cs.CVarXiv:2306.08571v22023Learning Woody Clearing With Loss Alignment for Zero-Shot Regrowth and Woody Segmentation
Kal Backman, Jared Wood, Adam Roff
cs.CVarXiv:2608.26489v12026Hyperspectral Image Denoising Employing a Spatial-Spectral Deep Residual Convolutional Neural Network
Qiangqiang Yuan, Qiang Zhang, Jie Li +2
cs.CVarXiv:1806.00183v32018Text2Mesh: Text-Driven Neural Stylization for Meshes
Oscar Michel, Roi Bar-On, Richard Liu +2
cs.CVcs.CLcs.GRarXiv:2112.03221v12021From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation
Haowen Gu, Gensheng Pei, Junzhu Mao +3
cs.CVcs.AIarXiv:2608.26856v12026Summaries:简体中文DiffIR: Efficient Diffusion Model for Image Restoration
Bin Xia, Yulun Zhang, Shiyin Wang +5
cs.CVarXiv:2303.09472v32023A survey of the Vision Transformers and their CNN-Transformer based Variants
Asifullah Khan, Zunaira Rauf, Anabia Sohail +4
cs.CVarXiv:2305.09880v42023RECAP-Forcing: Retaining Content Appearances for Long Video Generation
Haiyang Xu, Zheng Ding, Zhuowen Tu
cs.CVarXiv:2608.26671v12026MobileNeRF: Exploiting the Polygon Rasterization Pipeline for Efficient Neural Field Rendering on Mobile Architectures
Zhiqin Chen, Thomas Funkhouser, Peter Hedman +1
cs.CVcs.GRcs.LGarXiv:2208.00277v52022Disentangled Person Image Generation
Liqian Ma, Qianru Sun, Stamatios Georgoulis +3
cs.CVarXiv:1712.02621v42017FaceForensics: A Large-scale Video Dataset for Forgery Detection in Human Faces
Andreas Rössler, Davide Cozzolino, Luisa Verdoliva +3
cs.CVarXiv:1803.09179v12018MedMNIST Classification Decathlon: A Lightweight AutoML Benchmark for Medical Image Analysis
Jiancheng Yang, Rui Shi, Bingbing Ni
cs.CVcs.AIcs.LGarXiv:2010.14925v42020PhySG: Inverse Rendering with Spherical Gaussians for Physics-based Material Editing and Relighting
Kai Zhang, Fujun Luan, Qianqian Wang +2
cs.CVcs.GRarXiv:2104.00674v12021Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object Detection
Shihao Wang, Yingfei Liu, Tiancai Wang +2
cs.CVarXiv:2303.11926v22023GRNet: Gridding Residual Network for Dense Point Cloud Completion
Haozhe Xie, Hongxun Yao, Shangchen Zhou +3
cs.CVcs.LGeess.IVarXiv:2006.03761v42020COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis
Yansong Tang, Dajun Ding, Yongming Rao +5
cs.CVarXiv:1903.02874v12019CGS-SLAM: Collaborative Gaussian Splatting based SLAM for Multi-Agent Reconstruction
Jean-Daniel de Ambrogi, Aladine Chetouani, Vincent Nguyen +1
cs.CVcs.ROarXiv:2608.26868v12026Plug-and-Play Methods Provably Converge with Properly Trained Denoisers
Ernest K. Ryu, Jialin Liu, Sicheng Wang +3
cs.CVeess.IVarXiv:1905.05406v12019G2D: Generative-to-Discriminative Collaborative Inference for Zero-Shot Image Classification
Zehua Hao, Fang Liu, Qinliang Wang +3
cs.CVarXiv:2608.26744v12026Analyzing and Improving the Training Dynamics of Diffusion Models
Tero Karras, Miika Aittala, Jaakko Lehtinen +3
cs.CVcs.AIcs.LGarXiv:2312.02696v22023Transfer Learning from Deep Features for Remote Sensing and Poverty Mapping
Michael Xie, Neal Jean, Marshall Burke +2
cs.CVcs.CYarXiv:1510.00098v22015STDP-based spiking deep convolutional neural networks for object recognition
Saeed Reza Kheradpisheh, Mohammad Ganjtabesh, Simon J Thorpe +1
cs.CVarXiv:1611.01421v32016PhysDiff: Physics-Guided Human Motion Diffusion Model
Ye Yuan, Jiaming Song, Umar Iqbal +2
cs.CVcs.AIcs.GRarXiv:2212.02500v32022SMIL: Multimodal Learning with Severely Missing Modality
Mengmeng Ma, Jian Ren, Long Zhao +3
cs.CVarXiv:2103.05677v12021PandaGPT: One Model To Instruction-Follow Them All
Yixuan Su, Tian Lan, Huayang Li +3
cs.CLcs.CVarXiv:2305.16355v12023Hyperspectral Diffusion Equivariant Imaging (HyDiff-EI): A Self-supervised Framework for Hyperspectral Image Inpainting
Shuo Li, Mike Davies, Mehrdad Yaghoobi
cs.CVcs.LGarXiv:2608.26812v12026Deep Convolutional Ranking for Multilabel Image Annotation
Yunchao Gong, Yangqing Jia, Thomas Leung +2
cs.CVarXiv:1312.4894v22013PyHST2: an hybrid distributed code for high speed tomographic reconstruction with iterative reconstruction and a priori knowledge capabilities
Alessandro Mirone, Emmanuelle Gouillart, Emmanuel Brun +2
math.NAcs.CVarXiv:1306.1392v12013