Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

2,581 to 2,640 of 18,815

  1. SpectFormer: Frequency and Attention is what you need in a Vision Transformer

    Badri N. Patro, Vinay P. Namboodiri, Vijay Srinivas Agneeswaran

    cs.CVcs.AIcs.CLarXiv:2304.06446v22023
  2. A Survey of Stealth Malware: Attacks, Mitigation Measures, and Steps Toward Autonomous Open World Solutions

    Ethan M. Rudd, Andras Rozsa, Manuel Günther +1

    cs.CRcs.CVarXiv:1603.06028v22016
  3. Thinking Fast and Slow: Efficient Text-to-Visual Retrieval with Transformers

    Antoine Miech, Jean-Baptiste Alayrac, Ivan Laptev +2

    cs.CVarXiv:2103.16553v12021
  4. Feature Pyramid Network for Multi-Class Land Segmentation

    Selim S. Seferbekov, Vladimir I. Iglovikov, Alexander V. Buslaev +1

    cs.CVarXiv:1806.03510v22018
  5. Neural Compatibility Modeling with Attentive Knowledge Distillation

    Xuemeng Song, Fuli Feng, Xianjing Han +3

    cs.CVcs.MMarXiv:1805.00313v12018
  6. Imposing Hard Constraints on Deep Networks: Promises and Limitations

    Pablo Márquez-Neila, Mathieu Salzmann, Pascal Fua

    cs.CVarXiv:1706.02025v12017
  7. Eliciting Self-Verification in Multimodal Reasoning Agents with Reinforcement Learning

    Vishwas Sathish, Viresh Ranjan, Xinliang Zhu +2

    cs.AIcs.CLcs.CVarXiv:2609.08025v12026
  8. Towards the Detection of Diffusion Model Deepfakes

    Jonas Ricker, Simon Damm, Thorsten Holz +1

    cs.CVarXiv:2210.14571v42022
  9. ManipulaTHOR: A Framework for Visual Object Manipulation

    Kiana Ehsani, Winson Han, Alvaro Herrasti +5

    cs.CVcs.AIcs.LGarXiv:2104.11213v12021
  10. RevalExo: A Functional Daily-Activity Benchmark for Inertial and Visual Locomotion Mode Recognition in Older Adults and Clinical Cohorts

    Diwas Lamsal, Juha Carlon, Reinhard Claeys +8

    cs.AIcs.CVarXiv:2609.08090v12026
  11. Unsupervised Change Detection in Multi-temporal VHR Images Based on Deep Kernel PCA Convolutional Mapping Network

    Chen Wu, Hongruixuan Chen, Bo Do +1

    eess.IVcs.CVarXiv:1912.08628v12019
  12. Review of Deep Learning

    Rong Zhang, Weiping Li, Tong Mo

    cs.LGcs.CVcs.NEarXiv:1804.01653v22018
  13. Diffusion-SDF: Text-to-Shape via Voxelized Diffusion

    Muheng Li, Yueqi Duan, Jie Zhou +1

    cs.CVcs.AIcs.GRarXiv:2212.03293v22022
  14. Counterfactual Critic Multi-Agent Training for Scene Graph Generation

    Long Chen, Hanwang Zhang, Jun Xiao +3

    cs.CVarXiv:1812.02347v32018
  15. Interpretable and Accurate Fine-grained Recognition via Region Grouping

    Zixuan Huang, Yin Li

    cs.CVcs.AIcs.LGarXiv:2005.10411v12020
  16. BEVBert: Multimodal Map Pre-training for Language-guided Navigation

    Dong An, Yuankai Qi, Yangguang Li +4

    cs.CVcs.AIcs.CLarXiv:2212.04385v22022
  17. What is a salient object? A dataset and a baseline model for salient object detection

    Ali Borji

    cs.CVarXiv:1412.5027v12014
  18. Uncovering convolutional neural network decisions for diagnosing multiple sclerosis on conventional MRI using layer-wise relevance propagation

    Fabian Eitel, Emily Soehler, Judith Bellmann-Strobl +10

    cs.CVarXiv:1904.08771v12019
  19. OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web

    Raghav Kapoor, Yash Parag Butala, Melisa Russak +4

    cs.AIcs.CLcs.CVarXiv:2402.17553v32024
  20. Towards Transferable Adversarial Attacks on Vision Transformers

    Zhipeng Wei, Jingjing Chen, Micah Goldblum +3

    cs.CVcs.AIarXiv:2109.04176v32021
  21. SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation

    Soroush Mehraban, Xin Lei Lin, Vida Adeli +5

    cs.CVarXiv:2609.08108v12026
  22. Incomplete Contrastive Multi-View Clustering with High-Confidence Guiding

    Guoqing Chao, Yi Jiang, Dianhui Chu

    cs.CVcs.LGarXiv:2312.08697v12023
  23. FastDeRain: A Novel Video Rain Streak Removal Method Using Directional Gradient Priors

    Tai-Xiang Jiang, Ting-Zhu Huang, Xi-Le Zhao +2

    cs.CVarXiv:1803.07487v32018
  24. Learning Multi-level Deep Representations for Image Emotion Classification

    Tianrong Rao, Min Xu, Dong Xu

    cs.CVarXiv:1611.07145v22016
  25. Generalized Jensen-Shannon Divergence Loss for Learning with Noisy Labels

    Erik Englesson, Hossein Azizpour

    cs.LGcs.CVstat.MLarXiv:2105.04522v42021
  26. Language Conditioned Spatial Relation Reasoning for 3D Object Grounding

    Shizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi +2

    cs.CVarXiv:2211.09646v12022
  27. LightenDiffusion: Unsupervised Low-Light Image Enhancement with Latent-Retinex Diffusion Models

    Hai Jiang, Ao Luo, Xiaohong Liu +2

    cs.CVarXiv:2407.08939v12024
  28. An Implementation of Faster RCNN with Study for Region Sampling

    Xinlei Chen, Abhinav Gupta

    cs.CVarXiv:1702.02138v22017
  29. Localization in the Crowd with Topological Constraints

    Shahira Abousamra, Minh Hoai, Dimitris Samaras +1

    cs.CVarXiv:2012.12482v12020
  30. Lightweight Salient Object Detection in Optical Remote-Sensing Images via Semantic Matching and Edge Alignment

    Gongyang Li, Zhi Liu, Xinpeng Zhang +1

    cs.CVarXiv:2301.02778v22023
  31. OpenEarthMap: A Benchmark Dataset for Global High-Resolution Land Cover Mapping

    Junshi Xia, Naoto Yokoya, Bruno Adriano +1

    cs.CVcs.LGarXiv:2210.10732v12022
  32. PatchNet: A Simple Face Anti-Spoofing Framework via Fine-Grained Patch Recognition

    Chien-Yi Wang, Yu-Ding Lu, Shang-Ta Yang +1

    cs.CVarXiv:2203.14325v12022
  33. SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding

    Baoxiong Jia, Yixin Chen, Huangyue Yu +5

    cs.CVcs.AIcs.CLarXiv:2401.09340v32024
  34. Sparse Adversarial Perturbations for Videos

    Xingxing Wei, Jun Zhu, Hang Su

    cs.CVarXiv:1803.02536v12018
  35. Emu: Generative Pretraining in Multimodality

    Quan Sun, Qiying Yu, Yufeng Cui +7

    cs.CVarXiv:2307.05222v22023
  36. It's Moving! A Probabilistic Model for Causal Motion Segmentation in Moving Camera Videos

    Pia Bideau, Erik Learned-Miller

    cs.CVarXiv:1604.00136v12016
  37. EDCNN: Edge enhancement-based Densely Connected Network with Compound Loss for Low-Dose CT Denoising

    Tengfei Liang, Yi Jin, Yidong Li +3

    eess.IVcs.CVarXiv:2011.00139v12020
  38. Generalizing Dataset Distillation via Deep Generative Prior

    George Cazenavette, Tongzhou Wang, Antonio Torralba +2

    cs.CVcs.AIcs.LGarXiv:2305.01649v22023
  39. MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs

    Sheng-Chieh Lin, Chankyu Lee, Mohammad Shoeybi +3

    cs.CLcs.AIcs.CVarXiv:2411.02571v22024
  40. Learning for Video Compression with Recurrent Auto-Encoder and Recurrent Probability Model

    Ren Yang, Fabian Mentzer, Luc Van Gool +1

    eess.IVcs.CVarXiv:2006.13560v42020
  41. VideoDex: Learning Dexterity from Internet Videos

    Kenneth Shaw, Shikhar Bahl, Deepak Pathak

    cs.ROcs.AIcs.CVarXiv:2212.04498v12022
  42. MEGANet: Multi-Scale Edge-Guided Attention Network for Weak Boundary Polyp Segmentation

    Nhat-Tan Bui, Dinh-Hieu Hoang, Quang-Thuc Nguyen +2

    cs.CVarXiv:2309.03329v32023
  43. Diffuse, Attend, and Segment: Unsupervised Zero-Shot Segmentation using Stable Diffusion

    Junjiao Tian, Lavisha Aggarwal, Andrea Colaco +2

    cs.CVarXiv:2308.12469v32023
  44. Labelling unlabelled videos from scratch with multi-modal self-supervision

    Yuki M. Asano, Mandela Patrick, Christian Rupprecht +1

    cs.CVcs.LGarXiv:2006.13662v32020
  45. Mean-Shifted Contrastive Loss for Anomaly Detection

    Tal Reiss, Yedid Hoshen

    cs.CVcs.LGarXiv:2106.03844v22021
  46. Inverse Compositional Spatial Transformer Networks

    Chen-Hsuan Lin, Simon Lucey

    cs.CVcs.LGarXiv:1612.03897v12016
  47. Dynamic-structured Semantic Propagation Network

    Xiaodan Liang, Hongfei Zhou, Eric Xing

    cs.CVarXiv:1803.06067v12018
  48. On GANs and GMMs

    Eitan Richardson, Yair Weiss

    cs.CVcs.LGarXiv:1805.12462v22018
  49. Multiple Video Frame Interpolation via Enhanced Deformable Separable Convolution

    Xianhang Cheng, Zhenzhong Chen

    cs.CVeess.IVarXiv:2006.08070v22020
  50. MobileStereoNet: Towards Lightweight Deep Networks for Stereo Matching

    Faranak Shamsafar, Samuel Woerz, Rafia Rahim +1

    cs.CVarXiv:2108.09770v12021
  51. FOIL it! Find One mismatch between Image and Language caption

    Ravi Shekhar, Sandro Pezzelle, Yauhen Klimovich +4

    cs.CVcs.CLcs.MMarXiv:1705.01359v12017
  52. MDU-Net: Multi-scale Densely Connected U-Net for biomedical image segmentation

    Jiawei Zhang, Yuzhen Jin, Jilan Xu +2

    cs.CVarXiv:1812.00352v32018
  53. A survey on computational spectral reconstruction methods from RGB to hyperspectral imaging

    Jingang Zhang, Runmu Su, Wenqi Ren +3

    eess.IVcs.CVarXiv:2106.15944v22021
  54. Active Domain Adaptation via Clustering Uncertainty-weighted Embeddings

    Viraj Prabhu, Arjun Chandrasekaran, Kate Saenko +1

    cs.CVcs.LGarXiv:2010.08666v32020
  55. Just Go with the Flow: Self-Supervised Scene Flow Estimation

    Himangi Mittal, Brian Okorn, David Held

    cs.CVcs.LGcs.ROarXiv:1912.00497v22019
  56. DVDnet: A Fast Network for Deep Video Denoising

    Matias Tassano, Julie Delon, Thomas Veit

    eess.IVcs.CVarXiv:1906.11890v12019
  57. Medical Image Registration Using Deep Neural Networks: A Comprehensive Review

    Hamid Reza Boveiri, Raouf Khayami, Reza Javidan +1

    eess.IVcs.CVcs.LGarXiv:2002.03401v12020
  58. Early-detection and classification of live bacteria using time-lapse coherent imaging and deep learning

    Hongda Wang, Hatice Ceylan Koydemir, Yunzhe Qiu +8

    physics.ins-detcs.CVphysics.app-pharXiv:2001.10695v12020
  59. Dense Depth Estimation in Monocular Endoscopy with Self-supervised Learning Methods

    Xingtong Liu, Ayushi Sinha, Masaru Ishii +4

    cs.CVstat.MLarXiv:1902.07766v22019
  60. BppAttack: Stealthy and Efficient Trojan Attacks against Deep Neural Networks via Image Quantization and Contrastive Adversarial Learning

    Zhenting Wang, Juan Zhai, Shiqing Ma

    cs.CVcs.CRcs.LGarXiv:2205.13383v12022