Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
17,341 to 17,400 of 18,867
Deep Facial Expression Recognition: A Survey
Shan Li, Weihong Deng
cs.CVarXiv:1804.08348v22018BiSeNet V2: Bilateral Network with Guided Aggregation for Real-time Semantic Segmentation
Changqian Yu, Changxin Gao, Jingbo Wang +3
cs.CVarXiv:2004.02147v12020DeblurGAN: Blind Motion Deblurring Using Conditional Adversarial Networks
Orest Kupyn, Volodymyr Budzan, Mykola Mykhailych +2
cs.CVarXiv:1711.07064v42017Wild Patterns: Ten Years After the Rise of Adversarial Machine Learning
Battista Biggio, Fabio Roli
cs.CVcs.CRcs.GTarXiv:1712.03141v22017ImageBind: One Embedding Space To Bind Them All
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu +4
cs.CVcs.AIcs.LGarXiv:2305.05665v22023IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
Hu Ye, Jun Zhang, Sibo Liu +2
cs.CVcs.AIarXiv:2308.06721v12023MVSNet: Depth Inference for Unstructured Multi-view Stereo
Yao Yao, Zixin Luo, Shiwei Li +2
cs.CVarXiv:1804.02505v22018Multiscale Vision Transformers
Haoqi Fan, Bo Xiong, Karttikeya Mangalam +4
cs.CVcs.AIcs.LGarXiv:2104.11227v12021Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Zhe Chen, Weiyun Wang, Yue Cao +39
cs.CVarXiv:2412.05271v52024Generative Adversarial Network in Medical Imaging: A Review
Xin Yi, Ekta Walia, Paul Babyn
cs.CVcs.LGarXiv:1809.07294v42018Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training
Yujun Lin, Song Han, Huizi Mao +2
cs.CVcs.DCcs.LGarXiv:1712.01887v32017Coupled Generative Adversarial Networks
Ming-Yu Liu, Oncel Tuzel
cs.CVarXiv:1606.07536v22016Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Bin Lin, Yang Ye, Bin Zhu +4
cs.CVarXiv:2311.10122v32023COCO-Stuff: Thing and Stuff Classes in Context
Holger Caesar, Jasper Uijlings, Vittorio Ferrari
cs.CVarXiv:1612.03716v42016Segmentation Transformer: Object-Contextual Representations for Semantic Segmentation
Yuhui Yuan, Xiaokang Chen, Xilin Chen +1
cs.CVarXiv:1909.11065v62019LinkNet: Exploiting Encoder Representations for Efficient Semantic Segmentation
Abhishek Chaurasia, Eugenio Culurciello
cs.CVcs.LGarXiv:1707.03718v12017Feature Squeezing: Detecting Adversarial Examples in Deep Neural Networks
Weilin Xu, David Evans, Yanjun Qi
cs.CVcs.CRcs.LGarXiv:1704.01155v22017Learning to Adapt Structured Output Space for Semantic Segmentation
Yi-Hsuan Tsai, Wei-Chih Hung, Samuel Schulter +3
cs.CVarXiv:1802.10349v32018Category-Level 3D Correspondence in Camera Space via Morphable Object Priors
Leonhard Sommer, Artur Jesslen, Basavaraj Sunagad +1
cs.CVarXiv:2605.28257v12026Beyond Correlation Filters: Learning Continuous Convolution Operators for Visual Tracking
Martin Danelljan, Andreas Robinson, Fahad Shahbaz Khan +1
cs.CVarXiv:1608.03773v22016iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning
Chang-Bin Zhang, Yujie Zhong, Qiang Zhang +1
cs.CVarXiv:2605.31096v12026How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
Qian Yang, Ankur Sikarwar, Huy Le +4
cs.CVarXiv:2605.27310v12026OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants
Xudong Lu, Xueying Li, Annan Wang +8
cs.CVcs.CLarXiv:2605.26485v12026MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Chaoyou Fu, Peixian Chen, Yunhang Shen +11
cs.CVarXiv:2306.13394v52023Inverting Gradients -- How easy is it to break privacy in federated learning?
Jonas Geiping, Hartmut Bauermeister, Hannah Dröge +1
cs.CVcs.CRcs.LGarXiv:2003.14053v22020QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation
Dmitry Kalashnikov, Alex Irpan, Peter Pastor +8
cs.LGcs.AIcs.CVarXiv:1806.10293v32018ResNeSt: Split-Attention Networks
Hang Zhang, Chongruo Wu, Zhongyue Zhang +9
cs.CVarXiv:2004.08955v22020Panoptic Segmentation
Alexander Kirillov, Kaiming He, Ross Girshick +2
cs.CVarXiv:1801.00868v32018Grounded Language-Image Pre-training
Liunian Harold Li, Pengchuan Zhang, Haotian Zhang +9
cs.CVcs.AIcs.CLarXiv:2112.03857v22021Target-driven Visual Navigation in Indoor Scenes using Deep Reinforcement Learning
Yuke Zhu, Roozbeh Mottaghi, Eric Kolve +4
cs.CVarXiv:1609.05143v12016MemNet: A Persistent Memory Network for Image Restoration
Ying Tai, Jian Yang, Xiaoming Liu +1
cs.CVarXiv:1708.02209v12017Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
Zifeng Wu, Chunhua Shen, Anton van den Hengel
cs.CVarXiv:1611.10080v12016Interpretable Explanations of Black Boxes by Meaningful Perturbation
Ruth Fong, Andrea Vedaldi
cs.CVcs.AIcs.LGarXiv:1704.03296v42017GANomaly: Semi-Supervised Anomaly Detection via Adversarial Training
Samet Akcay, Amir Atapour-Abarghouei, Toby P. Breckon
cs.CVarXiv:1805.06725v32018Hierarchical Question-Image Co-Attention for Visual Question Answering
Jiasen Lu, Jianwei Yang, Dhruv Batra +1
cs.CVcs.CLarXiv:1606.00061v52016SceneFrom3D: Geometry-Conditioned Outdoor 3D Scene Generation via View Scheduling with Object-Level Control
Geonung Kim, Jeongeun Park, Nuri Ryu +2
cs.GRcs.CVarXiv:2607.04540v12026Summaries:한국어ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation
Anindya Mondal, Sauradip Nag, Anjan Dutta
cs.CVeess.IVarXiv:2606.23835v12026Summaries:한국어PhyCo: Learning Controllable Physical Priors for Generative Motion
Sriram Narayanan, Ziyu Jiang, Srinivasa Narasimhan +1
cs.CVcs.AIcs.LGarXiv:2604.28169v12026Summaries:한국어LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
Chunyuan Li, Cliff Wong, Sheng Zhang +6
cs.CVcs.CLarXiv:2306.00890v12023The Cityscapes Dataset for Semantic Urban Scene Understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos +6
cs.CVarXiv:1604.01685v22016Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
Bryan A. Plummer, Liwei Wang, Chris M. Cervantes +3
cs.CVcs.CLarXiv:1505.04870v42015Video Enhancement with Task-Oriented Flow
Tianfan Xue, Baian Chen, Jiajun Wu +2
cs.CVarXiv:1711.09078v32017Large-Scale Evolution of Image Classifiers
Esteban Real, Sherry Moore, Andrew Selle +5
cs.NEcs.AIcs.CVarXiv:1703.01041v22017Generation and Comprehension of Unambiguous Object Descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev +3
cs.CVcs.CLcs.LGarXiv:1511.02283v32015MesoNet: a Compact Facial Video Forgery Detection Network
Darius Afchar, Vincent Nozick, Junichi Yamagishi +1
cs.CVeess.IVarXiv:1809.00888v12018Deep Metric Learning via Lifted Structured Feature Embedding
Hyun Oh Song, Yu Xiang, Stefanie Jegelka +1
cs.CVcs.LGarXiv:1511.06452v12015Objaverse: A Universe of Annotated 3D Objects
Matt Deitke, Dustin Schwenk, Jordi Salvador +7
cs.CVcs.AIcs.GRarXiv:2212.08051v12022Image Captioning with Semantic Attention
Quanzeng You, Hailin Jin, Zhaowen Wang +2
cs.CVarXiv:1603.03925v12016Fully Convolutional Siamese Networks for Change Detection
Rodrigo Caye Daudt, Bertrand Le Saux, Alexandre Boulch
cs.CVcs.LGarXiv:1810.08462v12018NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding
Jun Liu, Amir Shahroudy, Mauricio Perez +3
cs.CVarXiv:1905.04757v22019RMPE: Regional Multi-person Pose Estimation
Hao-Shu Fang, Shuqin Xie, Yu-Wing Tai +1
cs.CVarXiv:1612.00137v52016Memorizing Normality to Detect Anomaly: Memory-augmented Deep Autoencoder for Unsupervised Anomaly Detection
Dong Gong, Lingqiao Liu, Vuong Le +4
cs.CVarXiv:1904.02639v22019TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios
Xingkui Zhu, Shuchang Lyu, Xu Wang +1
cs.CVcs.AIarXiv:2108.11539v12021Kindling the Darkness: A Practical Low-light Image Enhancer
Yonghua Zhang, Jiawan Zhang, Xiaojie Guo
cs.CVarXiv:1905.04161v12019Towards Total Recall in Industrial Anomaly Detection
Karsten Roth, Latha Pemula, Joaquin Zepeda +3
cs.CVarXiv:2106.08265v22021A Survey on Contrastive Self-supervised Learning
Ashish Jaiswal, Ashwin Ramesh Babu, Mohammad Zaki Zadeh +2
cs.CVarXiv:2011.00362v32020Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models
Haozhan Shen, Tiancheng Zhao, Kangjia Zhao +1
cs.CVarXiv:2605.28132v12026Deep Anomaly Detection with Outlier Exposure
Dan Hendrycks, Mantas Mazeika, Thomas Dietterich
cs.LGcs.CLcs.CVarXiv:1812.04606v32018Deeper, Broader and Artier Domain Generalization
Da Li, Yongxin Yang, Yi-Zhe Song +1
cs.CVarXiv:1710.03077v12017LoMo: Local Modality Substitution for Deeper Vision-Language Fusion
Feng Han, Zhixiong Zhang, Zheming Liang +2
cs.CVcs.CLarXiv:2605.30265v12026