Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

11,581 to 11,640 of 18,971

  1. Unsupervised Point Cloud Pre-Training via Occlusion Completion

    Hanchen Wang, Qi Liu, Xiangyu Yue +2

    cs.CVcs.LGarXiv:2010.01089v32020
  2. iCAN: Instance-Centric Attention Network for Human-Object Interaction Detection

    Chen Gao, Yuliang Zou, Jia-Bin Huang

    cs.CVarXiv:1808.10437v12018
  3. SCSA: Exploring the Synergistic Effects Between Spatial and Channel Attention

    Yunzhong Si, Huiying Xu, Xinzhong Zhu +4

    cs.CVarXiv:2407.05128v22024
  4. Attentional Pooling for Action Recognition

    Rohit Girdhar, Deva Ramanan

    cs.CVarXiv:1711.01467v32017
  5. LSDA: Large Scale Detection Through Adaptation

    Judy Hoffman, Sergio Guadarrama, Eric Tzeng +5

    cs.CVarXiv:1407.5035v32014
  6. Advances in adversarial attacks and defenses in computer vision: A survey

    Naveed Akhtar, Ajmal Mian, Navid Kardan +1

    cs.CVcs.CRcs.CYarXiv:2108.00401v22021
  7. DeepFake Detection by Analyzing Convolutional Traces

    Luca Guarnera, Oliver Giudice, Sebastiano Battiato

    cs.CVarXiv:2004.10448v12020
  8. Panoptic Neural Fields: A Semantic Object-Aware Neural Scene Representation

    Abhijit Kundu, Kyle Genova, Xiaoqi Yin +6

    cs.CVarXiv:2205.04334v12022
  9. Harmonizing Transferability and Discriminability for Adapting Object Detectors

    Chaoqi Chen, Zebiao Zheng, Xinghao Ding +2

    cs.CVarXiv:2003.06297v12020
  10. Texture Fields: Learning Texture Representations in Function Space

    Michael Oechsle, Lars Mescheder, Michael Niemeyer +2

    cs.CVarXiv:1905.07259v12019
  11. Simple Baseline for Visual Question Answering

    Bolei Zhou, Yuandong Tian, Sainbayar Sukhbaatar +2

    cs.CVcs.CLarXiv:1512.02167v22015
  12. Combining Self-Supervised Learning and Imitation for Vision-Based Rope Manipulation

    Ashvin Nair, Dian Chen, Pulkit Agrawal +4

    cs.CVcs.LGcs.ROarXiv:1703.02018v12017
  13. Towards Automatic Wild Animal Monitoring: Identification of Animal Species in Camera-trap Images using Very Deep Convolutional Neural Networks

    Alexander Gomez, Augusto Salazar, Francisco Vargas

    cs.CVarXiv:1603.06169v22016
  14. Paris-Lille-3D: a large and high-quality ground truth urban point cloud dataset for automatic segmentation and classification

    Xavier Roynard, Jean-Emmanuel Deschaud, François Goulette

    cs.LGcs.CVstat.MLarXiv:1712.00032v22017
  15. Unsupervised Learning of Probably Symmetric Deformable 3D Objects from Images in the Wild

    Shangzhe Wu, Christian Rupprecht, Andrea Vedaldi

    cs.CVarXiv:1911.11130v22019
  16. Nonconvex Nonsmooth Low-Rank Minimization via Iteratively Reweighted Nuclear Norm

    Canyi Lu, Jinhui Tang, Shuicheng Yan +1

    cs.LGcs.CVmath.NAarXiv:1510.06895v12015
  17. Self-supervised Learning of Adversarial Example: Towards Good Generalizations for Deepfake Detection

    Liang Chen, Yong Zhang, Yibing Song +2

    cs.CVarXiv:2203.12208v32022
  18. GRES: Generalized Referring Expression Segmentation

    Chang Liu, Henghui Ding, Xudong Jiang

    cs.CVarXiv:2306.00968v12023
  19. BigNAS: Scaling Up Neural Architecture Search with Big Single-Stage Models

    Jiahui Yu, Pengchong Jin, Hanxiao Liu +7

    cs.CVarXiv:2003.11142v32020
  20. DiffBIR: Towards Blind Image Restoration with Generative Diffusion Prior

    Xinqi Lin, Jingwen He, Ziyan Chen +6

    cs.CVarXiv:2308.15070v32023
  21. Relay Backpropagation for Effective Learning of Deep Convolutional Neural Networks

    Li Shen, Zhouchen Lin, Qingming Huang

    cs.CVcs.LGarXiv:1512.05830v22015
  22. Deep Imbalanced Learning for Face Recognition and Attribute Prediction

    Chen Huang, Yining Li, Chen Change Loy +1

    cs.CVarXiv:1806.00194v22018
  23. Legged Locomotion in Challenging Terrains using Egocentric Vision

    Ananye Agarwal, Ashish Kumar, Jitendra Malik +1

    cs.ROcs.AIcs.CVarXiv:2211.07638v12022
  24. Feature Distillation: DNN-Oriented JPEG Compression Against Adversarial Examples

    Zihao Liu, Qi Liu, Tao Liu +4

    cs.CVcs.CRarXiv:1803.05787v22018
  25. Invisible for both Camera and LiDAR: Security of Multi-Sensor Fusion based Perception in Autonomous Driving Under Physical-World Attacks

    Yulong Cao*, Ningfei Wang*, Chaowei Xiao* +6

    cs.CRcs.CVcs.LGarXiv:2106.09249v12021
  26. PP-OCR: A Practical Ultra Lightweight OCR System

    Yuning Du, Chenxia Li, Ruoyu Guo +8

    cs.CVarXiv:2009.09941v32020
  27. SwinNet: Swin Transformer drives edge-aware RGB-D and RGB-T salient object detection

    Zhengyi Liu, Yacheng Tan, Qian He +1

    cs.CVarXiv:2204.05585v12022
  28. LST: Ladder Side-Tuning for Parameter and Memory Efficient Transfer Learning

    Yi-Lin Sung, Jaemin Cho, Mohit Bansal

    cs.CLcs.AIcs.CVarXiv:2206.06522v22022
  29. Bit-Flip Attack: Crushing Neural Network with Progressive Bit Search

    Adnan Siraj Rakin, Zhezhi He, Deliang Fan

    cs.CVcs.CRarXiv:1903.12269v22019
  30. RDD2022: A multi-national image dataset for automatic Road Damage Detection

    Deeksha Arya, Hiroya Maeda, Sanjay Kumar Ghosh +2

    cs.CVcs.AIcs.LGarXiv:2209.08538v12022
  31. Deep Learning for Multi-Task Medical Image Segmentation in Multiple Modalities

    Pim Moeskops, Jelmer M. Wolterink, Bas H. M. van der Velden +4

    cs.CVarXiv:1704.03379v12017
  32. Scaling Up Vision-Language Pre-training for Image Captioning

    Xiaowei Hu, Zhe Gan, Jianfeng Wang +4

    cs.CVcs.CLarXiv:2111.12233v22021
  33. Scalable Person Re-identification on Supervised Smoothed Manifold

    Song Bai, Xiang Bai, Qi Tian

    cs.CVarXiv:1703.08359v12017
  34. Rosetta: Large scale system for text detection and recognition in images

    Fedor Borisyuk, Albert Gordo, Viswanath Sivakumar

    cs.CVarXiv:1910.05085v12019
  35. COIN: COmpression with Implicit Neural representations

    Emilien Dupont, Adam Goliński, Milad Alizadeh +2

    eess.IVcs.CVcs.LGarXiv:2103.03123v22021
  36. Pixel-Adaptive Convolutional Neural Networks

    Hang Su, Varun Jampani, Deqing Sun +3

    cs.CVcs.GRcs.LGarXiv:1904.05373v12019
  37. Feature Importance-aware Transferable Adversarial Attacks

    Zhibo Wang, Hengchang Guo, Zhifei Zhang +3

    cs.CVarXiv:2107.14185v32021
  38. Distribution-Free, Risk-Controlling Prediction Sets

    Stephen Bates, Anastasios Angelopoulos, Lihua Lei +2

    cs.LGcs.AIcs.CVarXiv:2101.02703v32021
  39. Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation

    Shuai Yang, Yifan Zhou, Ziwei Liu +1

    cs.CVarXiv:2306.07954v22023
  40. Large Language Model Guided Tree-of-Thought

    Jieyi Long

    cs.AIcs.CLcs.CVarXiv:2305.08291v12023
  41. How would surround vehicles move? A Unified Framework for Maneuver Classification and Motion Prediction

    Nachiket Deo, Akshay Rangesh, Mohan M. Trivedi

    cs.CVarXiv:1801.06523v12018
  42. Deep Reinforcement Learning-based Image Captioning with Embedding Reward

    Zhou Ren, Xiaoyu Wang, Ning Zhang +2

    cs.CVcs.AIarXiv:1704.03899v12017
  43. Multi-Oriented Scene Text Detection via Corner Localization and Region Segmentation

    Pengyuan Lyu, Cong Yao, Wenhao Wu +2

    cs.CVarXiv:1802.08948v22018
  44. CellViT: Vision Transformers for Precise Cell Segmentation and Classification

    Fabian Hörst, Moritz Rempe, Lukas Heine +8

    eess.IVcs.CVcs.LGarXiv:2306.15350v22023
  45. Diversify and Match: A Domain Adaptive Representation Learning Paradigm for Object Detection

    Taekyung Kim, Minki Jeong, Seunghyeon Kim +2

    cs.CVarXiv:1905.05396v12019
  46. Comparing Kullback-Leibler Divergence and Mean Squared Error Loss in Knowledge Distillation

    Taehyeon Kim, Jaehoon Oh, NakYil Kim +2

    cs.LGcs.CVarXiv:2105.08919v12021
  47. Efficiently Identifying Task Groupings for Multi-Task Learning

    Christopher Fifty, Ehsan Amid, Zhe Zhao +3

    cs.LGcs.AIcs.CVarXiv:2109.04617v22021
  48. Mastering Atari Games with Limited Data

    Weirui Ye, Shaohuai Liu, Thanard Kurutach +2

    cs.LGcs.AIcs.CVarXiv:2111.00210v22021
  49. Semi-Supervised Medical Image Segmentation via Cross Teaching between CNN and Transformer

    Xiangde Luo, Minhao Hu, Tao Song +2

    eess.IVcs.CVarXiv:2112.04894v22021
  50. Video Captioning with Transferred Semantic Attributes

    Yingwei Pan, Ting Yao, Houqiang Li +1

    cs.CVarXiv:1611.07675v12016
  51. Super-Resolution with Deep Convolutional Sufficient Statistics

    Joan Bruna, Pablo Sprechmann, Yann LeCun

    cs.CVarXiv:1511.05666v42015
  52. Automatic Moth Detection from Trap Images for Pest Management

    Weiguang Ding, Graham Taylor

    cs.CVcs.LGcs.NEarXiv:1602.07383v12016
  53. FVC: A New Framework towards Deep Video Compression in Feature Space

    Zhihao Hu, Guo Lu, Dong Xu

    eess.IVcs.CVarXiv:2105.09600v22021
  54. Skin Lesion Classification Using Ensembles of Multi-Resolution EfficientNets with Meta Data

    Nils Gessert, Maximilian Nielsen, Mohsin Shaikh +2

    cs.CVarXiv:1910.03910v12019
  55. Deep Exemplar-based Colorization

    Mingming He, Dongdong Chen, Jing Liao +2

    cs.CVarXiv:1807.06587v22018
  56. Patch-based Probabilistic Image Quality Assessment for Face Selection and Improved Video-based Face Recognition

    Yongkang Wong, Shaokang Chen, Sandra Mau +2

    cs.CVstat.AParXiv:1304.0869v22013
  57. Learning the Best Pooling Strategy for Visual Semantic Embedding

    Jiacheng Chen, Hexiang Hu, Hao Wu +2

    cs.CVarXiv:2011.04305v52020
  58. DilateFormer: Multi-Scale Dilated Transformer for Visual Recognition

    Jiayu Jiao, Yu-Ming Tang, Kun-Yu Lin +4

    cs.CVarXiv:2302.01791v12023
  59. Unsupervised Misaligned Infrared and Visible Image Fusion via Cross-Modality Image Generation and Registration

    Di Wang, Jinyuan Liu, Xin Fan +1

    cs.CVarXiv:2205.11876v12022
  60. In Defense of Classical Image Processing: Fast Depth Completion on the CPU

    Jason Ku, Ali Harakeh, Steven L. Waslander

    cs.CVarXiv:1802.00036v12018