Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

10,981 to 11,040 of 18,867

  1. A Deep Learning-based Radar and Camera Sensor Fusion Architecture for Object Detection

    Felix Nobis, Maximilian Geisslinger, Markus Weber +2

    cs.CVarXiv:2005.07431v12020
  2. Hi-CMD: Hierarchical Cross-Modality Disentanglement for Visible-Infrared Person Re-Identification

    Seokeon Choi, Sumin Lee, Youngeun Kim +2

    cs.CVarXiv:1912.01230v32019
  3. Probabilistic Embeddings for Cross-Modal Retrieval

    Sanghyuk Chun, Seong Joon Oh, Rafael Sampaio de Rezende +2

    cs.CVarXiv:2101.05068v22021
  4. Exploring Hate Speech Detection in Multimodal Publications

    Raul Gomez, Jaume Gibert, Lluis Gomez +1

    cs.CVcs.CLarXiv:1910.03814v12019
  5. TGIF: A New Dataset and Benchmark on Animated GIF Description

    Yuncheng Li, Yale Song, Liangliang Cao +4

    cs.CVarXiv:1604.02748v22016
  6. An Image is Worth 32 Tokens for Reconstruction and Generation

    Qihang Yu, Mark Weber, Xueqing Deng +3

    cs.CVarXiv:2406.07550v12024
  7. ROAD: Reality Oriented Adaptation for Semantic Segmentation of Urban Scenes

    Yuhua Chen, Wen Li, Luc Van Gool

    cs.CVarXiv:1711.11556v22017
  8. Diffusion models as plug-and-play priors

    Alexandros Graikos, Nikolay Malkin, Nebojsa Jojic +1

    cs.LGcs.CVarXiv:2206.09012v32022
  9. Universal Adversarial Perturbations Against Semantic Image Segmentation

    Jan Hendrik Metzen, Mummadi Chaithanya Kumar, Thomas Brox +1

    stat.MLcs.AIcs.CVarXiv:1704.05712v32017
  10. Do Convnets Learn Correspondence?

    Jonathan Long, Ning Zhang, Trevor Darrell

    cs.CVcs.LGcs.NEarXiv:1411.1091v12014
  11. Auxiliary Tasks in Multi-task Learning

    Lukas Liebel, Marco Körner

    cs.CVcs.LGarXiv:1805.06334v22018
  12. Transferring Rich Feature Hierarchies for Robust Visual Tracking

    Naiyan Wang, Siyi Li, Abhinav Gupta +1

    cs.CVcs.NEarXiv:1501.04587v22015
  13. Loss Functions for Neural Networks for Image Processing

    Hang Zhao, Orazio Gallo, Iuri Frosio +1

    cs.CVarXiv:1511.08861v32015
  14. A-NeRF: Articulated Neural Radiance Fields for Learning Human Shape, Appearance, and Pose

    Shih-Yang Su, Frank Yu, Michael Zollhoefer +1

    cs.CVcs.GRarXiv:2102.06199v32021
  15. Projected GANs Converge Faster

    Axel Sauer, Kashyap Chitta, Jens Müller +1

    cs.CVcs.LGarXiv:2111.01007v12021
  16. TopFormer: Token Pyramid Transformer for Mobile Semantic Segmentation

    Wenqiang Zhang, Zilong Huang, Guozhong Luo +5

    cs.CVarXiv:2204.05525v12022
  17. Robust Subspace Learning: Robust PCA, Robust Subspace Tracking, and Robust Subspace Recovery

    Namrata Vaswani, Thierry Bouwmans, Sajid Javed +1

    cs.ITcs.CVstat.MEarXiv:1711.09492v42017
  18. Latent Action Pretraining from Videos

    Seonghyeon Ye, Joel Jang, Byeongguk Jeon +13

    cs.ROcs.CLcs.CVarXiv:2410.11758v22024
  19. ASpanFormer: Detector-Free Image Matching with Adaptive Span Transformer

    Hongkai Chen, Zixin Luo, Lei Zhou +6

    cs.CVarXiv:2208.14201v12022
  20. LSQ+: Improving low-bit quantization through learnable offsets and better initialization

    Yash Bhalgat, Jinwon Lee, Markus Nagel +2

    cs.CVcs.LGstat.MLarXiv:2004.09576v12020
  21. DeepSentiBank: Visual Sentiment Concept Classification with Deep Convolutional Neural Networks

    Tao Chen, Damian Borth, Trevor Darrell +1

    cs.CVcs.LGcs.MMarXiv:1410.8586v12014
  22. Robot Parkour Learning

    Ziwen Zhuang, Zipeng Fu, Jianren Wang +4

    cs.ROcs.AIcs.CVarXiv:2309.05665v22023
  23. Generating Natural Questions About an Image

    Nasrin Mostafazadeh, Ishan Misra, Jacob Devlin +3

    cs.CLcs.AIcs.CVarXiv:1603.06059v32016
  24. Single Shot Text Detector with Regional Attention

    Pan He, Weilin Huang, Tong He +3

    cs.CVarXiv:1709.00138v12017
  25. Learning N:M Fine-grained Structured Sparse Neural Networks From Scratch

    Aojun Zhou, Yukun Ma, Junnan Zhu +5

    cs.CVcs.ARarXiv:2102.04010v22021
  26. Efficient Diffusion Training via Min-SNR Weighting Strategy

    Tiankai Hang, Shuyang Gu, Chen Li +5

    cs.CVarXiv:2303.09556v32023
  27. End-to-end Audio-visual Speech Recognition with Conformers

    Pingchuan Ma, Stavros Petridis, Maja Pantic

    cs.CVeess.ASarXiv:2102.06657v12021
  28. TailorNet: Predicting Clothing in 3D as a Function of Human Pose, Shape and Garment Style

    Chaitanya Patel, Zhouyingcheng Liao, Gerard Pons-Moll

    cs.CVcs.GRarXiv:2003.04583v22020
  29. Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis

    Jian Han, Jinlai Liu, Yi Jiang +5

    cs.CVarXiv:2412.04431v22024
  30. Weakly-supervised Disentangling with Recurrent Transformations for 3D View Synthesis

    Jimei Yang, Scott Reed, Ming-Hsuan Yang +1

    cs.LGcs.AIcs.CVarXiv:1601.00706v12016
  31. MSR-net:Low-light Image Enhancement Using Deep Convolutional Network

    Liang Shen, Zihan Yue, Fan Feng +3

    cs.CVarXiv:1711.02488v12017
  32. Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning

    Zhicheng Huang, Zhaoyang Zeng, Yupan Huang +3

    cs.CVarXiv:2104.03135v22021
  33. IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning

    Pan Lu, Liang Qiu, Jiaqi Chen +6

    cs.CVcs.AIcs.CLarXiv:2110.13214v42021
  34. SCTransNet: Spatial-channel Cross Transformer Network for Infrared Small Target Detection

    Shuai Yuan, Hanlin Qin, Xiang Yan +2

    cs.CVarXiv:2401.15583v32024
  35. Single-source Domain Expansion Network for Cross-Scene Hyperspectral Image Classification

    Yuxiang Zhang, Wei Li, Weidong Sun +2

    cs.CVarXiv:2209.01634v12022
  36. SteganoGAN: High Capacity Image Steganography with GANs

    Kevin Alex Zhang, Alfredo Cuesta-Infante, Lei Xu +1

    cs.CVcs.LGcs.MMarXiv:1901.03892v22019
  37. DeepSeeNet: A deep learning model for automated classification of patient-based age-related macular degeneration severity from color fundus photographs

    Yifan Peng, Shazia Dharssi, Qingyu Chen +5

    cs.CVarXiv:1811.07492v22018
  38. Quaternion Fourier Transform on Quaternion Fields and Generalizations

    Eckhard Hitzer

    math.RAcs.CVmath-pharXiv:1306.1023v12013
  39. Deep Multimodal Fusion by Channel Exchanging

    Yikai Wang, Wenbing Huang, Fuchun Sun +3

    cs.CVcs.LGarXiv:2011.05005v22020
  40. Multimodal Residual Learning for Visual QA

    Jin-Hwa Kim, Sang-Woo Lee, Dong-Hyun Kwak +4

    cs.CVarXiv:1606.01455v22016
  41. ZegCLIP: Towards Adapting CLIP for Zero-shot Semantic Segmentation

    Ziqin Zhou, Bowen Zhang, Yinjie Lei +2

    cs.CVarXiv:2212.03588v32022
  42. Focal Sparse Convolutional Networks for 3D Object Detection

    Yukang Chen, Yanwei Li, Xiangyu Zhang +2

    cs.CVcs.LGarXiv:2204.12463v12022
  43. MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

    Shanghua Gao, Pan Zhou, Ming-Ming Cheng +1

    cs.CVarXiv:2303.14389v22023
  44. AMP-Net: Denoising based Deep Unfolding for Compressive Image Sensing

    Zhonghao Zhang, Yipeng Liu, Jiani Liu +2

    eess.IVcs.CVarXiv:2004.10078v22020
  45. Cluster Contrast for Unsupervised Person Re-Identification

    Zuozhuo Dai, Guangyuan Wang, Weihao Yuan +3

    cs.CVarXiv:2103.11568v42021
  46. Review: deep learning on 3D point clouds

    Saifullahi Aminu Bello, Shangshu Yu, Cheng Wang

    cs.CVarXiv:2001.06280v12020
  47. Going Deeper into First-Person Activity Recognition

    Minghuang Ma, Haoqi Fan, Kris M. Kitani

    cs.CVarXiv:1605.03688v12016
  48. MetaFormer Baselines for Vision

    Weihao Yu, Chenyang Si, Pan Zhou +5

    cs.CVcs.AIcs.LGarXiv:2210.13452v42022
  49. CFA: Coupled-hypersphere-based Feature Adaptation for Target-Oriented Anomaly Localization

    Sungwook Lee, Seunghyun Lee, Byung Cheol Song

    cs.CVcs.LGarXiv:2206.04325v12022
  50. Fast Hyperspectral Image Denoising and Inpainting Based on Low-Rank and Sparse Representations

    Lina Zhuang, Jose M. Bioucas-Dias

    eess.IVcs.CVarXiv:2103.06842v12021
  51. Learning Depth-Guided Convolutions for Monocular 3D Object Detection

    Mingyu Ding, Yuqi Huo, Hongwei Yi +4

    cs.CVarXiv:1912.04799v22019
  52. DeepHuman: 3D Human Reconstruction from a Single Image

    Zerong Zheng, Tao Yu, Yixuan Wei +2

    cs.CVarXiv:1903.06473v22019
  53. GSPN: Generative Shape Proposal Network for 3D Instance Segmentation in Point Cloud

    Li Yi, Wang Zhao, He Wang +2

    cs.CVarXiv:1812.03320v12018
  54. Neural Descriptor Fields: SE(3)-Equivariant Object Representations for Manipulation

    Anthony Simeonov, Yilun Du, Andrea Tagliasacchi +4

    cs.ROcs.AIcs.CVarXiv:2112.05124v12021
  55. Language-Grounded Indoor 3D Semantic Segmentation in the Wild

    David Rozenberszki, Or Litany, Angela Dai

    cs.CVarXiv:2204.07761v22022
  56. V2V4Real: A Real-world Large-scale Dataset for Vehicle-to-Vehicle Cooperative Perception

    Runsheng Xu, Xin Xia, Jinlong Li +10

    cs.CVarXiv:2303.07601v22023
  57. A Review and Comparative Study on Probabilistic Object Detection in Autonomous Driving

    Di Feng, Ali Harakeh, Steven Waslander +1

    cs.CVcs.ROarXiv:2011.10671v22020
  58. Interpreting Deep Visual Representations via Network Dissection

    Bolei Zhou, David Bau, Aude Oliva +1

    cs.CVarXiv:1711.05611v22017
  59. ConvNet Architecture Search for Spatiotemporal Feature Learning

    Du Tran, Jamie Ray, Zheng Shou +2

    cs.CVarXiv:1708.05038v12017
  60. SafetyNet: Detecting and Rejecting Adversarial Examples Robustly

    Jiajun Lu, Theerasit Issaranon, David Forsyth

    cs.CVcs.LGarXiv:1704.00103v22017