Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

8,761 to 8,820 of 18,866

  1. A Poisson-Gaussian Denoising Dataset with Real Fluorescence Microscopy Images

    Yide Zhang, Yinhao Zhu, Evan Nichols +4

    cs.CVcs.LGeess.IVarXiv:1812.10366v22018
  2. TAO: A Large-Scale Benchmark for Tracking Any Object

    Achal Dave, Tarasha Khurana, Pavel Tokmakov +2

    cs.CVarXiv:2005.10356v12020
  3. Uncertainty-based Traffic Accident Anticipation with Spatio-Temporal Relational Learning

    Wentao Bao, Qi Yu, Yu Kong

    cs.CVarXiv:2008.00334v12020
  4. PHOCNet: A Deep Convolutional Neural Network for Word Spotting in Handwritten Documents

    Sebastian Sudholt, Gernot A. Fink

    cs.CVarXiv:1604.00187v32016
  5. StyleDiffusion: Controllable Disentangled Style Transfer via Diffusion Models

    Zhizhong Wang, Lei Zhao, Wei Xing

    cs.CVarXiv:2308.07863v12023
  6. Deconfounded Video Moment Retrieval with Causal Intervention

    Xun Yang, Fuli Feng, Wei Ji +2

    cs.CVarXiv:2106.01534v12021
  7. Synthesis of Compositional Animations from Textual Descriptions

    Anindita Ghosh, Noshaba Cheema, Cennet Oguz +2

    cs.CVcs.LGarXiv:2103.14675v62021
  8. Fully Convolutional Architectures for Multi-Class Segmentation in Chest Radiographs

    Alexey A. Novikov, Dimitrios Lenis, David Major +3

    cs.CVcs.LGarXiv:1701.08816v42017
  9. Human Motion Prediction via Spatio-Temporal Inpainting

    Alejandro Hernandez Ruiz, Juergen Gall, Francesc Moreno-Noguer

    cs.CVarXiv:1812.05478v22018
  10. Towards a Joint Khmer Text Recognition and Word Segmentation

    Marry Kong, Rina Buoy, Sovisal Chenda +3

    cs.CVcs.CLarXiv:2608.30213v12026
  11. VIBE: Video Instruction-aligned Background music gEneration

    Aryan Vijay Bhosale, Vaibhavi Lokegaonkar, Vishnu Raj +5

    cs.SDcs.AIcs.CLarXiv:2608.30125v12026
  12. Lot Machine: Multimodal Lot Extraction from Auction Catalogs

    Mathias Zinnen, Alisha Mund, Sabine Lang +3

    cs.CVcs.AIcs.CLarXiv:2608.30510v12026
  13. Activated Gradients for Deep Neural Networks

    Mei Liu, Liangming Chen, Xiaohao Du +2

    cs.CVcs.AIarXiv:2107.04228v12021
  14. Visual Representations for Semantic Target Driven Navigation

    Arsalan Mousavian, Alexander Toshev, Marek Fiser +3

    cs.CVarXiv:1805.06066v32018
  15. An Empirical Evaluation of Similarity Measures for Time Series Classification

    Joan Serrà, Josep Lluis Arcos

    cs.LGcs.CVstat.MLarXiv:1401.3973v12014
  16. Oriented Objects as pairs of Middle Lines

    Haoran Wei, Yue Zhang, Zhonghan Chang +3

    cs.CVarXiv:1912.10694v32019
  17. CRNet: Cross-Reference Networks for Few-Shot Segmentation

    Weide Liu, Chi Zhang, Guosheng Lin +1

    cs.CVarXiv:2003.10658v12020
  18. Snapshot Distillation: Teacher-Student Optimization in One Generation

    Chenglin Yang, Lingxi Xie, Chi Su +1

    cs.CVarXiv:1812.00123v12018
  19. Deep Roto-Translation Scattering for Object Classification

    Edouard Oyallon, Stéphane Mallat

    cs.CVarXiv:1412.8659v22014
  20. Deep Cascaded Bi-Network for Face Hallucination

    Shizhan Zhu, Sifei Liu, Chen Change Loy +1

    cs.CVarXiv:1607.05046v12016
  21. Deep Supervised Hashing with Triplet Labels

    Xiaofang Wang, Yi Shi, Kris M. Kitani

    cs.CVarXiv:1612.03900v12016
  22. VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time

    Sicheng Xu, Guojun Chen, Yu-Xiao Guo +6

    cs.CVarXiv:2404.10667v22024
  23. Deep Imbalanced Attribute Classification using Visual Attention Aggregation

    Nikolaos Sarafianos, Xiang Xu, Ioannis A. Kakadiaris

    cs.CVarXiv:1807.03903v22018
  24. Invisible Steganography via Generative Adversarial Networks

    Ru Zhang, Shiqi Dong, Jianyi Liu

    cs.MMcs.CVarXiv:1807.08571v32018
  25. Deep Temporal Linear Encoding Networks

    Ali Diba, Vivek Sharma, Luc Van Gool

    cs.CVarXiv:1611.06678v12016
  26. Multi-Modal Domain Adaptation for Fine-Grained Action Recognition

    Jonathan Munro, Dima Damen

    cs.CVarXiv:2001.09691v22020
  27. MANTIS: Interleaved Multi-Image Instruction Tuning

    Dongfu Jiang, Xuan He, Huaye Zeng +4

    cs.CVcs.AIcs.CLarXiv:2405.01483v32024
  28. f-BRS: Rethinking Backpropagating Refinement for Interactive Segmentation

    Konstantin Sofiiuk, Ilia Petrov, Olga Barinova +1

    cs.CVarXiv:2001.10331v32020
  29. CIAGAN: Conditional Identity Anonymization Generative Adversarial Networks

    Maxim Maximov, Ismail Elezi, Laura Leal-Taixé

    cs.CVarXiv:2005.09544v22020
  30. Eigenspectra optoacoustic tomography achieves quantitative blood oxygenation imaging deep in tissues

    Stratis Tzoumas, Antonio Nunes, Ivan Olefir +6

    physics.med-phcs.CVphysics.opticsarXiv:1511.05846v12015
  31. Image Super-Resolution via Dual-State Recurrent Networks

    Wei Han, Shiyu Chang, Ding Liu +3

    cs.CVarXiv:1805.02704v12018
  32. Fine-grained Categorization and Dataset Bootstrapping using Deep Metric Learning with Humans in the Loop

    Yin Cui, Feng Zhou, Yuanqing Lin +1

    cs.CVarXiv:1512.05227v22015
  33. Grounding Visual Explanations

    Lisa Anne Hendricks, Ronghang Hu, Trevor Darrell +1

    cs.CVarXiv:1807.09685v22018
  34. Summit: Scaling Deep Learning Interpretability by Visualizing Activation and Attribution Summarizations

    Fred Hohman, Haekyu Park, Caleb Robinson +1

    cs.HCcs.CVcs.LGarXiv:1904.02323v32019
  35. Spatiotemporal Recurrent Convolutional Networks for Recognizing Spontaneous Micro-expressions

    Zhaoqiang Xia, Xiaopeng Hong, Xingyu Gao +2

    cs.CVarXiv:1901.04656v12019
  36. Towards Alzheimer's Disease Classification through Transfer Learning

    Marcia Hon, Naimul Khan

    cs.CVarXiv:1711.11117v12017
  37. A Visual Question Answering Model to Automate Nondestructive Evaluation Image Analysis

    Mehrdad Shafiei Dizaji, Hoda Azari

    cs.CVcs.LGarXiv:2608.29408v12026
  38. Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models

    George Stein, Jesse C. Cresswell, Rasa Hosseinzadeh +7

    cs.LGcs.CVstat.MLarXiv:2306.04675v22023
  39. GestureDiffuCLIP: Gesture Diffusion Model with CLIP Latents

    Tenglong Ao, Zeyi Zhang, Libin Liu

    cs.CVcs.GRarXiv:2303.14613v42023
  40. Using Channel Representations in Regularization Terms: A Case Study on Image Diffusion

    Christian Heinemann, Freddie Åström, George Baravdish +3

    cs.CVarXiv:2608.29227v12026
  41. ATGS: Anchored Temporal Gaussian Splatting for Long Volumetric Video Representation

    Jiahao Wu, Jie Liang, Die Hu +6

    cs.CVarXiv:2608.30184v12026
  42. FrameScope: Temporal Data Valuation for Stream Active Learning in Autonomous Vehicle Systems

    Yuheng Zhu, Man-Ki Yoon

    cs.CVcs.LGarXiv:2608.28672v12026
  43. Objectron: A Large Scale Dataset of Object-Centric Videos in the Wild with Pose Annotations

    Adel Ahmadyan, Liangkai Zhang, Jianing Wei +2

    cs.CVarXiv:2012.09988v12020
  44. Towards High-Resolution Salient Object Detection

    Yi Zeng, Pingping Zhang, Jianming Zhang +2

    cs.CVarXiv:1908.07274v12019
  45. Applying Guidance in a Limited Interval Improves Sample and Distribution Quality in Diffusion Models

    Tuomas Kynkäänniemi, Miika Aittala, Tero Karras +3

    cs.CVcs.AIcs.LGarXiv:2404.07724v22024
  46. Making Deep Heatmaps Robust to Partial Occlusions for 3D Object Pose Estimation

    Markus Oberweger, Mahdi Rad, Vincent Lepetit

    cs.CVarXiv:1804.03959v32018
  47. Effective Graph and Rank-based Contextual Embeddings for Textual and Multimedia Data

    Thiago César Castilho Almeida, Gustavo Rosseto Letício, Lucas Pascotti Valem +2

    cs.LGcs.CVcs.IRarXiv:2608.29001v12026
  48. Building Statistical Shape Spaces for 3D Human Modeling

    Leonid Pishchulin, Stefanie Wuhrer, Thomas Helten +2

    cs.CVarXiv:1503.05860v22015
  49. MetaPoison: Practical General-purpose Clean-label Data Poisoning

    W. Ronny Huang, Jonas Geiping, Liam Fowl +2

    cs.LGcs.AIcs.CRarXiv:2004.00225v22020
  50. Deep Hough Transform for Semantic Line Detection

    Kai Zhao, Qi Han, Chang-Bin Zhang +2

    cs.CVarXiv:2003.04676v42020
  51. SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation

    Huaishao Luo, Junwei Bao, Youzheng Wu +2

    cs.CVcs.AIarXiv:2211.14813v22022
  52. Visual to Sound: Generating Natural Sound for Videos in the Wild

    Yipin Zhou, Zhaowen Wang, Chen Fang +2

    cs.CVarXiv:1712.01393v22017
  53. PanopticFusion: Online Volumetric Semantic Mapping at the Level of Stuff and Things

    Gaku Narita, Takashi Seno, Tomoya Ishikawa +1

    cs.CVcs.ROarXiv:1903.01177v22019
  54. InstaBoost: Boosting Instance Segmentation via Probability Map Guided Copy-Pasting

    Hao-Shu Fang, Jianhua Sun, Runzhong Wang +3

    cs.CVarXiv:1908.07801v12019
  55. SESF-Fuse: An Unsupervised Deep Model for Multi-Focus Image Fusion

    Boyuan Ma, Xiaojuan Ban, Haiyou Huang +1

    cs.CVarXiv:1908.01703v22019
  56. AdaptAV: Continuous Adaption of Vision Models for Autonomous Vehicles Using Cloud-based Oracle

    Yuheng Zhu, Dhruva Ungrupulithaya, Boluo Ge +1

    cs.CVcs.LGarXiv:2608.28673v12026
  57. CIA-Net: Robust Nuclei Instance Segmentation with Contour-aware Information Aggregation

    Yanning Zhou, Omer Fahri Onder, Qi Dou +3

    cs.CVarXiv:1903.05358v12019
  58. A Tensor Variational Formulation of Gradient Energy Total Variation

    Freddie Åström, George Baravdish, Michael Felsberg

    cs.CVarXiv:2608.29172v12026
  59. A Fully Convolutional Neural Network based Structured Prediction Approach Towards the Retinal Vessel Segmentation

    Avijit Dasgupta, Sonam Singh

    cs.CVarXiv:1611.02064v22016
  60. Natural Language Does Not Emerge 'Naturally' in Multi-Agent Dialog

    Satwik Kottur, José M. F. Moura, Stefan Lee +1

    cs.CLcs.AIcs.CVarXiv:1706.08502v32017