Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

14,521 to 14,580 of 18,830

  1. Baking Neural Radiance Fields for Real-Time View Synthesis

    Peter Hedman, Pratul P. Srinivasan, Ben Mildenhall +2

    cs.CVcs.GRarXiv:2103.14645v12021
  2. Mind the Class Weight Bias: Weighted Maximum Mean Discrepancy for Unsupervised Domain Adaptation

    Hongliang Yan, Yukang Ding, Peihua Li +3

    cs.CVarXiv:1705.00609v12017
  3. Joint Distribution Alignment for Universal Domain Adaptation

    Shizhe Li, Hongshan Pu, Mengying Xie +2

    cs.LGcs.CVarXiv:2608.24429v12026
  4. Model Effect or Label Effect? Refined Annotations and a Human-Referenced Benchmark for Pulmonary Embolism Segmentation

    Qihang Sun, Zhongxiao Liu, Bailiang Jian +6

    eess.IVcs.CVarXiv:2608.24486v12026
  5. Rethinking RGB-D Salient Object Detection: Models, Data Sets, and Large-Scale Benchmarks

    Deng-Ping Fan, Zheng Lin, Jia-Xing Zhao +5

    cs.CVarXiv:1907.06781v22019
  6. Correlation Congruence for Knowledge Distillation

    Baoyun Peng, Xiao Jin, Jiaheng Liu +5

    cs.CVarXiv:1904.01802v12019
  7. Bayesian Loss for Crowd Count Estimation with Point Supervision

    Zhiheng Ma, Xing Wei, Xiaopeng Hong +1

    cs.CVarXiv:1908.03684v12019
  8. SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite Imagery

    Yezhen Cong, Samar Khanna, Chenlin Meng +6

    cs.CVcs.AIarXiv:2207.08051v32022
  9. Detecting and Recognizing Human-Object Interactions

    Georgia Gkioxari, Ross Girshick, Piotr Dollár +1

    cs.CVarXiv:1704.07333v32017
  10. EXPANSE: A Deep Continual / Progressive Learning System for Deep Transfer Learning

    Mohammadreza Iman, John A. Miller, Khaled Rasheed +2

    cs.LGcs.CVarXiv:2205.10356v22022
  11. When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs

    Zhengxiang Wang, Owen Rambow

    cs.AIcs.CVarXiv:2608.23978v12026
  12. A Fourier Perspective on Model Robustness in Computer Vision

    Dong Yin, Raphael Gontijo Lopes, Jonathon Shlens +2

    cs.LGcs.CVstat.MLarXiv:1906.08988v32019
  13. OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

    Anas Awadalla, Irena Gao, Josh Gardner +13

    cs.CVcs.AIcs.LGarXiv:2308.01390v22023
  14. High-Fidelity Generative Image Compression

    Fabian Mentzer, George Toderici, Michael Tschannen +1

    eess.IVcs.CVcs.LGarXiv:2006.09965v32020
  15. Towards Real-World Blind Face Restoration with Generative Facial Prior

    Xintao Wang, Yu Li, Honglun Zhang +1

    cs.CVarXiv:2101.04061v22021
  16. Boosting Image Captioning with Attributes

    Ting Yao, Yingwei Pan, Yehao Li +2

    cs.CVarXiv:1611.01646v12016
  17. Generating Visual Explanations

    Lisa Anne Hendricks, Zeynep Akata, Marcus Rohrbach +3

    cs.CVcs.AIcs.CLarXiv:1603.08507v12016
  18. Detecting and classifying lesions in mammograms with Deep Learning

    Dezső Ribli, Anna Horváth, Zsuzsa Unger +2

    cs.CVarXiv:1707.08401v32017
  19. Learning Synergies between Pushing and Grasping with Self-supervised Deep Reinforcement Learning

    Andy Zeng, Shuran Song, Stefan Welker +3

    cs.ROcs.AIcs.CVarXiv:1803.09956v32018
  20. Bi-Real Net: Enhancing the Performance of 1-bit CNNs With Improved Representational Capability and Advanced Training Algorithm

    Zechun Liu, Baoyuan Wu, Wenhan Luo +3

    cs.CVarXiv:1808.00278v52018
  21. Paraphrasing Complex Network: Network Compression via Factor Transfer

    Jangho Kim, SeongUk Park, Nojun Kwak

    cs.CVarXiv:1802.04977v32018
  22. MotionGPT: Human Motion as a Foreign Language

    Biao Jiang, Xin Chen, Wen Liu +3

    cs.CVcs.CLcs.GRarXiv:2306.14795v22023
  23. Closed-Form Factorization of Latent Semantics in GANs

    Yujun Shen, Bolei Zhou

    cs.CVarXiv:2007.06600v42020
  24. Video Classification with Channel-Separated Convolutional Networks

    Du Tran, Heng Wang, Lorenzo Torresani +1

    cs.CVcs.AIarXiv:1904.02811v42019
  25. Deep Stacked Hierarchical Multi-patch Network for Image Deblurring

    Hongguang Zhang, Yuchao Dai, Hongdong Li +1

    cs.CVarXiv:1904.03468v12019
  26. A Comprehensive Survey of Image Augmentation Techniques for Deep Learning

    Mingle Xu, Sook Yoon, Alvaro Fuentes +1

    cs.CVarXiv:2205.01491v22022
  27. Vehicle Detection from 3D Lidar Using Fully Convolutional Network

    Bo Li, Tianlei Zhang, Tian Xia

    cs.CVcs.ROarXiv:1608.07916v12016
  28. Towards Open World Recognition

    Abhijit Bendale, Terrance Boult

    cs.CVarXiv:1412.5687v12014
  29. Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

    Sihyun Yu, Sangkyung Kwak, Huiwon Jang +4

    cs.CVcs.LGarXiv:2410.06940v42024
  30. TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency Distillation

    Xiaoda Yang, Yuxiang Liu, Kaiwen Zheng +10

    cs.CVarXiv:2608.24674v12026
  31. LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

    Yanwei Li, Chengyao Wang, Jiaya Jia

    cs.CVcs.CLarXiv:2311.17043v12023
  32. Metadata-Aware Adaptation of a Generative Foundation Model for Conditional CMR Synthesis

    Marc Rodríguez, Grzegorz Skorupko, Nay Aung +3

    cs.CVcs.AIarXiv:2608.24342v12026
  33. The Reversible Residual Network: Backpropagation Without Storing Activations

    Aidan N. Gomez, Mengye Ren, Raquel Urtasun +1

    cs.CVcs.LGarXiv:1707.04585v12017
  34. Weakly Supervised Learning of Instance Segmentation with Inter-pixel Relations

    Jiwoon Ahn, Sunghyun Cho, Suha Kwak

    cs.CVcs.LGarXiv:1904.05044v32019
  35. Interpretable 3D Human Action Analysis with Temporal Convolutional Networks

    Tae Soo Kim, Austin Reiter

    cs.CVarXiv:1704.04516v12017
  36. In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised Learning

    Mamshad Nayeem Rizve, Kevin Duarte, Yogesh S Rawat +1

    cs.LGcs.CVarXiv:2101.06329v32021
  37. Fast Training of Convolutional Networks through FFTs

    Michael Mathieu, Mikael Henaff, Yann LeCun

    cs.CVcs.LGcs.NEarXiv:1312.5851v52013
  38. Fast AutoAugment

    Sungbin Lim, Ildoo Kim, Taesup Kim +2

    cs.LGcs.CVstat.MLarXiv:1905.00397v22019
  39. DeepPhys: Video-Based Physiological Measurement Using Convolutional Attention Networks

    Weixuan Chen, Daniel McDuff

    cs.CVcs.HCarXiv:1805.07888v22018
  40. Zero-shot Image-to-Image Translation

    Gaurav Parmar, Krishna Kumar Singh, Richard Zhang +3

    cs.CVcs.GRcs.LGarXiv:2302.03027v12023
  41. Generating Natural Adversarial Examples

    Zhengli Zhao, Dheeru Dua, Sameer Singh

    cs.LGcs.AIcs.CLarXiv:1710.11342v22017
  42. A Novel Performance Evaluation Methodology for Single-Target Trackers

    Matej Kristan, Jiri Matas, Ales Leonardis +6

    cs.CVarXiv:1503.01313v32015
  43. Paint by Example: Exemplar-based Image Editing with Diffusion Models

    Binxin Yang, Shuyang Gu, Bo Zhang +5

    cs.CVarXiv:2211.13227v12022
  44. Deep Feature Flow for Video Recognition

    Xizhou Zhu, Yuwen Xiong, Jifeng Dai +2

    cs.CVarXiv:1611.07715v22016
  45. Human-Inspired Social Engagement Analysis via Interpretable Mutual Visual Attention

    Urwa Fatima, Mohammad Zohaib, Francesca Odone +1

    cs.CVarXiv:2608.24580v12026
  46. DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion Frames

    Erik Wijmans, Abhishek Kadian, Ari Morcos +5

    cs.CVcs.AIcs.LGarXiv:1911.00357v22019
  47. One-Shot Free-View Neural Talking-Head Synthesis for Video Conferencing

    Ting-Chun Wang, Arun Mallya, Ming-Yu Liu

    cs.CVarXiv:2011.15126v32020
  48. Learning Category-Specific Mesh Reconstruction from Image Collections

    Angjoo Kanazawa, Shubham Tulsiani, Alexei A. Efros +1

    cs.CVarXiv:1803.07549v22018
  49. BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment

    Kelvin C. K. Chan, Shangchen Zhou, Xiangyu Xu +1

    cs.CVarXiv:2104.13371v12021
  50. HumanNeRF: Free-viewpoint Rendering of Moving People from Monocular Video

    Chung-Yi Weng, Brian Curless, Pratul P. Srinivasan +2

    cs.CVcs.GRarXiv:2201.04127v22022
  51. Face Detection using Deep Learning: An Improved Faster RCNN Approach

    Xudong Sun, Pengcheng Wu, Steven C. H. Hoi

    cs.CVarXiv:1701.08289v12017
  52. Graph-Supervised Hierarchical Clinical Alignment for Radiology Report Generation with Large Language Models

    Yingshu Li, Yunyi Liu, Zhanyu Wang +4

    cs.CVarXiv:2608.24121v12026
  53. DRRG: A Discrete Diffusion Framework for Radiology Report Generation

    Shaoyang Zhoua, Yingshu Li, Yunyi Liu +4

    cs.CVarXiv:2608.24105v12026
  54. Learning joint reconstruction of hands and manipulated objects

    Yana Hasson, Gül Varol, Dimitrios Tzionas +4

    cs.CVarXiv:1904.05767v12019
  55. Gold-YOLO: Efficient Object Detector via Gather-and-Distribute Mechanism

    Chengcheng Wang, Wei He, Ying Nie +4

    cs.CVcs.AIarXiv:2309.11331v52023
  56. TorchMorph: CUDA-accelerated Morphological Transforms

    Kai Zhao

    cs.CVarXiv:2608.24738v12026
  57. Evolving Deep Convolutional Neural Networks for Image Classification

    Yanan Sun, Bing Xue, Mengjie Zhang +1

    cs.NEcs.CVarXiv:1710.10741v32017
  58. Quantifying the effects of data augmentation and stain color normalization in convolutional neural networks for computational pathology

    David Tellez, Geert Litjens, Peter Bandi +4

    cs.CVarXiv:1902.06543v22019
  59. Detecting and Simulating Artifacts in GAN Fake Images

    Xu Zhang, Svebor Karaman, Shih-Fu Chang

    cs.CVeess.IVarXiv:1907.06515v22019
  60. Diverse Beam Search: Decoding Diverse Solutions from Neural Sequence Models

    Ashwin K Vijayakumar, Michael Cogswell, Ramprasath R. Selvaraju +4

    cs.AIcs.CLcs.CVarXiv:1610.02424v22016