Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

11,161 to 11,220 of 18,822

  1. Hardware-oriented Approximation of Convolutional Neural Networks

    Philipp Gysel, Mohammad Motamedi, Soheil Ghiasi

    cs.CVarXiv:1604.03168v32016
  2. Artificial Fingerprinting for Generative Models: Rooting Deepfake Attribution in Training Data

    Ning Yu, Vladislav Skripniuk, Sahar Abdelnabi +1

    cs.CRcs.CVcs.CYarXiv:2007.08457v72020
  3. Unsupervised Training for 3D Morphable Model Regression

    Kyle Genova, Forrester Cole, Aaron Maschinot +3

    cs.CVarXiv:1806.06098v12018
  4. DSSL: Deep Surroundings-person Separation Learning for Text-based Person Retrieval

    Aichun Zhu, Zijie Wang, Yifeng Li +5

    cs.CVarXiv:2109.05534v12021
  5. Mix-of-Show: Decentralized Low-Rank Adaptation for Multi-Concept Customization of Diffusion Models

    Yuchao Gu, Xintao Wang, Jay Zhangjie Wu +10

    cs.CVarXiv:2305.18292v22023
  6. Deep Spatial Feature Reconstruction for Partial Person Re-identification: Alignment-Free Approach

    Lingxiao He, Jian Liang, Haiqing Li +1

    cs.CVarXiv:1801.00881v32018
  7. Quality Aware Network for Set to Set Recognition

    Yu Liu, Junjie Yan, Wanli Ouyang

    cs.CVcs.AIarXiv:1704.03373v12017
  8. Q-Diffusion: Quantizing Diffusion Models

    Xiuyu Li, Yijiang Liu, Long Lian +5

    cs.CVcs.LGarXiv:2302.04304v32023
  9. Learning to Perform Physics Experiments via Deep Reinforcement Learning

    Misha Denil, Pulkit Agrawal, Tejas D Kulkarni +3

    stat.MLcs.AIcs.CVarXiv:1611.01843v32016
  10. Weakly Supervised Deep Learning for COVID-19 Infection Detection and Classification from CT Images

    Shaoping Hu, Yuan Gao, Zhangming Niu +9

    eess.IVcs.CVcs.LGarXiv:2004.06689v12020
  11. Contrastive Learning with Stronger Augmentations

    Xiao Wang, Guo-Jun Qi

    cs.CVcs.AIcs.LGarXiv:2104.07713v22021
  12. Poisson noise reduction with non-local PCA

    Joseph Salmon, Zachary Harmany, Charles-Alban Deledalle +1

    cs.CVcs.LGstat.COarXiv:1206.0338v42012
  13. Deep Appearance Models for Face Rendering

    Stephen Lombardi, Jason Saragih, Tomas Simon +1

    cs.GRcs.CVarXiv:1808.00362v12018
  14. Unsupervised Domain Adaptation of Object Detectors: A Survey

    Poojan Oza, Vishwanath A. Sindagi, Vibashan VS +1

    cs.CVcs.LGarXiv:2105.13502v22021
  15. Making an Invisibility Cloak: Real World Adversarial Attacks on Object Detectors

    Zuxuan Wu, Ser-Nam Lim, Larry Davis +1

    cs.CVcs.CRcs.LGarXiv:1910.14667v22019
  16. Towards Multi-class Object Detection in Unconstrained Remote Sensing Imagery

    Seyed Majid Azimi, Eleonora Vig, Reza Bahmanyar +2

    cs.CVarXiv:1807.02700v32018
  17. SC-FEGAN: Face Editing Generative Adversarial Network with User's Sketch and Color

    Youngjoo Jo, Jongyoul Park

    cs.CVarXiv:1902.06838v12019
  18. Predicting the Driver's Focus of Attention: the DR(eye)VE Project

    Andrea Palazzi, Davide Abati, Simone Calderara +2

    cs.CVarXiv:1705.03854v32017
  19. Learning to Recover 3D Scene Shape from a Single Image

    Wei Yin, Jianming Zhang, Oliver Wang +4

    cs.CVarXiv:2012.09365v12020
  20. Temporal Cycle-Consistency Learning

    Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2

    cs.CVcs.LGarXiv:1904.07846v12019
  21. Deep learning is effective for the classification of OCT images of normal versus Age-related Macular Degeneration

    Cecilia S. Lee, Doug M. Baughman, Aaron Y. Lee

    stat.MLcs.CVcs.LGarXiv:1612.04891v12016
  22. ExFuse: Enhancing Feature Fusion for Semantic Segmentation

    Zhenli Zhang, Xiangyu Zhang, Chao Peng +2

    cs.CVarXiv:1804.03821v12018
  23. MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

    Tao Gong, Chengqi Lyu, Shilong Zhang +7

    cs.CVcs.CLarXiv:2305.04790v32023
  24. Geo-Neus: Geometry-Consistent Neural Implicit Surfaces Learning for Multi-view Reconstruction

    Qiancheng Fu, Qingshan Xu, Yew-Soon Ong +1

    cs.CVcs.GRarXiv:2205.15848v12022
  25. BungeeNeRF: Progressive Neural Radiance Field for Extreme Multi-scale Scene Rendering

    Yuanbo Xiangli, Linning Xu, Xingang Pan +5

    cs.CVcs.AIarXiv:2112.05504v42021
  26. Capturing Hands in Action using Discriminative Salient Points and Physics Simulation

    Dimitrios Tzionas, Luca Ballan, Abhilash Srikantha +3

    cs.CVarXiv:1506.02178v42015
  27. Synthesizing Images of Humans in Unseen Poses

    Guha Balakrishnan, Amy Zhao, Adrian V. Dalca +2

    cs.CVarXiv:1804.07739v12018
  28. A Generative Model For Zero Shot Learning Using Conditional Variational Autoencoders

    Ashish Mishra, M Shiva Krishna Reddy, Anurag Mittal +1

    cs.CVarXiv:1709.00663v22017
  29. Learning Semantic Concepts and Order for Image and Sentence Matching

    Yan Huang, Qi Wu, Liang Wang

    cs.CVarXiv:1712.02036v12017
  30. Unsupervised Learning from Narrated Instruction Videos

    Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal +3

    cs.CVcs.LGarXiv:1506.09215v42015
  31. Style Transfer by Relaxed Optimal Transport and Self-Similarity

    Nicholas Kolkin, Jason Salavon, Greg Shakhnarovich

    cs.CVarXiv:1904.12785v22019
  32. Image Restoration with Mean-Reverting Stochastic Differential Equations

    Ziwei Luo, Fredrik K. Gustafsson, Zheng Zhao +2

    cs.LGcs.CVarXiv:2301.11699v32023
  33. GeoWizard: Unleashing the Diffusion Priors for 3D Geometry Estimation from a Single Image

    Xiao Fu, Wei Yin, Mu Hu +6

    cs.CVarXiv:2403.12013v12024
  34. Parting with Misconceptions about Learning-based Vehicle Motion Planning

    Daniel Dauner, Marcel Hallgarten, Andreas Geiger +1

    cs.ROcs.AIcs.CVarXiv:2306.07962v22023
  35. Avatar-Net: Multi-scale Zero-shot Style Transfer by Feature Decoration

    Lu Sheng, Ziyi Lin, Jing Shao +1

    cs.CVarXiv:1805.03857v22018
  36. Deep learning for smart fish farming: applications, opportunities and challenges

    Xinting Yang, Song Zhang, Jintao Liu +3

    cs.CVcs.LGeess.IVarXiv:2004.11848v22020
  37. Deep learning-based transformation of the H&E stain into special stains

    Kevin de Haan, Yijie Zhang, Jonathan E. Zuckerman +11

    eess.IVcs.CVcs.LGarXiv:2008.08871v22020
  38. SSMB: Self-Supervised Local Feature Detection under Motion Blur

    Zhenjun Zhao, Fabio Bellavia, Wenting Wang +6

    cs.CVarXiv:2608.27181v12026
  39. V-Link: Recovering Lost Visual Representations in Action DiT for Vision-Language-Action Models

    Yehao Lu, Jiarui Yang, Yuning Su +10

    cs.CVarXiv:2608.25308v12026
  40. Deep Analog: Open-Set Film Emulation with Reference-Conditioned 3D LUTs

    Yitong Mu

    cs.CVcs.GRcs.LGarXiv:2608.14702v12026
  41. Projection Pursuit CPCANet for Domain Generalization

    Yu-Hsi Chen, Abd-Krim Seghouane

    cs.CVarXiv:2607.22117v12026
  42. Edge-Aware Thermal Infrared UAV Swarm Tracking

    Yu-Hsi Chen

    cs.CVcs.ROarXiv:2607.12544v12026
  43. Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking

    Zekun Qi, Xuchuan Chen, Dairu Liu +10

    cs.ROcs.AIcs.CVarXiv:2606.03985v12026
  44. YoCausal: How Far is Video Generation from World Model? A Causality Perspective

    You-Zhe Xie, Yu-Hsuan Li, Jie-Ying Lee +3

    cs.CVarXiv:2605.30346v12026
  45. InstructSAM: Segment Any Instance with Any Instructions

    Yuqian Yuan, Wentong Li, Zhaocheng Li +6

    cs.CVarXiv:2605.26102v22026
  46. PhysBrain 1.0 Technical Report

    Shijie Lian, Bin Yu, Xiaopeng Lin +10

    cs.ROcs.AIcs.CLarXiv:2605.15298v12026
  47. RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework

    Hao Gao, Shaoyu Chen, Yifan Zhu +4

    cs.CVarXiv:2604.15308v12026
  48. Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision

    Zhixiang Wei, Yi Li, Zhehan Kan +38

    cs.CVarXiv:2601.19798v12026
  49. SkyReels-V3 Technique Report

    Debang Li, Zhengcong Fei, Tuanhui Li +19

    cs.CVarXiv:2601.17323v22026
  50. PubMed-OCR: PMC Open Access OCR Annotations

    Hunter Heidenreich, Yosheb Getachew, Olivia Dinica +1

    cs.CVcs.CLcs.DLarXiv:2601.11425v12026
  51. YOLO-World: Real-Time Open-Vocabulary Object Detection

    Tianheng Cheng, Lin Song, Yixiao Ge +3

    cs.CVarXiv:2401.17270v32024
  52. DUSt3R: Geometric 3D Vision Made Easy

    Shuzhe Wang, Vincent Leroy, Yohann Cabon +2

    cs.CVarXiv:2312.14132v32023
  53. Metric3D: Towards Zero-shot Metric 3D Prediction from A Single Image

    Wei Yin, Chi Zhang, Hao Chen +5

    cs.CVcs.AIarXiv:2307.10984v12023
  54. EVA-CLIP: Improved Training Techniques for CLIP at Scale

    Quan Sun, Yuxin Fang, Ledell Wu +2

    cs.CVarXiv:2303.15389v12023
  55. Tri-Perspective View for Vision-Based 3D Semantic Occupancy Prediction

    Yuanhui Huang, Wenzhao Zheng, Yunpeng Zhang +2

    cs.CVcs.AIcs.LGarXiv:2302.07817v22023
  56. RTMDet: An Empirical Study of Designing Real-Time Object Detectors

    Chengqi Lyu, Wenwei Zhang, Haian Huang +5

    cs.CVarXiv:2212.07784v22022
  57. Magic3D: High-Resolution Text-to-3D Content Creation

    Chen-Hsuan Lin, Jun Gao, Luming Tang +7

    cs.CVcs.GRcs.LGarXiv:2211.10440v22022
  58. ViM: Out-Of-Distribution with Virtual-logit Matching

    Haoqi Wang, Zhizhong Li, Litong Feng +1

    cs.CVarXiv:2203.10807v12022
  59. Point-NeRF: Point-based Neural Radiance Fields

    Qiangeng Xu, Zexiang Xu, Julien Philip +4

    cs.CVarXiv:2201.08845v72022
  60. Vision Transformer with Deformable Attention

    Zhuofan Xia, Xuran Pan, Shiji Song +2

    cs.CVarXiv:2201.00520v32022