Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
11,161 to 11,220 of 18,822
Hardware-oriented Approximation of Convolutional Neural Networks
Philipp Gysel, Mohammad Motamedi, Soheil Ghiasi
cs.CVarXiv:1604.03168v32016Artificial Fingerprinting for Generative Models: Rooting Deepfake Attribution in Training Data
Ning Yu, Vladislav Skripniuk, Sahar Abdelnabi +1
cs.CRcs.CVcs.CYarXiv:2007.08457v72020Unsupervised Training for 3D Morphable Model Regression
Kyle Genova, Forrester Cole, Aaron Maschinot +3
cs.CVarXiv:1806.06098v12018DSSL: Deep Surroundings-person Separation Learning for Text-based Person Retrieval
Aichun Zhu, Zijie Wang, Yifeng Li +5
cs.CVarXiv:2109.05534v12021Mix-of-Show: Decentralized Low-Rank Adaptation for Multi-Concept Customization of Diffusion Models
Yuchao Gu, Xintao Wang, Jay Zhangjie Wu +10
cs.CVarXiv:2305.18292v22023Deep Spatial Feature Reconstruction for Partial Person Re-identification: Alignment-Free Approach
Lingxiao He, Jian Liang, Haiqing Li +1
cs.CVarXiv:1801.00881v32018Quality Aware Network for Set to Set Recognition
Yu Liu, Junjie Yan, Wanli Ouyang
cs.CVcs.AIarXiv:1704.03373v12017Q-Diffusion: Quantizing Diffusion Models
Xiuyu Li, Yijiang Liu, Long Lian +5
cs.CVcs.LGarXiv:2302.04304v32023Learning to Perform Physics Experiments via Deep Reinforcement Learning
Misha Denil, Pulkit Agrawal, Tejas D Kulkarni +3
stat.MLcs.AIcs.CVarXiv:1611.01843v32016Weakly Supervised Deep Learning for COVID-19 Infection Detection and Classification from CT Images
Shaoping Hu, Yuan Gao, Zhangming Niu +9
eess.IVcs.CVcs.LGarXiv:2004.06689v12020Contrastive Learning with Stronger Augmentations
Xiao Wang, Guo-Jun Qi
cs.CVcs.AIcs.LGarXiv:2104.07713v22021Poisson noise reduction with non-local PCA
Joseph Salmon, Zachary Harmany, Charles-Alban Deledalle +1
cs.CVcs.LGstat.COarXiv:1206.0338v42012Deep Appearance Models for Face Rendering
Stephen Lombardi, Jason Saragih, Tomas Simon +1
cs.GRcs.CVarXiv:1808.00362v12018Unsupervised Domain Adaptation of Object Detectors: A Survey
Poojan Oza, Vishwanath A. Sindagi, Vibashan VS +1
cs.CVcs.LGarXiv:2105.13502v22021Making an Invisibility Cloak: Real World Adversarial Attacks on Object Detectors
Zuxuan Wu, Ser-Nam Lim, Larry Davis +1
cs.CVcs.CRcs.LGarXiv:1910.14667v22019Towards Multi-class Object Detection in Unconstrained Remote Sensing Imagery
Seyed Majid Azimi, Eleonora Vig, Reza Bahmanyar +2
cs.CVarXiv:1807.02700v32018SC-FEGAN: Face Editing Generative Adversarial Network with User's Sketch and Color
Youngjoo Jo, Jongyoul Park
cs.CVarXiv:1902.06838v12019Predicting the Driver's Focus of Attention: the DR(eye)VE Project
Andrea Palazzi, Davide Abati, Simone Calderara +2
cs.CVarXiv:1705.03854v32017Learning to Recover 3D Scene Shape from a Single Image
Wei Yin, Jianming Zhang, Oliver Wang +4
cs.CVarXiv:2012.09365v12020Temporal Cycle-Consistency Learning
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2
cs.CVcs.LGarXiv:1904.07846v12019Deep learning is effective for the classification of OCT images of normal versus Age-related Macular Degeneration
Cecilia S. Lee, Doug M. Baughman, Aaron Y. Lee
stat.MLcs.CVcs.LGarXiv:1612.04891v12016ExFuse: Enhancing Feature Fusion for Semantic Segmentation
Zhenli Zhang, Xiangyu Zhang, Chao Peng +2
cs.CVarXiv:1804.03821v12018MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Tao Gong, Chengqi Lyu, Shilong Zhang +7
cs.CVcs.CLarXiv:2305.04790v32023Geo-Neus: Geometry-Consistent Neural Implicit Surfaces Learning for Multi-view Reconstruction
Qiancheng Fu, Qingshan Xu, Yew-Soon Ong +1
cs.CVcs.GRarXiv:2205.15848v12022BungeeNeRF: Progressive Neural Radiance Field for Extreme Multi-scale Scene Rendering
Yuanbo Xiangli, Linning Xu, Xingang Pan +5
cs.CVcs.AIarXiv:2112.05504v42021Capturing Hands in Action using Discriminative Salient Points and Physics Simulation
Dimitrios Tzionas, Luca Ballan, Abhilash Srikantha +3
cs.CVarXiv:1506.02178v42015Synthesizing Images of Humans in Unseen Poses
Guha Balakrishnan, Amy Zhao, Adrian V. Dalca +2
cs.CVarXiv:1804.07739v12018A Generative Model For Zero Shot Learning Using Conditional Variational Autoencoders
Ashish Mishra, M Shiva Krishna Reddy, Anurag Mittal +1
cs.CVarXiv:1709.00663v22017Learning Semantic Concepts and Order for Image and Sentence Matching
Yan Huang, Qi Wu, Liang Wang
cs.CVarXiv:1712.02036v12017Unsupervised Learning from Narrated Instruction Videos
Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal +3
cs.CVcs.LGarXiv:1506.09215v42015Style Transfer by Relaxed Optimal Transport and Self-Similarity
Nicholas Kolkin, Jason Salavon, Greg Shakhnarovich
cs.CVarXiv:1904.12785v22019Image Restoration with Mean-Reverting Stochastic Differential Equations
Ziwei Luo, Fredrik K. Gustafsson, Zheng Zhao +2
cs.LGcs.CVarXiv:2301.11699v32023GeoWizard: Unleashing the Diffusion Priors for 3D Geometry Estimation from a Single Image
Xiao Fu, Wei Yin, Mu Hu +6
cs.CVarXiv:2403.12013v12024Parting with Misconceptions about Learning-based Vehicle Motion Planning
Daniel Dauner, Marcel Hallgarten, Andreas Geiger +1
cs.ROcs.AIcs.CVarXiv:2306.07962v22023Avatar-Net: Multi-scale Zero-shot Style Transfer by Feature Decoration
Lu Sheng, Ziyi Lin, Jing Shao +1
cs.CVarXiv:1805.03857v22018Deep learning for smart fish farming: applications, opportunities and challenges
Xinting Yang, Song Zhang, Jintao Liu +3
cs.CVcs.LGeess.IVarXiv:2004.11848v22020Deep learning-based transformation of the H&E stain into special stains
Kevin de Haan, Yijie Zhang, Jonathan E. Zuckerman +11
eess.IVcs.CVcs.LGarXiv:2008.08871v22020SSMB: Self-Supervised Local Feature Detection under Motion Blur
Zhenjun Zhao, Fabio Bellavia, Wenting Wang +6
cs.CVarXiv:2608.27181v12026V-Link: Recovering Lost Visual Representations in Action DiT for Vision-Language-Action Models
Yehao Lu, Jiarui Yang, Yuning Su +10
cs.CVarXiv:2608.25308v12026Deep Analog: Open-Set Film Emulation with Reference-Conditioned 3D LUTs
Yitong Mu
cs.CVcs.GRcs.LGarXiv:2608.14702v12026Projection Pursuit CPCANet for Domain Generalization
Yu-Hsi Chen, Abd-Krim Seghouane
cs.CVarXiv:2607.22117v12026Edge-Aware Thermal Infrared UAV Swarm Tracking
Yu-Hsi Chen
cs.CVcs.ROarXiv:2607.12544v12026Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking
Zekun Qi, Xuchuan Chen, Dairu Liu +10
cs.ROcs.AIcs.CVarXiv:2606.03985v12026YoCausal: How Far is Video Generation from World Model? A Causality Perspective
You-Zhe Xie, Yu-Hsuan Li, Jie-Ying Lee +3
cs.CVarXiv:2605.30346v12026InstructSAM: Segment Any Instance with Any Instructions
Yuqian Yuan, Wentong Li, Zhaocheng Li +6
cs.CVarXiv:2605.26102v22026PhysBrain 1.0 Technical Report
Shijie Lian, Bin Yu, Xiaopeng Lin +10
cs.ROcs.AIcs.CLarXiv:2605.15298v12026RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework
Hao Gao, Shaoyu Chen, Yifan Zhu +4
cs.CVarXiv:2604.15308v12026Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
Zhixiang Wei, Yi Li, Zhehan Kan +38
cs.CVarXiv:2601.19798v12026SkyReels-V3 Technique Report
Debang Li, Zhengcong Fei, Tuanhui Li +19
cs.CVarXiv:2601.17323v22026PubMed-OCR: PMC Open Access OCR Annotations
Hunter Heidenreich, Yosheb Getachew, Olivia Dinica +1
cs.CVcs.CLcs.DLarXiv:2601.11425v12026YOLO-World: Real-Time Open-Vocabulary Object Detection
Tianheng Cheng, Lin Song, Yixiao Ge +3
cs.CVarXiv:2401.17270v32024DUSt3R: Geometric 3D Vision Made Easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon +2
cs.CVarXiv:2312.14132v32023Metric3D: Towards Zero-shot Metric 3D Prediction from A Single Image
Wei Yin, Chi Zhang, Hao Chen +5
cs.CVcs.AIarXiv:2307.10984v12023EVA-CLIP: Improved Training Techniques for CLIP at Scale
Quan Sun, Yuxin Fang, Ledell Wu +2
cs.CVarXiv:2303.15389v12023Tri-Perspective View for Vision-Based 3D Semantic Occupancy Prediction
Yuanhui Huang, Wenzhao Zheng, Yunpeng Zhang +2
cs.CVcs.AIcs.LGarXiv:2302.07817v22023RTMDet: An Empirical Study of Designing Real-Time Object Detectors
Chengqi Lyu, Wenwei Zhang, Haian Huang +5
cs.CVarXiv:2212.07784v22022Magic3D: High-Resolution Text-to-3D Content Creation
Chen-Hsuan Lin, Jun Gao, Luming Tang +7
cs.CVcs.GRcs.LGarXiv:2211.10440v22022ViM: Out-Of-Distribution with Virtual-logit Matching
Haoqi Wang, Zhizhong Li, Litong Feng +1
cs.CVarXiv:2203.10807v12022Point-NeRF: Point-based Neural Radiance Fields
Qiangeng Xu, Zexiang Xu, Julien Philip +4
cs.CVarXiv:2201.08845v72022Vision Transformer with Deformable Attention
Zhuofan Xia, Xuran Pan, Shiji Song +2
cs.CVarXiv:2201.00520v32022