Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
5,821 to 5,880 of 18,817
Convolutional Bypasses Are Better Vision Transformer Adapters
Shibo Jie, Zhi-Hong Deng
cs.CVarXiv:2207.07039v32022Feature Decomposition and Reconstruction Learning for Effective Facial Expression Recognition
Delian Ruan, Yan Yan, Shenqi Lai +3
cs.CVarXiv:2104.05160v22021DragDiffusion: Harnessing Diffusion Models for Interactive Point-based Image Editing
Yujun Shi, Chuhui Xue, Jun Hao Liew +5
cs.CVcs.LGarXiv:2306.14435v62023Unified Vision-Language-Action Model
Yuqi Wang, Xinghang Li, Wenxuan Wang +5
cs.CVcs.ROarXiv:2506.19850v12025mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Qinghao Ye, Haiyang Xu, Guohai Xu +15
cs.CLcs.CVcs.LGarXiv:2304.14178v32023Interactive and Explainable Region-guided Radiology Report Generation
Tim Tanida, Philip Müller, Georgios Kaissis +1
cs.CVcs.CLcs.LGarXiv:2304.08295v12023Navigating to Objects in the Real World
Theophile Gervet, Soumith Chintala, Dhruv Batra +2
cs.ROcs.CVcs.LGarXiv:2212.00922v12022Open-vocabulary Queryable Scene Representations for Real World Planning
Boyuan Chen, Fei Xia, Brian Ichter +5
cs.ROcs.AIcs.CVarXiv:2209.09874v22022A high performance fingerprint liveness detection method based on quality related features
Javier Galbally, Fernando Alonso-Fernandez, Julian Fierrez +1
cs.CVarXiv:2111.01898v12021Latent Image Animator: Learning to Animate Images via Latent Space Navigation
Yaohui Wang, Di Yang, Francois Bremond +1
cs.CVarXiv:2203.09043v12022Self-paced Contrastive Learning with Hybrid Memory for Domain Adaptive Object Re-ID
Yixiao Ge, Feng Zhu, Dapeng Chen +2
cs.CVarXiv:2006.02713v22020Object Detection in Aerial Images: A Large-Scale Benchmark and Challenges
Jian Ding, Nan Xue, Gui-Song Xia +8
cs.CVarXiv:2102.12219v22021Deep learning on edge: extracting field boundaries from satellite images with a convolutional neural network
François Waldner, Foivos I. Diakogiannis
cs.CVarXiv:1910.12023v22019Hyperspectral Image Classification-Traditional to Deep Models: A Survey for Future Prospects
Muhammad Ahmad, Sidrah Shabbir, Swalpa Kumar Roy +7
eess.IVcs.CVarXiv:2101.06116v32021Class-incremental learning: survey and performance evaluation on image classification
Marc Masana, Xialei Liu, Bartlomiej Twardowski +3
cs.LGcs.CVarXiv:2010.15277v32020Automatic Classification of Defective Photovoltaic Module Cells in Electroluminescence Images
Sergiu Deitsch, Vincent Christlein, Stephan Berger +4
cs.CVarXiv:1807.02894v32018Grad-CAM++: Improved Visual Explanations for Deep Convolutional Networks
Aditya Chattopadhyay, Anirban Sarkar, Prantik Howlader +1
cs.CVarXiv:1710.11063v32017Backdoor Learning: A Survey
Yiming Li, Yong Jiang, Zhifeng Li +1
cs.CRcs.CVcs.LGarXiv:2007.08745v52020Blockchain-Federated-Learning and Deep Learning Models for COVID-19 detection using CT Imaging
Rajesh Kumar, Abdullah Aman Khan, Sinmin Zhang +7
eess.IVcs.CVcs.LGarXiv:2007.06537v22020Demographic Bias in Biometrics: A Survey on an Emerging Challenge
P. Drozdowski, C. Rathgeb, A. Dantcheva +2
cs.CYcs.CRcs.CVarXiv:2003.02488v22020Locality Preserving Joint Transfer for Domain Adaptation
Li Jingjing, Jing Mengmeng, Lu Ke +2
cs.CVarXiv:1906.07441v12019Locality and Structure Regularized Low Rank Representation for Hyperspectral Image Classification
Qi Wang, Xiange He, Xuelong Li
cs.CVeess.IVarXiv:1905.02488v12019FakeCatcher: Detection of Synthetic Portrait Videos using Biological Signals
Umur Aybars Ciftci, Ilke Demir
cs.CVarXiv:1901.02212v32019The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions
Philipp Tschandl, Cliff Rosendahl, Harald Kittler
cs.CVarXiv:1803.10417v32018An Introduction to Image Synthesis with Generative Adversarial Nets
He Huang, Philip S. Yu, Changhu Wang
cs.CVarXiv:1803.04469v22018MedSAM2: Segment Anything in 3D Medical Images and Videos
Jun Ma, Zongxin Yang, Sumin Kim +6
eess.IVcs.AIcs.CVarXiv:2504.03600v12025Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
Junyan Ye, Dongzhi Jiang, Zihao Wang +9
cs.CVcs.AIcs.CLarXiv:2508.09987v12025OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
Yifei Li, Junbo Niu, Ziyang Miao +12
cs.CVcs.AIarXiv:2501.05510v22025TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System
Yanjie Ze, Siheng Zhao, Weizhuo Wang +6
cs.ROcs.CVcs.LGarXiv:2511.02832v12025Crowded Scene Analysis: A Survey
Teng Li, Huan Chang, Meng Wang +3
cs.CVarXiv:1502.01812v12015Do generative video models understand physical principles?
Saman Motamed, Laura Culp, Kevin Swersky +2
cs.CVcs.AIcs.GRarXiv:2501.09038v32025A Primer on Motion Capture with Deep Learning: Principles, Pitfalls and Perspectives
Alexander Mathis, Steffen Schneider, Jessy Lauer +1
cs.CVcs.LGq-bio.NCarXiv:2009.00564v22020Investigating Bias and Fairness in Facial Expression Recognition
Tian Xu, Jennifer White, Sinan Kalkan +1
cs.CVarXiv:2007.10075v32020See What You Are Told: Visual Attention Sink in Large Multimodal Models
Seil Kang, Jinyeong Kim, Junhyeok Kim +1
cs.CVcs.AIarXiv:2503.03321v12025Counting Animals in Camera-Traps Image Sequences without Count Labels: Winning Solution to the iWildCam 2021 Challenge
Fagner Cunha, Juan G. Colonna, Eulanda M. dos Santos
cs.CVarXiv:2609.03233v12026Multi-Level Bottom-Top and Top-Bottom Feature Fusion for Crowd Counting
Vishwanath A Sindagi, Vishal M. Patel
cs.CVarXiv:1908.10937v12019VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
Runjia Li, Philip Torr, Andrea Vedaldi +1
cs.CVarXiv:2506.18903v32025Real-time Unsupervised Object Discovery from Asynchronous Event Streams
Pratham G. Shenwai, Hemant Kumar Singh, Sridhar Ravi
cs.CVarXiv:2608.26644v12026A Comparative Study of Label-free Representation Quality Metrics in Deep Learning
Daniel Richards Arputharaj, Daniel Jönsson, Gabriel Eilertsen
cs.LGcs.CVarXiv:2608.23182v12026ArtiMo: Agent-Driven Articulated Mesh Animation
Chunyu Zou, Peng Dai, Yi-Hua Huang +4
cs.CVarXiv:2608.20699v12026CRS-Bench: A Reference-Relative Reliability Benchmark for Medical Image Encoders
Xingtao Lin, Hangqi Ren, Caiwan Sun +1
eess.IVcs.AIcs.CVarXiv:2608.22059v12026RECOUNT: Reference-guided Counting with Synthetic Visual Exemplars
Adriano D'Alessandro, Ali Mahdavi-Amiri, Ghassan Hamarneh
cs.CVarXiv:2608.20621v12026Maximum Entropy Encoding of Energy-Weighted Spherical Moments
Jiaze Sun
cs.GRcs.CVarXiv:2608.20429v12026SparseGS: Sparse View Synthesis using 3D Gaussian Splatting
Haolin Xiong, Sairisheek Muttukuru, Hanyuan Xiao +4
cs.CVcs.LGeess.IVarXiv:2312.00206v42023ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Language Models
Jihae Jeong, Junha Choi, Hwanjo Yu
cs.CVcs.AIcs.CLarXiv:2608.19075v12026SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation
Keyu Tu, Zhuowei Chen, Mengqi Huang +4
cs.CVcs.AIarXiv:2608.17426v12026Harnessing Magnitude-Only and Complex Measurements for Improved Dynamic MRI Reconstruction with Learned Priors
Mahdi Saberi, Yaşar Utku Alçalar, Merve Gülle +2
eess.IVcs.AIcs.CVarXiv:2608.18036v12026Bridging the Gap between Labeled and Unlabeled Data via Unified Flow with Feature Memory Bank
Shanwen Wang, Xin Sun, Danfeng Hong +2
cs.CVcs.AIarXiv:2608.16681v12026MMVU: Measuring Expert-Level Multi-Discipline Video Understanding
Yilun Zhao, Lujing Xie, Haowei Zhang +16
cs.CVcs.AIcs.CLarXiv:2501.12380v12025Three-Body Scattering for Generative Modeling
Peng Sun, Zhenglin Cheng, Deyuan Liu +3
cs.LGcs.CVarXiv:2607.18198v12026FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Mining
Jinghong Lan, Wei Cheng, Yunuo Chen +10
cs.CVcs.AIarXiv:2606.20506v12026REBASE: Reference-Background Subspace Elimination for Training-Free In-Context Segmentation
Mantha Sai Gopal, Jaison Saji Chacko, Harsh Nandwana +3
cs.CVarXiv:2607.09082v12026Autonomous Scientific Discovery via Iterative Meta-Reflection
Bingchen Zhao, Sara Beery, Oisin Mac Aodha
cs.CVcs.AIarXiv:2607.01131v12026The FID Lottery: Quantifying Hidden Randomness in Generative-Model Evaluation
Nicolas Dufour, Alexei A. Efros, Patrick Pérez
cs.CVarXiv:2606.20536v12026SHAMISA: SHAped Modeling of Implicit Structural Associations for Self-supervised No-Reference Image Quality Assessment
Mahdi Naseri, Zhou Wang
cs.CVcs.AIcs.LGarXiv:2603.13669v22026Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs
Haozhe Zhao, Shuzheng Si, Zhenhailong Wang +6
cs.CVcs.AIcs.CLarXiv:2605.30611v12026Bernini: Latent Semantic Planning for Video Diffusion
Bernini Team, Chenchen Liu, Junyi Chen +9
cs.CVcs.AIcs.MMarXiv:2605.22344v12026Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling
Keming Wu, Zuhao Yang, Kaichen Zhang +24
cs.CVarXiv:2604.28185v22026Linking spatial biology and clinical histology via Haiku
Yan Cui, Jacob S. Leiby, Wenhui Lei +6
cs.LGcs.CVq-bio.QMarXiv:2605.00925v12026Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models
Fabian Morelli, Arnas Uselis, Ankit Sonthalia +1
cs.CVarXiv:2605.15961v12026