Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

5,821 to 5,880 of 18,817

  1. Convolutional Bypasses Are Better Vision Transformer Adapters

    Shibo Jie, Zhi-Hong Deng

    cs.CVarXiv:2207.07039v32022
  2. Feature Decomposition and Reconstruction Learning for Effective Facial Expression Recognition

    Delian Ruan, Yan Yan, Shenqi Lai +3

    cs.CVarXiv:2104.05160v22021
  3. DragDiffusion: Harnessing Diffusion Models for Interactive Point-based Image Editing

    Yujun Shi, Chuhui Xue, Jun Hao Liew +5

    cs.CVcs.LGarXiv:2306.14435v62023
  4. Unified Vision-Language-Action Model

    Yuqi Wang, Xinghang Li, Wenxuan Wang +5

    cs.CVcs.ROarXiv:2506.19850v12025
  5. mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

    Qinghao Ye, Haiyang Xu, Guohai Xu +15

    cs.CLcs.CVcs.LGarXiv:2304.14178v32023
  6. Interactive and Explainable Region-guided Radiology Report Generation

    Tim Tanida, Philip Müller, Georgios Kaissis +1

    cs.CVcs.CLcs.LGarXiv:2304.08295v12023
  7. Navigating to Objects in the Real World

    Theophile Gervet, Soumith Chintala, Dhruv Batra +2

    cs.ROcs.CVcs.LGarXiv:2212.00922v12022
  8. Open-vocabulary Queryable Scene Representations for Real World Planning

    Boyuan Chen, Fei Xia, Brian Ichter +5

    cs.ROcs.AIcs.CVarXiv:2209.09874v22022
  9. A high performance fingerprint liveness detection method based on quality related features

    Javier Galbally, Fernando Alonso-Fernandez, Julian Fierrez +1

    cs.CVarXiv:2111.01898v12021
  10. Latent Image Animator: Learning to Animate Images via Latent Space Navigation

    Yaohui Wang, Di Yang, Francois Bremond +1

    cs.CVarXiv:2203.09043v12022
  11. Self-paced Contrastive Learning with Hybrid Memory for Domain Adaptive Object Re-ID

    Yixiao Ge, Feng Zhu, Dapeng Chen +2

    cs.CVarXiv:2006.02713v22020
  12. Object Detection in Aerial Images: A Large-Scale Benchmark and Challenges

    Jian Ding, Nan Xue, Gui-Song Xia +8

    cs.CVarXiv:2102.12219v22021
  13. Deep learning on edge: extracting field boundaries from satellite images with a convolutional neural network

    François Waldner, Foivos I. Diakogiannis

    cs.CVarXiv:1910.12023v22019
  14. Hyperspectral Image Classification-Traditional to Deep Models: A Survey for Future Prospects

    Muhammad Ahmad, Sidrah Shabbir, Swalpa Kumar Roy +7

    eess.IVcs.CVarXiv:2101.06116v32021
  15. Class-incremental learning: survey and performance evaluation on image classification

    Marc Masana, Xialei Liu, Bartlomiej Twardowski +3

    cs.LGcs.CVarXiv:2010.15277v32020
  16. Automatic Classification of Defective Photovoltaic Module Cells in Electroluminescence Images

    Sergiu Deitsch, Vincent Christlein, Stephan Berger +4

    cs.CVarXiv:1807.02894v32018
  17. Grad-CAM++: Improved Visual Explanations for Deep Convolutional Networks

    Aditya Chattopadhyay, Anirban Sarkar, Prantik Howlader +1

    cs.CVarXiv:1710.11063v32017
  18. Backdoor Learning: A Survey

    Yiming Li, Yong Jiang, Zhifeng Li +1

    cs.CRcs.CVcs.LGarXiv:2007.08745v52020
  19. Blockchain-Federated-Learning and Deep Learning Models for COVID-19 detection using CT Imaging

    Rajesh Kumar, Abdullah Aman Khan, Sinmin Zhang +7

    eess.IVcs.CVcs.LGarXiv:2007.06537v22020
  20. Demographic Bias in Biometrics: A Survey on an Emerging Challenge

    P. Drozdowski, C. Rathgeb, A. Dantcheva +2

    cs.CYcs.CRcs.CVarXiv:2003.02488v22020
  21. Locality Preserving Joint Transfer for Domain Adaptation

    Li Jingjing, Jing Mengmeng, Lu Ke +2

    cs.CVarXiv:1906.07441v12019
  22. Locality and Structure Regularized Low Rank Representation for Hyperspectral Image Classification

    Qi Wang, Xiange He, Xuelong Li

    cs.CVeess.IVarXiv:1905.02488v12019
  23. FakeCatcher: Detection of Synthetic Portrait Videos using Biological Signals

    Umur Aybars Ciftci, Ilke Demir

    cs.CVarXiv:1901.02212v32019
  24. The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions

    Philipp Tschandl, Cliff Rosendahl, Harald Kittler

    cs.CVarXiv:1803.10417v32018
  25. An Introduction to Image Synthesis with Generative Adversarial Nets

    He Huang, Philip S. Yu, Changhu Wang

    cs.CVarXiv:1803.04469v22018
  26. MedSAM2: Segment Anything in 3D Medical Images and Videos

    Jun Ma, Zongxin Yang, Sumin Kim +6

    eess.IVcs.AIcs.CVarXiv:2504.03600v12025
  27. Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

    Junyan Ye, Dongzhi Jiang, Zihao Wang +9

    cs.CVcs.AIcs.CLarXiv:2508.09987v12025
  28. OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?

    Yifei Li, Junbo Niu, Ziyang Miao +12

    cs.CVcs.AIarXiv:2501.05510v22025
  29. TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System

    Yanjie Ze, Siheng Zhao, Weizhuo Wang +6

    cs.ROcs.CVcs.LGarXiv:2511.02832v12025
  30. Crowded Scene Analysis: A Survey

    Teng Li, Huan Chang, Meng Wang +3

    cs.CVarXiv:1502.01812v12015
  31. Do generative video models understand physical principles?

    Saman Motamed, Laura Culp, Kevin Swersky +2

    cs.CVcs.AIcs.GRarXiv:2501.09038v32025
  32. A Primer on Motion Capture with Deep Learning: Principles, Pitfalls and Perspectives

    Alexander Mathis, Steffen Schneider, Jessy Lauer +1

    cs.CVcs.LGq-bio.NCarXiv:2009.00564v22020
  33. Investigating Bias and Fairness in Facial Expression Recognition

    Tian Xu, Jennifer White, Sinan Kalkan +1

    cs.CVarXiv:2007.10075v32020
  34. See What You Are Told: Visual Attention Sink in Large Multimodal Models

    Seil Kang, Jinyeong Kim, Junhyeok Kim +1

    cs.CVcs.AIarXiv:2503.03321v12025
  35. Counting Animals in Camera-Traps Image Sequences without Count Labels: Winning Solution to the iWildCam 2021 Challenge

    Fagner Cunha, Juan G. Colonna, Eulanda M. dos Santos

    cs.CVarXiv:2609.03233v12026
  36. Multi-Level Bottom-Top and Top-Bottom Feature Fusion for Crowd Counting

    Vishwanath A Sindagi, Vishal M. Patel

    cs.CVarXiv:1908.10937v12019
  37. VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory

    Runjia Li, Philip Torr, Andrea Vedaldi +1

    cs.CVarXiv:2506.18903v32025
  38. Real-time Unsupervised Object Discovery from Asynchronous Event Streams

    Pratham G. Shenwai, Hemant Kumar Singh, Sridhar Ravi

    cs.CVarXiv:2608.26644v12026
  39. A Comparative Study of Label-free Representation Quality Metrics in Deep Learning

    Daniel Richards Arputharaj, Daniel Jönsson, Gabriel Eilertsen

    cs.LGcs.CVarXiv:2608.23182v12026
  40. ArtiMo: Agent-Driven Articulated Mesh Animation

    Chunyu Zou, Peng Dai, Yi-Hua Huang +4

    cs.CVarXiv:2608.20699v12026
  41. CRS-Bench: A Reference-Relative Reliability Benchmark for Medical Image Encoders

    Xingtao Lin, Hangqi Ren, Caiwan Sun +1

    eess.IVcs.AIcs.CVarXiv:2608.22059v12026
  42. RECOUNT: Reference-guided Counting with Synthetic Visual Exemplars

    Adriano D'Alessandro, Ali Mahdavi-Amiri, Ghassan Hamarneh

    cs.CVarXiv:2608.20621v12026
  43. Maximum Entropy Encoding of Energy-Weighted Spherical Moments

    Jiaze Sun

    cs.GRcs.CVarXiv:2608.20429v12026
  44. SparseGS: Sparse View Synthesis using 3D Gaussian Splatting

    Haolin Xiong, Sairisheek Muttukuru, Hanyuan Xiao +4

    cs.CVcs.LGeess.IVarXiv:2312.00206v42023
  45. ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Language Models

    Jihae Jeong, Junha Choi, Hwanjo Yu

    cs.CVcs.AIcs.CLarXiv:2608.19075v12026
  46. SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

    Keyu Tu, Zhuowei Chen, Mengqi Huang +4

    cs.CVcs.AIarXiv:2608.17426v12026
  47. Harnessing Magnitude-Only and Complex Measurements for Improved Dynamic MRI Reconstruction with Learned Priors

    Mahdi Saberi, Yaşar Utku Alçalar, Merve Gülle +2

    eess.IVcs.AIcs.CVarXiv:2608.18036v12026
  48. Bridging the Gap between Labeled and Unlabeled Data via Unified Flow with Feature Memory Bank

    Shanwen Wang, Xin Sun, Danfeng Hong +2

    cs.CVcs.AIarXiv:2608.16681v12026
  49. MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

    Yilun Zhao, Lujing Xie, Haowei Zhang +16

    cs.CVcs.AIcs.CLarXiv:2501.12380v12025
  50. Three-Body Scattering for Generative Modeling

    Peng Sun, Zhenglin Cheng, Deyuan Liu +3

    cs.LGcs.CVarXiv:2607.18198v12026
  51. FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Mining

    Jinghong Lan, Wei Cheng, Yunuo Chen +10

    cs.CVcs.AIarXiv:2606.20506v12026
  52. REBASE: Reference-Background Subspace Elimination for Training-Free In-Context Segmentation

    Mantha Sai Gopal, Jaison Saji Chacko, Harsh Nandwana +3

    cs.CVarXiv:2607.09082v12026
  53. Autonomous Scientific Discovery via Iterative Meta-Reflection

    Bingchen Zhao, Sara Beery, Oisin Mac Aodha

    cs.CVcs.AIarXiv:2607.01131v12026
  54. The FID Lottery: Quantifying Hidden Randomness in Generative-Model Evaluation

    Nicolas Dufour, Alexei A. Efros, Patrick Pérez

    cs.CVarXiv:2606.20536v12026
  55. SHAMISA: SHAped Modeling of Implicit Structural Associations for Self-supervised No-Reference Image Quality Assessment

    Mahdi Naseri, Zhou Wang

    cs.CVcs.AIcs.LGarXiv:2603.13669v22026
  56. Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs

    Haozhe Zhao, Shuzheng Si, Zhenhailong Wang +6

    cs.CVcs.AIcs.CLarXiv:2605.30611v12026
  57. Bernini: Latent Semantic Planning for Video Diffusion

    Bernini Team, Chenchen Liu, Junyi Chen +9

    cs.CVcs.AIcs.MMarXiv:2605.22344v12026
  58. Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling

    Keming Wu, Zuhao Yang, Kaichen Zhang +24

    cs.CVarXiv:2604.28185v22026
  59. Linking spatial biology and clinical histology via Haiku

    Yan Cui, Jacob S. Leiby, Wenhui Lei +6

    cs.LGcs.CVq-bio.QMarXiv:2605.00925v12026
  60. Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models

    Fabian Morelli, Arnas Uselis, Ankit Sonthalia +1

    cs.CVarXiv:2605.15961v12026