Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

4,681 to 4,740 of 18,795

  1. GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding

    Fei Tang, Zhangxuan Gu, Zhengxi Lu +9

    cs.LGcs.AIcs.CLarXiv:2507.15846v32025
  2. Automatic Detection of Solar Photovoltaic Arrays in High Resolution Aerial Imagery

    Jordan M. Malof, Kyle Bradbury, Leslie M. Collins +1

    cs.CVarXiv:1607.06029v12016
  3. Frequency Dynamic Convolution for Dense Image Prediction

    Linwei Chen, Lin Gu, Liang Li +2

    cs.CVcs.AIarXiv:2503.18783v22025
  4. How I Warped Your Noise: a Temporally-Correlated Noise Prior for Diffusion Models

    Pascal Chang, Jingwei Tang, Markus Gross +1

    cs.CVcs.LGarXiv:2504.03072v12025
  5. Self-Supervised Learning for Domain Adaptation on Point-Clouds

    Idan Achituve, Haggai Maron, Gal Chechik

    cs.CVcs.LGeess.IVarXiv:2003.12641v52020
  6. BlobBoards: Robust Markers for Accurate Pose

    James Pritts, Till Sittart, Hendrik Sauer +4

    cs.CVarXiv:2608.28830v12026
  7. A Dataset and Benchmark for Large-scale Multi-modal Face Anti-spoofing

    Shifeng Zhang, Xiaobo Wang, Ajian Liu +6

    cs.CVarXiv:1812.00408v32018
  8. AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors

    Ruoxuan Feng, Jiangyu Hu, Wenke Xia +5

    cs.LGcs.CVcs.ROarXiv:2502.12191v32025
  9. Back to MLP: A Simple Baseline for Human Motion Prediction

    Wen Guo, Yuming Du, Xi Shen +3

    cs.CVcs.AIarXiv:2207.01567v32022
  10. Self-Supervised Scene De-occlusion

    Xiaohang Zhan, Xingang Pan, Bo Dai +3

    cs.CVarXiv:2004.02788v12020
  11. MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation

    Henghui Ding, Chang Liu, Shuting He +4

    cs.CVarXiv:2512.10945v12025
  12. Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene Priors

    Wonbong Jang, Philippe Weinzaepfel, Vincent Leroy +2

    cs.CVarXiv:2503.17316v12025
  13. GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices

    Quanfeng Lu, Wenqi Shao, Zitao Liu +7

    cs.CVarXiv:2406.08451v22024
  14. nnInteractive: Redefining 3D Promptable Segmentation

    Fabian Isensee, Maximilian Rokuss, Lars Krämer +10

    cs.CVarXiv:2503.08373v12025
  15. CompletionFormer: Depth Completion with Convolutions and Vision Transformers

    Zhang Youmin, Guo Xianda, Poggi Matteo +3

    cs.CVarXiv:2304.13030v12023
  16. VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators

    Hengtao Li, Pengxiang Ding, Runze Suo +8

    cs.ROcs.CVarXiv:2510.00406v12025
  17. UAV-DETR: Efficient End-to-End Object Detection for Unmanned Aerial Vehicle Imagery

    Huaxiang Zhang, Kai Liu, Zhongxue Gan +1

    cs.CVarXiv:2501.01855v32025
  18. Generative Adversarial Networks for Image and Video Synthesis: Algorithms and Applications

    Ming-Yu Liu, Xun Huang, Jiahui Yu +2

    cs.CVarXiv:2008.02793v22020
  19. LiveCap: Real-time Human Performance Capture from Monocular Video

    Marc Habermann, Weipeng Xu, Michael Zollhoefer +2

    cs.CVarXiv:1810.02648v32018
  20. Hybrid Approach of Relation Network and Localized Graph Convolutional Filtering for Breast Cancer Subtype Classification

    Sungmin Rhee, Seokjun Seo, Sun Kim

    cs.CVcs.LGarXiv:1711.05859v32017
  21. Deep Learning Inversion of Electrical Resistivity Data

    Bin Liu, Qian Guo, Shucai Li +6

    cs.CVcs.AIphysics.geo-pharXiv:1904.05265v22019
  22. Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

    Zhiqi Li, Guo Chen, Shilong Liu +24

    cs.CVcs.AIcs.LGarXiv:2501.14818v12025
  23. MASQ: Mask-Aware Spatiotemporal Quantization for Unsupervised Skeleton Action Segmentation

    Xinyao Qin, Linxiang Peng, Youbao Ye +2

    cs.CVarXiv:2608.29891v12026
  24. Tracing Generated Samples to Training-Data Clusters in Flow-Matching Models

    Rania Briq, Ohad Fried, Michael Kamp +1

    cs.LGcs.CVarXiv:2608.30081v22026
  25. A Transfer-Learning Approach for Accelerated MRI using Deep Neural Networks

    Salman Ul Hassan Dar, Muzaffer Özbey, Ahmet Burak Çatlı +1

    cs.CVarXiv:1710.02615v32017
  26. DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving

    Anqing Jiang, Yu Gao, Zhigang Sun +11

    cs.AIcs.CVcs.ROarXiv:2505.19381v42025
  27. NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos

    Yuxue Yang, Lue Fan, Ziqi Shi +3

    cs.CVarXiv:2601.00393v22026
  28. OPAL: Orthonormal Prototype Alignment Learning for Interpretable Image Classification

    Ilán Carretero, Gustavo Jesús Angulo, Rocío del Amor +1

    cs.CVarXiv:2608.30003v12026
  29. FIS-OT: Feature-Induced Optimal Transport for Unsupervised Action Segmentation

    Linxiang Peng, Xinyao Qin, Jinhan Li +2

    cs.CVarXiv:2608.29980v12026
  30. Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models

    Thomas Fel, Ekdeep Singh Lubana, Jacob S. Prince +7

    cs.CVarXiv:2502.12892v22025
  31. RIDGE: Region-Informed Derivative-Guided Evidence Selection for Long Video Understanding

    Shanqing Xu, Meng Luo, Mengchen Qian +7

    cs.CVarXiv:2608.29958v12026
  32. V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning

    Zixu Cheng, Jian Hu, Ziquan Liu +3

    cs.CVarXiv:2503.11495v12025
  33. Dior: Drawing the Light of Image via Material-Decoupled Illumination Representation

    Xuanpu Zhang, Xuesong Niu, Haoxiang Cao +3

    cs.CVarXiv:2608.29925v12026
  34. PET image denoising based on denoising diffusion probabilistic models

    Kuang Gong, Keith A. Johnson, Georges El Fakhri +2

    eess.IVcs.CVphysics.med-pharXiv:2209.06167v22022
  35. VOID: Video Object and Interaction Deletion

    Saman Motamed, William Harvey, Benjamin Klein +3

    cs.CVcs.AIarXiv:2604.02296v12026
  36. VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking

    Limin Wang, Bingkun Huang, Zhiyu Zhao +5

    cs.CVcs.LGarXiv:2303.16727v22023
  37. TrivialAugment: Tuning-free Yet State-of-the-Art Data Augmentation

    Samuel G. Müller, Frank Hutter

    cs.CVcs.LGarXiv:2103.10158v22021
  38. Confidence Regularized Self-Training

    Yang Zou, Zhiding Yu, Xiaofeng Liu +2

    cs.CVcs.LGcs.MMarXiv:1908.09822v32019
  39. End-to-End Learning of Representations for Asynchronous Event-Based Data

    Daniel Gehrig, Antonio Loquercio, Konstantinos G. Derpanis +1

    cs.CVarXiv:1904.08245v42019
  40. Recurrent Back-Projection Network for Video Super-Resolution

    Muhammad Haris, Greg Shakhnarovich, Norimichi Ukita

    cs.CVarXiv:1903.10128v12019
  41. MILD-Net: Minimal Information Loss Dilated Network for Gland Instance Segmentation in Colon Histology Images

    Simon Graham, Hao Chen, Jevgenij Gamper +5

    cs.CVarXiv:1806.01963v42018
  42. Unified Panoramic Geometry Estimation via Multi-View Foundation Models

    Vukasin Bozic, Isidora Slavkovic, Dominik Narnhofer +4

    cs.CVcs.AIarXiv:2605.26368v22026
  43. The GAN is dead; long live the GAN! A Modern GAN Baseline

    Yiwen Huang, Aaron Gokaslan, Volodymyr Kuleshov +1

    cs.LGcs.CVarXiv:2501.05441v12025
  44. Unsupervised Learning of a Hierarchical Spiking Neural Network for Optical Flow Estimation: From Events to Global Motion Perception

    Federico Paredes-Vallés, Kirk Y. W. Scheper, Guido C. H. E. de Croon

    cs.CVarXiv:1807.10936v22018
  45. A2DINOv3: Rethinking Multi-Modal Object Detection via Socialized Collaboration

    Jiekang Feng, Zhihe Fan, Yunqi Zhu +5

    cs.CVcs.AIarXiv:2608.21099v12026
  46. Panoptic Pairwise Distortion Graph

    Muhammad Kamran Janjua, Abdul Wahab, Bahador Rashidi

    cs.CVcs.AIcs.LGarXiv:2604.11004v12026
  47. Solar Cell Surface Defect Inspection Based on Multispectral Convolutional Neural Network

    Haiyong Chen, Yue Pang, Qidi Hu +1

    cs.CVeess.IVarXiv:1812.06220v12018
  48. Revisiting Point Cloud Shape Classification with a Simple and Effective Baseline

    Ankit Goyal, Hei Law, Bowei Liu +2

    cs.CVcs.LGarXiv:2106.05304v12021
  49. Attention-Based Deep Neural Networks for Detection of Cancerous and Precancerous Esophagus Tissue on Histopathological Slides

    Naofumi Tomita, Behnaz Abdollahi, Jason Wei +3

    eess.IVcs.CVarXiv:1811.08513v22018
  50. RASID: A Robust WLAN Device-free Passive Motion Detection System

    Ahmed E. Kosba, Ahmed Saeed, Moustafa Youssef

    cs.NIcs.CVarXiv:1105.6084v22011
  51. End-to-End Multimodal Emotion Recognition using Deep Neural Networks

    Panagiotis Tzirakis, George Trigeorgis, Mihalis A. Nicolaou +2

    cs.CVcs.CLarXiv:1704.08619v12017
  52. Lossy Image Compression with Compressive Autoencoders

    Lucas Theis, Wenzhe Shi, Andrew Cunningham +1

    stat.MLcs.CVarXiv:1703.00395v12017
  53. BiosecurID: a multimodal biometric database

    Julian Fierrez, Javier Galbally, Javier Ortega-Garcia +22

    cs.CRcs.CVeess.IVarXiv:2111.03472v12021
  54. Anatomy-specific classification of medical images using deep convolutional nets

    Holger R. Roth, Christopher T. Lee, Hoo-Chang Shin +5

    cs.CVarXiv:1504.04003v12015
  55. YOLOE: Real-Time Seeing Anything

    Ao Wang, Lihao Liu, Hui Chen +3

    cs.CVarXiv:2503.07465v22025
  56. Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning

    Yue Ma, Yulong Liu, Qiyuan Zhu +8

    cs.CVarXiv:2506.05207v42025
  57. Pedestrian Archetypes Extension -- More Pedestrian Models for Autonomous Vehicle Safety Testing

    Taorui Huang, Namita Gaidhani, Ritvik Bansal +6

    cs.CVarXiv:2607.16922v12026
  58. Towards Automatic Threat Detection: A Survey of Advances of Deep Learning within X-ray Security Imaging

    Samet Akcay, Toby Breckon

    cs.CVarXiv:2001.01293v22020
  59. Lossy Event Compression: From Event Stream Distortion to Task Performance

    Zahra Rezaee, Catarina Brites, João Ascenso

    cs.CVeess.IVarXiv:2608.28429v12026
  60. Multi-Scale Temporal Domain Alignment for Federated Video Domain Adaptation

    Lee En-Yi Hannah, Haozhi Cao, Yuecong Xu

    cs.CVarXiv:2608.29186v12026