Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
4,681 to 4,740 of 18,795
GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding
Fei Tang, Zhangxuan Gu, Zhengxi Lu +9
cs.LGcs.AIcs.CLarXiv:2507.15846v32025Automatic Detection of Solar Photovoltaic Arrays in High Resolution Aerial Imagery
Jordan M. Malof, Kyle Bradbury, Leslie M. Collins +1
cs.CVarXiv:1607.06029v12016Frequency Dynamic Convolution for Dense Image Prediction
Linwei Chen, Lin Gu, Liang Li +2
cs.CVcs.AIarXiv:2503.18783v22025How I Warped Your Noise: a Temporally-Correlated Noise Prior for Diffusion Models
Pascal Chang, Jingwei Tang, Markus Gross +1
cs.CVcs.LGarXiv:2504.03072v12025Self-Supervised Learning for Domain Adaptation on Point-Clouds
Idan Achituve, Haggai Maron, Gal Chechik
cs.CVcs.LGeess.IVarXiv:2003.12641v52020BlobBoards: Robust Markers for Accurate Pose
James Pritts, Till Sittart, Hendrik Sauer +4
cs.CVarXiv:2608.28830v12026A Dataset and Benchmark for Large-scale Multi-modal Face Anti-spoofing
Shifeng Zhang, Xiaobo Wang, Ajian Liu +6
cs.CVarXiv:1812.00408v32018AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors
Ruoxuan Feng, Jiangyu Hu, Wenke Xia +5
cs.LGcs.CVcs.ROarXiv:2502.12191v32025Back to MLP: A Simple Baseline for Human Motion Prediction
Wen Guo, Yuming Du, Xi Shen +3
cs.CVcs.AIarXiv:2207.01567v32022Self-Supervised Scene De-occlusion
Xiaohang Zhan, Xingang Pan, Bo Dai +3
cs.CVarXiv:2004.02788v12020MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation
Henghui Ding, Chang Liu, Shuting He +4
cs.CVarXiv:2512.10945v12025Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene Priors
Wonbong Jang, Philippe Weinzaepfel, Vincent Leroy +2
cs.CVarXiv:2503.17316v12025GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices
Quanfeng Lu, Wenqi Shao, Zitao Liu +7
cs.CVarXiv:2406.08451v22024nnInteractive: Redefining 3D Promptable Segmentation
Fabian Isensee, Maximilian Rokuss, Lars Krämer +10
cs.CVarXiv:2503.08373v12025CompletionFormer: Depth Completion with Convolutions and Vision Transformers
Zhang Youmin, Guo Xianda, Poggi Matteo +3
cs.CVarXiv:2304.13030v12023VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
Hengtao Li, Pengxiang Ding, Runze Suo +8
cs.ROcs.CVarXiv:2510.00406v12025UAV-DETR: Efficient End-to-End Object Detection for Unmanned Aerial Vehicle Imagery
Huaxiang Zhang, Kai Liu, Zhongxue Gan +1
cs.CVarXiv:2501.01855v32025Generative Adversarial Networks for Image and Video Synthesis: Algorithms and Applications
Ming-Yu Liu, Xun Huang, Jiahui Yu +2
cs.CVarXiv:2008.02793v22020LiveCap: Real-time Human Performance Capture from Monocular Video
Marc Habermann, Weipeng Xu, Michael Zollhoefer +2
cs.CVarXiv:1810.02648v32018Hybrid Approach of Relation Network and Localized Graph Convolutional Filtering for Breast Cancer Subtype Classification
Sungmin Rhee, Seokjun Seo, Sun Kim
cs.CVcs.LGarXiv:1711.05859v32017Deep Learning Inversion of Electrical Resistivity Data
Bin Liu, Qian Guo, Shucai Li +6
cs.CVcs.AIphysics.geo-pharXiv:1904.05265v22019Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
Zhiqi Li, Guo Chen, Shilong Liu +24
cs.CVcs.AIcs.LGarXiv:2501.14818v12025MASQ: Mask-Aware Spatiotemporal Quantization for Unsupervised Skeleton Action Segmentation
Xinyao Qin, Linxiang Peng, Youbao Ye +2
cs.CVarXiv:2608.29891v12026Tracing Generated Samples to Training-Data Clusters in Flow-Matching Models
Rania Briq, Ohad Fried, Michael Kamp +1
cs.LGcs.CVarXiv:2608.30081v22026A Transfer-Learning Approach for Accelerated MRI using Deep Neural Networks
Salman Ul Hassan Dar, Muzaffer Özbey, Ahmet Burak Çatlı +1
cs.CVarXiv:1710.02615v32017DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving
Anqing Jiang, Yu Gao, Zhigang Sun +11
cs.AIcs.CVcs.ROarXiv:2505.19381v42025NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos
Yuxue Yang, Lue Fan, Ziqi Shi +3
cs.CVarXiv:2601.00393v22026OPAL: Orthonormal Prototype Alignment Learning for Interpretable Image Classification
Ilán Carretero, Gustavo Jesús Angulo, Rocío del Amor +1
cs.CVarXiv:2608.30003v12026FIS-OT: Feature-Induced Optimal Transport for Unsupervised Action Segmentation
Linxiang Peng, Xinyao Qin, Jinhan Li +2
cs.CVarXiv:2608.29980v12026Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models
Thomas Fel, Ekdeep Singh Lubana, Jacob S. Prince +7
cs.CVarXiv:2502.12892v22025RIDGE: Region-Informed Derivative-Guided Evidence Selection for Long Video Understanding
Shanqing Xu, Meng Luo, Mengchen Qian +7
cs.CVarXiv:2608.29958v12026V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
Zixu Cheng, Jian Hu, Ziquan Liu +3
cs.CVarXiv:2503.11495v12025Dior: Drawing the Light of Image via Material-Decoupled Illumination Representation
Xuanpu Zhang, Xuesong Niu, Haoxiang Cao +3
cs.CVarXiv:2608.29925v12026PET image denoising based on denoising diffusion probabilistic models
Kuang Gong, Keith A. Johnson, Georges El Fakhri +2
eess.IVcs.CVphysics.med-pharXiv:2209.06167v22022VOID: Video Object and Interaction Deletion
Saman Motamed, William Harvey, Benjamin Klein +3
cs.CVcs.AIarXiv:2604.02296v12026VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking
Limin Wang, Bingkun Huang, Zhiyu Zhao +5
cs.CVcs.LGarXiv:2303.16727v22023TrivialAugment: Tuning-free Yet State-of-the-Art Data Augmentation
Samuel G. Müller, Frank Hutter
cs.CVcs.LGarXiv:2103.10158v22021Confidence Regularized Self-Training
Yang Zou, Zhiding Yu, Xiaofeng Liu +2
cs.CVcs.LGcs.MMarXiv:1908.09822v32019End-to-End Learning of Representations for Asynchronous Event-Based Data
Daniel Gehrig, Antonio Loquercio, Konstantinos G. Derpanis +1
cs.CVarXiv:1904.08245v42019Recurrent Back-Projection Network for Video Super-Resolution
Muhammad Haris, Greg Shakhnarovich, Norimichi Ukita
cs.CVarXiv:1903.10128v12019MILD-Net: Minimal Information Loss Dilated Network for Gland Instance Segmentation in Colon Histology Images
Simon Graham, Hao Chen, Jevgenij Gamper +5
cs.CVarXiv:1806.01963v42018Unified Panoramic Geometry Estimation via Multi-View Foundation Models
Vukasin Bozic, Isidora Slavkovic, Dominik Narnhofer +4
cs.CVcs.AIarXiv:2605.26368v22026The GAN is dead; long live the GAN! A Modern GAN Baseline
Yiwen Huang, Aaron Gokaslan, Volodymyr Kuleshov +1
cs.LGcs.CVarXiv:2501.05441v12025Unsupervised Learning of a Hierarchical Spiking Neural Network for Optical Flow Estimation: From Events to Global Motion Perception
Federico Paredes-Vallés, Kirk Y. W. Scheper, Guido C. H. E. de Croon
cs.CVarXiv:1807.10936v22018A2DINOv3: Rethinking Multi-Modal Object Detection via Socialized Collaboration
Jiekang Feng, Zhihe Fan, Yunqi Zhu +5
cs.CVcs.AIarXiv:2608.21099v12026Panoptic Pairwise Distortion Graph
Muhammad Kamran Janjua, Abdul Wahab, Bahador Rashidi
cs.CVcs.AIcs.LGarXiv:2604.11004v12026Solar Cell Surface Defect Inspection Based on Multispectral Convolutional Neural Network
Haiyong Chen, Yue Pang, Qidi Hu +1
cs.CVeess.IVarXiv:1812.06220v12018Revisiting Point Cloud Shape Classification with a Simple and Effective Baseline
Ankit Goyal, Hei Law, Bowei Liu +2
cs.CVcs.LGarXiv:2106.05304v12021Attention-Based Deep Neural Networks for Detection of Cancerous and Precancerous Esophagus Tissue on Histopathological Slides
Naofumi Tomita, Behnaz Abdollahi, Jason Wei +3
eess.IVcs.CVarXiv:1811.08513v22018RASID: A Robust WLAN Device-free Passive Motion Detection System
Ahmed E. Kosba, Ahmed Saeed, Moustafa Youssef
cs.NIcs.CVarXiv:1105.6084v22011End-to-End Multimodal Emotion Recognition using Deep Neural Networks
Panagiotis Tzirakis, George Trigeorgis, Mihalis A. Nicolaou +2
cs.CVcs.CLarXiv:1704.08619v12017Lossy Image Compression with Compressive Autoencoders
Lucas Theis, Wenzhe Shi, Andrew Cunningham +1
stat.MLcs.CVarXiv:1703.00395v12017BiosecurID: a multimodal biometric database
Julian Fierrez, Javier Galbally, Javier Ortega-Garcia +22
cs.CRcs.CVeess.IVarXiv:2111.03472v12021Anatomy-specific classification of medical images using deep convolutional nets
Holger R. Roth, Christopher T. Lee, Hoo-Chang Shin +5
cs.CVarXiv:1504.04003v12015YOLOE: Real-Time Seeing Anything
Ao Wang, Lihao Liu, Hui Chen +3
cs.CVarXiv:2503.07465v22025Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning
Yue Ma, Yulong Liu, Qiyuan Zhu +8
cs.CVarXiv:2506.05207v42025Pedestrian Archetypes Extension -- More Pedestrian Models for Autonomous Vehicle Safety Testing
Taorui Huang, Namita Gaidhani, Ritvik Bansal +6
cs.CVarXiv:2607.16922v12026Towards Automatic Threat Detection: A Survey of Advances of Deep Learning within X-ray Security Imaging
Samet Akcay, Toby Breckon
cs.CVarXiv:2001.01293v22020Lossy Event Compression: From Event Stream Distortion to Task Performance
Zahra Rezaee, Catarina Brites, João Ascenso
cs.CVeess.IVarXiv:2608.28429v12026Multi-Scale Temporal Domain Alignment for Federated Video Domain Adaptation
Lee En-Yi Hannah, Haozhi Cao, Yuecong Xu
cs.CVarXiv:2608.29186v12026