Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
4,741 to 4,800 of 18,815
MILD-Net: Minimal Information Loss Dilated Network for Gland Instance Segmentation in Colon Histology Images
Simon Graham, Hao Chen, Jevgenij Gamper +5
cs.CVarXiv:1806.01963v42018Unified Panoramic Geometry Estimation via Multi-View Foundation Models
Vukasin Bozic, Isidora Slavkovic, Dominik Narnhofer +4
cs.CVcs.AIarXiv:2605.26368v22026The GAN is dead; long live the GAN! A Modern GAN Baseline
Yiwen Huang, Aaron Gokaslan, Volodymyr Kuleshov +1
cs.LGcs.CVarXiv:2501.05441v12025Unsupervised Learning of a Hierarchical Spiking Neural Network for Optical Flow Estimation: From Events to Global Motion Perception
Federico Paredes-Vallés, Kirk Y. W. Scheper, Guido C. H. E. de Croon
cs.CVarXiv:1807.10936v22018A2DINOv3: Rethinking Multi-Modal Object Detection via Socialized Collaboration
Jiekang Feng, Zhihe Fan, Yunqi Zhu +5
cs.CVcs.AIarXiv:2608.21099v12026Panoptic Pairwise Distortion Graph
Muhammad Kamran Janjua, Abdul Wahab, Bahador Rashidi
cs.CVcs.AIcs.LGarXiv:2604.11004v12026Solar Cell Surface Defect Inspection Based on Multispectral Convolutional Neural Network
Haiyong Chen, Yue Pang, Qidi Hu +1
cs.CVeess.IVarXiv:1812.06220v12018Revisiting Point Cloud Shape Classification with a Simple and Effective Baseline
Ankit Goyal, Hei Law, Bowei Liu +2
cs.CVcs.LGarXiv:2106.05304v12021Attention-Based Deep Neural Networks for Detection of Cancerous and Precancerous Esophagus Tissue on Histopathological Slides
Naofumi Tomita, Behnaz Abdollahi, Jason Wei +3
eess.IVcs.CVarXiv:1811.08513v22018RASID: A Robust WLAN Device-free Passive Motion Detection System
Ahmed E. Kosba, Ahmed Saeed, Moustafa Youssef
cs.NIcs.CVarXiv:1105.6084v22011End-to-End Multimodal Emotion Recognition using Deep Neural Networks
Panagiotis Tzirakis, George Trigeorgis, Mihalis A. Nicolaou +2
cs.CVcs.CLarXiv:1704.08619v12017Lossy Image Compression with Compressive Autoencoders
Lucas Theis, Wenzhe Shi, Andrew Cunningham +1
stat.MLcs.CVarXiv:1703.00395v12017BiosecurID: a multimodal biometric database
Julian Fierrez, Javier Galbally, Javier Ortega-Garcia +22
cs.CRcs.CVeess.IVarXiv:2111.03472v12021Anatomy-specific classification of medical images using deep convolutional nets
Holger R. Roth, Christopher T. Lee, Hoo-Chang Shin +5
cs.CVarXiv:1504.04003v12015YOLOE: Real-Time Seeing Anything
Ao Wang, Lihao Liu, Hui Chen +3
cs.CVarXiv:2503.07465v22025Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning
Yue Ma, Yulong Liu, Qiyuan Zhu +8
cs.CVarXiv:2506.05207v42025Pedestrian Archetypes Extension -- More Pedestrian Models for Autonomous Vehicle Safety Testing
Taorui Huang, Namita Gaidhani, Ritvik Bansal +6
cs.CVarXiv:2607.16922v12026Towards Automatic Threat Detection: A Survey of Advances of Deep Learning within X-ray Security Imaging
Samet Akcay, Toby Breckon
cs.CVarXiv:2001.01293v22020Lossy Event Compression: From Event Stream Distortion to Task Performance
Zahra Rezaee, Catarina Brites, João Ascenso
cs.CVeess.IVarXiv:2608.28429v12026Multi-Scale Temporal Domain Alignment for Federated Video Domain Adaptation
Lee En-Yi Hannah, Haozhi Cao, Yuecong Xu
cs.CVarXiv:2608.29186v12026Object Detection Under Rainy Conditions for Autonomous Vehicles: A Review of State-of-the-Art and Emerging Techniques
Mazin Hnewa, Hayder Radha
cs.CVarXiv:2006.16471v42020CNN-based Density Estimation and Crowd Counting: A Survey
Guangshuai Gao, Junyu Gao, Qingjie Liu +2
cs.CVarXiv:2003.12783v12020Sim-to-Real Reinforcement Learning for Vision-Based Dexterous Manipulation on Humanoids
Toru Lin, Kartik Sachdev, Linxi Fan +2
cs.ROcs.AIcs.CVarXiv:2502.20396v22025SV4D 2.0: Enhancing Spatio-Temporal Consistency in Multi-View Video Diffusion for High-Quality 4D Generation
Chun-Han Yao, Yiming Xie, Vikram Voleti +2
cs.CVarXiv:2503.16396v32025LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Zuhao Yang, Sudong Wang, Kaichen Zhang +8
cs.CVarXiv:2511.20785v32025Multiple Object Tracking with Correlation Learning
Qiang Wang, Yun Zheng, Pan Pan +1
cs.CVarXiv:2104.03541v12021CAD-Llama: Leveraging Large Language Models for Computer-Aided Design Parametric 3D Model Generation
Jiahao Li, Weijian Ma, Xueyang Li +3
cs.CVarXiv:2505.04481v22025When Does Self-supervision Improve Few-shot Learning?
Jong-Chyi Su, Subhransu Maji, Bharath Hariharan
cs.CVcs.LGarXiv:1910.03560v22019Guardrail-Agnostic Societal Bias Evaluation in Large Vision-Language Models
Yusuke Hirota, Michael Ross Boone, Arun George Zachariah +4
cs.CVarXiv:2608.29590v12026HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation
Zunnan Xu, Zhentao Yu, Zixiang Zhou +10
cs.CVarXiv:2503.18860v22025Meta CLIP 2: A Worldwide Scaling Recipe
Yung-Sung Chuang, Yang Li, Dong Wang +13
cs.CVcs.CLarXiv:2507.22062v32025AssemblyNet: A large ensemble of CNNs for 3D Whole Brain MRI Segmentation
Pierrick Coupé, Boris Mansencal, Michaël Clément +5
eess.IVcs.CVcs.LGarXiv:1911.09098v12019Generative Feature Replay For Class-Incremental Learning
Xialei Liu, Chenshen Wu, Mikel Menta +5
cs.CVcs.LGarXiv:2004.09199v12020Med3DVLM: An Efficient Vision-Language Model for 3D Medical Image Analysis
Yu Xin, Gorkem Can Ates, Kuang Gong +1
cs.CVeess.IVarXiv:2503.20047v32025Input-Adaptive Gating of a Dehazing Front-End for On-Device Perception in Smoke-Obscured Environments
Seongjun Kang, Ishaan Garg, Vishnu Bharadwaj
cs.CVarXiv:2608.30034v12026LSNet: See Large, Focus Small
Ao Wang, Hui Chen, Zijia Lin +2
cs.CVarXiv:2503.23135v12025BSNet: Bi-Similarity Network for Few-shot Fine-grained Image Classification
Xiaoxu Li, Jijie Wu, Zhuo Sun +3
cs.CVarXiv:2011.14311v12020VLT: Vision-Language Transformer and Query Generation for Referring Segmentation
Henghui Ding, Chang Liu, Suchen Wang +1
cs.CVarXiv:2210.15871v12022DiffusionRenderer: Neural Inverse and Forward Rendering with Video Diffusion Models
Ruofan Liang, Zan Gojcic, Huan Ling +8
cs.CVcs.GRarXiv:2501.18590v22025RemoteSAM: Towards Segment Anything for Earth Observation
Liang Yao, Fan Liu, Delong Chen +6
cs.CVarXiv:2505.18022v32025Local Implicit Grid Representations for 3D Scenes
Chiyu Max Jiang, Avneesh Sud, Ameesh Makadia +3
cs.CVcs.CGcs.LGarXiv:2003.08981v12020RAFT-DVC: Resolution-Aware Machine Learning-Based Digital Volume Correlation
Zixiang Tong, Lehu Bu, Jin Yang
cs.CVcond-mat.mtrl-sciarXiv:2609.01876v12026A Comparative Study of Modern Inference Techniques for Structured Discrete Energy Minimization Problems
Jörg H. Kappes, Bjoern Andres, Fred A. Hamprecht +10
cs.CVarXiv:1404.0533v12014InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
Shuai Yang, Hao Li, Bin Wang +7
cs.ROcs.CVarXiv:2507.17520v22025DCEvo: Discriminative Cross-Dimensional Evolutionary Learning for Infrared and Visible Image Fusion
Jinyuan Liu, Bowei Zhang, Qingyun Mei +6
cs.CVarXiv:2503.17673v12025Any6D: Model-free 6D Pose Estimation of Novel Objects
Taeyeop Lee, Bowen Wen, Minjun Kang +3
cs.CVcs.AIcs.ROarXiv:2503.18673v22025Deep learning for predicting refractive error from retinal fundus images
Avinash V. Varadarajan, Ryan Poplin, Katy Blumer +7
cs.CVarXiv:1712.07798v12017A Cone-Constrained Bilinear Decomposition for Total Scaled-Gradient Variation Models
Haibin Su, Chunlin Wu, Huibin Chang +1
cs.CVarXiv:2609.00036v12026ViewAL: Active Learning with Viewpoint Entropy for Semantic Segmentation
Yawar Siddiqui, Julien Valentin, Matthias Nießner
cs.CVcs.LGarXiv:1911.11789v22019Long Short-Term Memory Kalman Filters:Recurrent Neural Estimators for Pose Regularization
Huseyin Coskun, Felix Achilles, Robert DiPietro +2
cs.CVarXiv:1708.01885v12017Recurrent Neural Network for (Un-)supervised Learning of Monocular VideoVisual Odometry and Depth
Rui Wang, Stephen M. Pizer, Jan-Michael Frahm
cs.CVarXiv:1904.07087v12019NTIRE 2020 Challenge on Real-World Image Super-Resolution: Methods and Results
Andreas Lugmayr, Martin Danelljan, Radu Timofte +43
eess.IVcs.CVarXiv:2005.01996v12020A Comprehensive Survey on Knowledge Distillation
Amir M. Mansourian, Rozhan Ahmadi, Masoud Ghafouri +8
cs.CVarXiv:2503.12067v22025Evidential Deep Learning for Multi-Modal Anti-UAV Detection
Dmitry Golovchits, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag
cs.CVarXiv:2609.01742v12026CMUNeXt: An Efficient Medical Image Segmentation Network based on Large Kernel and Skip Fusion
Fenghe Tang, Jianrui Ding, Lingtao Wang +2
eess.IVcs.CVarXiv:2308.01239v220234D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration
Jiahui Zhang, Yurui Chen, Yueming Xu +8
cs.CVarXiv:2506.22242v22025SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
Wufei Ma, Yu-Cheng Chou, Qihao Liu +4
cs.CVarXiv:2504.20024v22025Fine-Grained Action Retrieval Through Multiple Parts-of-Speech Embeddings
Michael Wray, Diane Larlus, Gabriela Csurka +1
cs.CVarXiv:1908.03477v12019Decoding Visual Neural Representations by Multimodal Learning of Brain-Visual-Linguistic Features
Changde Du, Kaicheng Fu, Jinpeng Li +1
cs.CVcs.AIcs.MMarXiv:2210.06756v22022A Recipe for Watermarking Diffusion Models
Yunqing Zhao, Tianyu Pang, Chao Du +3
cs.CVcs.CRcs.LGarXiv:2303.10137v22023