Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

4,741 to 4,800 of 18,815

  1. MILD-Net: Minimal Information Loss Dilated Network for Gland Instance Segmentation in Colon Histology Images

    Simon Graham, Hao Chen, Jevgenij Gamper +5

    cs.CVarXiv:1806.01963v42018
  2. Unified Panoramic Geometry Estimation via Multi-View Foundation Models

    Vukasin Bozic, Isidora Slavkovic, Dominik Narnhofer +4

    cs.CVcs.AIarXiv:2605.26368v22026
  3. The GAN is dead; long live the GAN! A Modern GAN Baseline

    Yiwen Huang, Aaron Gokaslan, Volodymyr Kuleshov +1

    cs.LGcs.CVarXiv:2501.05441v12025
  4. Unsupervised Learning of a Hierarchical Spiking Neural Network for Optical Flow Estimation: From Events to Global Motion Perception

    Federico Paredes-Vallés, Kirk Y. W. Scheper, Guido C. H. E. de Croon

    cs.CVarXiv:1807.10936v22018
  5. A2DINOv3: Rethinking Multi-Modal Object Detection via Socialized Collaboration

    Jiekang Feng, Zhihe Fan, Yunqi Zhu +5

    cs.CVcs.AIarXiv:2608.21099v12026
  6. Panoptic Pairwise Distortion Graph

    Muhammad Kamran Janjua, Abdul Wahab, Bahador Rashidi

    cs.CVcs.AIcs.LGarXiv:2604.11004v12026
  7. Solar Cell Surface Defect Inspection Based on Multispectral Convolutional Neural Network

    Haiyong Chen, Yue Pang, Qidi Hu +1

    cs.CVeess.IVarXiv:1812.06220v12018
  8. Revisiting Point Cloud Shape Classification with a Simple and Effective Baseline

    Ankit Goyal, Hei Law, Bowei Liu +2

    cs.CVcs.LGarXiv:2106.05304v12021
  9. Attention-Based Deep Neural Networks for Detection of Cancerous and Precancerous Esophagus Tissue on Histopathological Slides

    Naofumi Tomita, Behnaz Abdollahi, Jason Wei +3

    eess.IVcs.CVarXiv:1811.08513v22018
  10. RASID: A Robust WLAN Device-free Passive Motion Detection System

    Ahmed E. Kosba, Ahmed Saeed, Moustafa Youssef

    cs.NIcs.CVarXiv:1105.6084v22011
  11. End-to-End Multimodal Emotion Recognition using Deep Neural Networks

    Panagiotis Tzirakis, George Trigeorgis, Mihalis A. Nicolaou +2

    cs.CVcs.CLarXiv:1704.08619v12017
  12. Lossy Image Compression with Compressive Autoencoders

    Lucas Theis, Wenzhe Shi, Andrew Cunningham +1

    stat.MLcs.CVarXiv:1703.00395v12017
  13. BiosecurID: a multimodal biometric database

    Julian Fierrez, Javier Galbally, Javier Ortega-Garcia +22

    cs.CRcs.CVeess.IVarXiv:2111.03472v12021
  14. Anatomy-specific classification of medical images using deep convolutional nets

    Holger R. Roth, Christopher T. Lee, Hoo-Chang Shin +5

    cs.CVarXiv:1504.04003v12015
  15. YOLOE: Real-Time Seeing Anything

    Ao Wang, Lihao Liu, Hui Chen +3

    cs.CVarXiv:2503.07465v22025
  16. Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning

    Yue Ma, Yulong Liu, Qiyuan Zhu +8

    cs.CVarXiv:2506.05207v42025
  17. Pedestrian Archetypes Extension -- More Pedestrian Models for Autonomous Vehicle Safety Testing

    Taorui Huang, Namita Gaidhani, Ritvik Bansal +6

    cs.CVarXiv:2607.16922v12026
  18. Towards Automatic Threat Detection: A Survey of Advances of Deep Learning within X-ray Security Imaging

    Samet Akcay, Toby Breckon

    cs.CVarXiv:2001.01293v22020
  19. Lossy Event Compression: From Event Stream Distortion to Task Performance

    Zahra Rezaee, Catarina Brites, João Ascenso

    cs.CVeess.IVarXiv:2608.28429v12026
  20. Multi-Scale Temporal Domain Alignment for Federated Video Domain Adaptation

    Lee En-Yi Hannah, Haozhi Cao, Yuecong Xu

    cs.CVarXiv:2608.29186v12026
  21. Object Detection Under Rainy Conditions for Autonomous Vehicles: A Review of State-of-the-Art and Emerging Techniques

    Mazin Hnewa, Hayder Radha

    cs.CVarXiv:2006.16471v42020
  22. CNN-based Density Estimation and Crowd Counting: A Survey

    Guangshuai Gao, Junyu Gao, Qingjie Liu +2

    cs.CVarXiv:2003.12783v12020
  23. Sim-to-Real Reinforcement Learning for Vision-Based Dexterous Manipulation on Humanoids

    Toru Lin, Kartik Sachdev, Linxi Fan +2

    cs.ROcs.AIcs.CVarXiv:2502.20396v22025
  24. SV4D 2.0: Enhancing Spatio-Temporal Consistency in Multi-View Video Diffusion for High-Quality 4D Generation

    Chun-Han Yao, Yiming Xie, Vikram Voleti +2

    cs.CVarXiv:2503.16396v32025
  25. LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling

    Zuhao Yang, Sudong Wang, Kaichen Zhang +8

    cs.CVarXiv:2511.20785v32025
  26. Multiple Object Tracking with Correlation Learning

    Qiang Wang, Yun Zheng, Pan Pan +1

    cs.CVarXiv:2104.03541v12021
  27. CAD-Llama: Leveraging Large Language Models for Computer-Aided Design Parametric 3D Model Generation

    Jiahao Li, Weijian Ma, Xueyang Li +3

    cs.CVarXiv:2505.04481v22025
  28. When Does Self-supervision Improve Few-shot Learning?

    Jong-Chyi Su, Subhransu Maji, Bharath Hariharan

    cs.CVcs.LGarXiv:1910.03560v22019
  29. Guardrail-Agnostic Societal Bias Evaluation in Large Vision-Language Models

    Yusuke Hirota, Michael Ross Boone, Arun George Zachariah +4

    cs.CVarXiv:2608.29590v12026
  30. HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation

    Zunnan Xu, Zhentao Yu, Zixiang Zhou +10

    cs.CVarXiv:2503.18860v22025
  31. Meta CLIP 2: A Worldwide Scaling Recipe

    Yung-Sung Chuang, Yang Li, Dong Wang +13

    cs.CVcs.CLarXiv:2507.22062v32025
  32. AssemblyNet: A large ensemble of CNNs for 3D Whole Brain MRI Segmentation

    Pierrick Coupé, Boris Mansencal, Michaël Clément +5

    eess.IVcs.CVcs.LGarXiv:1911.09098v12019
  33. Generative Feature Replay For Class-Incremental Learning

    Xialei Liu, Chenshen Wu, Mikel Menta +5

    cs.CVcs.LGarXiv:2004.09199v12020
  34. Med3DVLM: An Efficient Vision-Language Model for 3D Medical Image Analysis

    Yu Xin, Gorkem Can Ates, Kuang Gong +1

    cs.CVeess.IVarXiv:2503.20047v32025
  35. Input-Adaptive Gating of a Dehazing Front-End for On-Device Perception in Smoke-Obscured Environments

    Seongjun Kang, Ishaan Garg, Vishnu Bharadwaj

    cs.CVarXiv:2608.30034v12026
  36. LSNet: See Large, Focus Small

    Ao Wang, Hui Chen, Zijia Lin +2

    cs.CVarXiv:2503.23135v12025
  37. BSNet: Bi-Similarity Network for Few-shot Fine-grained Image Classification

    Xiaoxu Li, Jijie Wu, Zhuo Sun +3

    cs.CVarXiv:2011.14311v12020
  38. VLT: Vision-Language Transformer and Query Generation for Referring Segmentation

    Henghui Ding, Chang Liu, Suchen Wang +1

    cs.CVarXiv:2210.15871v12022
  39. DiffusionRenderer: Neural Inverse and Forward Rendering with Video Diffusion Models

    Ruofan Liang, Zan Gojcic, Huan Ling +8

    cs.CVcs.GRarXiv:2501.18590v22025
  40. RemoteSAM: Towards Segment Anything for Earth Observation

    Liang Yao, Fan Liu, Delong Chen +6

    cs.CVarXiv:2505.18022v32025
  41. Local Implicit Grid Representations for 3D Scenes

    Chiyu Max Jiang, Avneesh Sud, Ameesh Makadia +3

    cs.CVcs.CGcs.LGarXiv:2003.08981v12020
  42. RAFT-DVC: Resolution-Aware Machine Learning-Based Digital Volume Correlation

    Zixiang Tong, Lehu Bu, Jin Yang

    cs.CVcond-mat.mtrl-sciarXiv:2609.01876v12026
  43. A Comparative Study of Modern Inference Techniques for Structured Discrete Energy Minimization Problems

    Jörg H. Kappes, Bjoern Andres, Fred A. Hamprecht +10

    cs.CVarXiv:1404.0533v12014
  44. InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation

    Shuai Yang, Hao Li, Bin Wang +7

    cs.ROcs.CVarXiv:2507.17520v22025
  45. DCEvo: Discriminative Cross-Dimensional Evolutionary Learning for Infrared and Visible Image Fusion

    Jinyuan Liu, Bowei Zhang, Qingyun Mei +6

    cs.CVarXiv:2503.17673v12025
  46. Any6D: Model-free 6D Pose Estimation of Novel Objects

    Taeyeop Lee, Bowen Wen, Minjun Kang +3

    cs.CVcs.AIcs.ROarXiv:2503.18673v22025
  47. Deep learning for predicting refractive error from retinal fundus images

    Avinash V. Varadarajan, Ryan Poplin, Katy Blumer +7

    cs.CVarXiv:1712.07798v12017
  48. A Cone-Constrained Bilinear Decomposition for Total Scaled-Gradient Variation Models

    Haibin Su, Chunlin Wu, Huibin Chang +1

    cs.CVarXiv:2609.00036v12026
  49. ViewAL: Active Learning with Viewpoint Entropy for Semantic Segmentation

    Yawar Siddiqui, Julien Valentin, Matthias Nießner

    cs.CVcs.LGarXiv:1911.11789v22019
  50. Long Short-Term Memory Kalman Filters:Recurrent Neural Estimators for Pose Regularization

    Huseyin Coskun, Felix Achilles, Robert DiPietro +2

    cs.CVarXiv:1708.01885v12017
  51. Recurrent Neural Network for (Un-)supervised Learning of Monocular VideoVisual Odometry and Depth

    Rui Wang, Stephen M. Pizer, Jan-Michael Frahm

    cs.CVarXiv:1904.07087v12019
  52. NTIRE 2020 Challenge on Real-World Image Super-Resolution: Methods and Results

    Andreas Lugmayr, Martin Danelljan, Radu Timofte +43

    eess.IVcs.CVarXiv:2005.01996v12020
  53. A Comprehensive Survey on Knowledge Distillation

    Amir M. Mansourian, Rozhan Ahmadi, Masoud Ghafouri +8

    cs.CVarXiv:2503.12067v22025
  54. Evidential Deep Learning for Multi-Modal Anti-UAV Detection

    Dmitry Golovchits, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag

    cs.CVarXiv:2609.01742v12026
  55. CMUNeXt: An Efficient Medical Image Segmentation Network based on Large Kernel and Skip Fusion

    Fenghe Tang, Jianrui Ding, Lingtao Wang +2

    eess.IVcs.CVarXiv:2308.01239v22023
  56. 4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration

    Jiahui Zhang, Yurui Chen, Yueming Xu +8

    cs.CVarXiv:2506.22242v22025
  57. SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning

    Wufei Ma, Yu-Cheng Chou, Qihao Liu +4

    cs.CVarXiv:2504.20024v22025
  58. Fine-Grained Action Retrieval Through Multiple Parts-of-Speech Embeddings

    Michael Wray, Diane Larlus, Gabriela Csurka +1

    cs.CVarXiv:1908.03477v12019
  59. Decoding Visual Neural Representations by Multimodal Learning of Brain-Visual-Linguistic Features

    Changde Du, Kaicheng Fu, Jinpeng Li +1

    cs.CVcs.AIcs.MMarXiv:2210.06756v22022
  60. A Recipe for Watermarking Diffusion Models

    Yunqing Zhao, Tianyu Pang, Chao Du +3

    cs.CVcs.CRcs.LGarXiv:2303.10137v22023