Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

7,081 to 7,140 of 18,830

  1. Dream3D: Zero-Shot Text-to-3D Synthesis Using 3D Shape Prior and Text-to-Image Diffusion Models

    Jiale Xu, Xintao Wang, Weihao Cheng +4

    cs.CVarXiv:2212.14704v22022
  2. DINOv3

    Oriane Siméoni, Huy V. Vo, Maximilian Seitzer +23

    cs.CVcs.LGarXiv:2508.10104v12025
  3. YOLOv12: Attention-Centric Real-Time Object Detectors

    Yunjie Tian, Qixiang Ye, David Doermann

    cs.CVcs.AIarXiv:2502.12524v12025
  4. SphereReID: Deep Hypersphere Manifold Embedding for Person Re-Identification

    Xing Fan, Wei Jiang, Hao Luo +1

    cs.CVarXiv:1807.00537v12018
  5. MIDR: Enrichment-Augmented Indexing for Multimodal Document Retrieval

    Debanjan Mahata, Atharva Tendle, Daniel Preotiuc-Pietro +2

    cs.IRcs.AIcs.CLarXiv:2609.01316v12026
  6. Wan: Open and Advanced Large-Scale Video Generative Models

    Team Wan, Ang Wang, Baole Ai +59

    cs.CVarXiv:2503.20314v22025
  7. Qwen2.5-VL Technical Report

    Shuai Bai, Keqin Chen, Xuejing Liu +24

    cs.CVcs.CLarXiv:2502.13923v12025
  8. Predicting Risk of Developing Diabetic Retinopathy using Deep Learning

    Ashish Bora, Siva Balasubramanian, Boris Babenko +13

    eess.IVcs.CVarXiv:2008.04370v12020
  9. NUWA-XL: Diffusion over Diffusion for eXtremely Long Video Generation

    Shengming Yin, Chenfei Wu, Huan Yang +13

    cs.CVcs.AIarXiv:2303.12346v12023
  10. UFOGen: You Forward Once Large Scale Text-to-Image Generation via Diffusion GANs

    Yanwu Xu, Yang Zhao, Zhisheng Xiao +1

    cs.CVarXiv:2311.09257v52023
  11. PolyTransform: Deep Polygon Transformer for Instance Segmentation

    Justin Liang, Namdar Homayounfar, Wei-Chiu Ma +3

    cs.CVarXiv:1912.02801v42019
  12. Joint Line Segmentation and Transcription for End-to-End Handwritten Paragraph Recognition

    Théodore Bluche

    cs.CVcs.LGcs.NEarXiv:1604.08352v12016
  13. HarmoFL: Harmonizing Local and Global Drifts in Federated Learning on Heterogeneous Medical Images

    Meirui Jiang, Zirui Wang, Qi Dou

    eess.IVcs.AIcs.CVarXiv:2112.10775v32021
  14. InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual Referring

    Zhihao Yuan, Xu Yan, Yinghong Liao +4

    cs.CVarXiv:2103.01128v22021
  15. Naturalistic Driver Intention and Path Prediction using Recurrent Neural Networks

    Alex Zyner, Stewart Worrall, Eduardo Nebot

    cs.CVarXiv:1807.09995v12018
  16. A Language Agent for Autonomous Driving

    Jiageng Mao, Junjie Ye, Yuxi Qian +2

    cs.CVcs.AIcs.CLarXiv:2311.10813v42023
  17. Pix2Rep-v2: Data-Efficient Representation Learning for Dense Medical Imaging Applications

    S. Sifaoui, E. Angelini, S. Toupin +2

    cs.CVarXiv:2609.01427v12026
  18. Object as Hotspots: An Anchor-Free 3D Object Detection Approach via Firing of Hotspots

    Qi Chen, Lin Sun, Zhixin Wang +2

    cs.CVarXiv:1912.12791v32019
  19. DenseReg: Fully Convolutional Dense Shape Regression In-the-Wild

    Riza Alp Guler, Yuxiang Zhou, George Trigeorgis +4

    cs.CVarXiv:1803.02188v22018
  20. Linguistic Structure Guided Context Modeling for Referring Image Segmentation

    Tianrui Hui, Si Liu, Shaofei Huang +4

    cs.CVcs.CLarXiv:2010.00515v32020
  21. Improving Chest X-Ray Report Generation by Leveraging Warm Starting

    Aaron Nicolson, Jason Dowling, Bevan Koopman

    cs.CVarXiv:2201.09405v22022
  22. Embedding Label Structures for Fine-Grained Feature Representation

    Xiaofan Zhang, Feng Zhou, Yuanqing Lin +1

    cs.CVarXiv:1512.02895v22015
  23. Deep Pictorial Gaze Estimation

    Seonwook Park, Adrian Spurr, Otmar Hilliges

    cs.CVarXiv:1807.10002v12018
  24. FASTER: Fast and Safe Trajectory Planner for Navigation in Unknown Environments

    Jesus Tordesillas, Brett T. Lopez, Michael Everett +1

    cs.ROcs.CVarXiv:2001.04420v22020
  25. Rethinking Depthwise Separable Convolutions: How Intra-Kernel Correlations Lead to Improved MobileNets

    Daniel Haase, Manuel Amthor

    cs.CVarXiv:2003.13549v32020
  26. Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration

    Daehwan Kim, Haejun Chung, Ikbeom Jang

    cs.LGcs.CVarXiv:2609.01072v22026
  27. BCNet: Learning Body and Cloth Shape from A Single Image

    Boyi Jiang, Juyong Zhang, Yang Hong +3

    cs.CVcs.GRarXiv:2004.00214v22020
  28. Region Normalization for Image Inpainting

    Tao Yu, Zongyu Guo, Xin Jin +5

    cs.CVarXiv:1911.10375v22019
  29. Highly accurate model for prediction of lung nodule malignancy with CT scans

    Jason Causey, Junyu Zhang, Shiqian Ma +6

    cs.CVq-bio.QMstat.MLarXiv:1802.01756v12018
    Summaries:한국어
  30. Adversarial Reprogramming of Neural Networks

    Gamaleldin F. Elsayed, Ian Goodfellow, Jascha Sohl-Dickstein

    cs.LGcs.CRcs.CVarXiv:1806.11146v22018
  31. Multi-granularity Generator for Temporal Action Proposal

    Yuan Liu, Lin Ma, Yifeng Zhang +2

    cs.CVarXiv:1811.11524v22018
  32. Video Panoptic Segmentation

    Dahun Kim, Sanghyun Woo, Joon-Young Lee +1

    cs.CVarXiv:2006.11339v12020
  33. Dense Classification and Implanting for Few-Shot Learning

    Yann Lifchitz, Yannis Avrithis, Sylvaine Picard +1

    cs.CVarXiv:1903.05050v12019
  34. Scale-based Approach for Active Wildfire Segmentation on Satellite Imagery

    Matheus F. Kovaleski, Cristiano Premebida, João Ruivo Paulo

    cs.CVarXiv:2609.01392v12026
  35. Beyond Local Search: Tracking Objects Everywhere with Instance-Specific Proposals

    Gao Zhu, Fatih Porikli, Hongdong Li

    cs.CVarXiv:1605.01839v12016
  36. Multimodal RGB-Infrared Combination for UAV-Based Wildfire Segmentation: A Comparative Study on FLAME3

    Matheus F. Kovaleski, Luís Garrote, Cristiano Premebida +2

    cs.CVarXiv:2609.01390v12026
  37. Pixel-wise Anomaly Detection in Complex Driving Scenes

    Giancarlo Di Biase, Hermann Blum, Roland Siegwart +1

    cs.CVarXiv:2103.05445v12021
  38. Learning Context Graph for Person Search

    Yichao Yan, Qiang Zhang, Bingbing Ni +3

    cs.CVarXiv:1904.01830v12019
  39. HomebrewedDB: RGB-D Dataset for 6D Pose Estimation of 3D Objects

    Roman Kaskman, Sergey Zakharov, Ivan Shugurov +1

    cs.CVcs.ROarXiv:1904.03167v22019
  40. Deep Gradient Projection Networks for Pan-sharpening

    Shuang Xu, Jiangshe Zhang, Zixiang Zhao +3

    cs.CVeess.IVarXiv:2103.04584v12021
  41. Deep Unsupervised Saliency Detection: A Multiple Noisy Labeling Perspective

    Jing Zhang, Tong Zhang, Yuchao Dai +2

    cs.CVarXiv:1803.10910v12018
  42. Maximum-Entropy Adversarial Data Augmentation for Improved Generalization and Robustness

    Long Zhao, Ting Liu, Xi Peng +1

    cs.LGcs.CVarXiv:2010.08001v22020
  43. Agentic Multimodal Models for Environmental Hyperspectral Unmixing

    Michał Cholewa, Luca Ciampi, Nicola Messina +2

    cs.CVarXiv:2609.01289v12026
  44. Level Playing Field for Million Scale Face Recognition

    Aaron Nech, Ira Kemelmacher-Shlizerman

    cs.CVarXiv:1705.00393v12017
  45. Fingerprint Spoof Buster

    Tarang Chugh, Kai Cao, Anil K. Jain

    cs.CVarXiv:1712.04489v12017
  46. Instant Volumetric Head Avatars

    Wojciech Zielonka, Timo Bolkart, Justus Thies

    cs.CVarXiv:2211.12499v22022
  47. Something-Else: Compositional Action Recognition with Spatial-Temporal Interaction Networks

    Joanna Materzynska, Tete Xiao, Roei Herzig +3

    cs.CVarXiv:1912.09930v32019
  48. Improved Automatic Target Recognition in Synthetic Aperture Sonar Imagery Using Large Deep Neural Networks

    C. J. Moore, Alex Hurt, Jordan Malof

    cs.CVarXiv:2609.01800v12026
  49. BEVSegFormer: Bird's Eye View Semantic Segmentation From Arbitrary Camera Rigs

    Lang Peng, Zhirong Chen, Zhangjie Fu +2

    cs.CVarXiv:2203.04050v32022
  50. Long-Tailed Recognition via Weight Balancing

    Shaden Alshammari, Yu-Xiong Wang, Deva Ramanan +1

    cs.CVarXiv:2203.14197v12022
  51. A Survey on Long-Tailed Visual Recognition

    Lu Yang, He Jiang, Qing Song +1

    cs.CVarXiv:2205.13775v12022
  52. An End-to-End Transformer Model for Crowd Localization

    Dingkang Liang, Wei Xu, Xiang Bai

    cs.CVarXiv:2202.13065v22022
  53. FeTrIL: Feature Translation for Exemplar-Free Class-Incremental Learning

    Grégoire Petit, Adrian Popescu, Hugo Schindler +2

    cs.CVcs.AIcs.LGarXiv:2211.13131v22022
  54. Remote Sensing Cross-Modal Text-Image Retrieval Based on Global and Local Information

    Zhiqiang Yuan, Wenkai Zhang, Changyuan Tian +5

    cs.CVcs.IRcs.MMarXiv:2204.09860v12022
  55. CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical Flow

    Philippe Weinzaepfel, Thomas Lucas, Vincent Leroy +7

    cs.CVarXiv:2211.10408v32022
  56. Occupancy Anticipation for Efficient Exploration and Navigation

    Santhosh K. Ramakrishnan, Ziad Al-Halah, Kristen Grauman

    cs.CVarXiv:2008.09285v22020
  57. Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental Learning

    Kai Zhu, Wei Zhai, Yang Cao +2

    cs.CVarXiv:2203.06359v22022
  58. Catching Both Gray and Black Swans: Open-set Supervised Anomaly Detection

    Choubo Ding, Guansong Pang, Chunhua Shen

    cs.CVarXiv:2203.14506v12022
  59. Fully Convolutional Networks for Continuous Sign Language Recognition

    Ka Leong Cheng, Zhaoyang Yang, Qifeng Chen +1

    cs.CVarXiv:2007.12402v12020
  60. MonoDETR: Depth-guided Transformer for Monocular 3D Object Detection

    Renrui Zhang, Han Qiu, Tai Wang +7

    cs.CVcs.AIeess.IVarXiv:2203.13310v52022