Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
7,081 to 7,140 of 18,830
Dream3D: Zero-Shot Text-to-3D Synthesis Using 3D Shape Prior and Text-to-Image Diffusion Models
Jiale Xu, Xintao Wang, Weihao Cheng +4
cs.CVarXiv:2212.14704v22022DINOv3
Oriane Siméoni, Huy V. Vo, Maximilian Seitzer +23
cs.CVcs.LGarXiv:2508.10104v12025YOLOv12: Attention-Centric Real-Time Object Detectors
Yunjie Tian, Qixiang Ye, David Doermann
cs.CVcs.AIarXiv:2502.12524v12025SphereReID: Deep Hypersphere Manifold Embedding for Person Re-Identification
Xing Fan, Wei Jiang, Hao Luo +1
cs.CVarXiv:1807.00537v12018MIDR: Enrichment-Augmented Indexing for Multimodal Document Retrieval
Debanjan Mahata, Atharva Tendle, Daniel Preotiuc-Pietro +2
cs.IRcs.AIcs.CLarXiv:2609.01316v12026Wan: Open and Advanced Large-Scale Video Generative Models
Team Wan, Ang Wang, Baole Ai +59
cs.CVarXiv:2503.20314v22025Qwen2.5-VL Technical Report
Shuai Bai, Keqin Chen, Xuejing Liu +24
cs.CVcs.CLarXiv:2502.13923v12025Predicting Risk of Developing Diabetic Retinopathy using Deep Learning
Ashish Bora, Siva Balasubramanian, Boris Babenko +13
eess.IVcs.CVarXiv:2008.04370v12020NUWA-XL: Diffusion over Diffusion for eXtremely Long Video Generation
Shengming Yin, Chenfei Wu, Huan Yang +13
cs.CVcs.AIarXiv:2303.12346v12023UFOGen: You Forward Once Large Scale Text-to-Image Generation via Diffusion GANs
Yanwu Xu, Yang Zhao, Zhisheng Xiao +1
cs.CVarXiv:2311.09257v52023PolyTransform: Deep Polygon Transformer for Instance Segmentation
Justin Liang, Namdar Homayounfar, Wei-Chiu Ma +3
cs.CVarXiv:1912.02801v42019Joint Line Segmentation and Transcription for End-to-End Handwritten Paragraph Recognition
Théodore Bluche
cs.CVcs.LGcs.NEarXiv:1604.08352v12016HarmoFL: Harmonizing Local and Global Drifts in Federated Learning on Heterogeneous Medical Images
Meirui Jiang, Zirui Wang, Qi Dou
eess.IVcs.AIcs.CVarXiv:2112.10775v32021InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual Referring
Zhihao Yuan, Xu Yan, Yinghong Liao +4
cs.CVarXiv:2103.01128v22021Naturalistic Driver Intention and Path Prediction using Recurrent Neural Networks
Alex Zyner, Stewart Worrall, Eduardo Nebot
cs.CVarXiv:1807.09995v12018A Language Agent for Autonomous Driving
Jiageng Mao, Junjie Ye, Yuxi Qian +2
cs.CVcs.AIcs.CLarXiv:2311.10813v42023Pix2Rep-v2: Data-Efficient Representation Learning for Dense Medical Imaging Applications
S. Sifaoui, E. Angelini, S. Toupin +2
cs.CVarXiv:2609.01427v12026Object as Hotspots: An Anchor-Free 3D Object Detection Approach via Firing of Hotspots
Qi Chen, Lin Sun, Zhixin Wang +2
cs.CVarXiv:1912.12791v32019DenseReg: Fully Convolutional Dense Shape Regression In-the-Wild
Riza Alp Guler, Yuxiang Zhou, George Trigeorgis +4
cs.CVarXiv:1803.02188v22018Linguistic Structure Guided Context Modeling for Referring Image Segmentation
Tianrui Hui, Si Liu, Shaofei Huang +4
cs.CVcs.CLarXiv:2010.00515v32020Improving Chest X-Ray Report Generation by Leveraging Warm Starting
Aaron Nicolson, Jason Dowling, Bevan Koopman
cs.CVarXiv:2201.09405v22022Embedding Label Structures for Fine-Grained Feature Representation
Xiaofan Zhang, Feng Zhou, Yuanqing Lin +1
cs.CVarXiv:1512.02895v22015Deep Pictorial Gaze Estimation
Seonwook Park, Adrian Spurr, Otmar Hilliges
cs.CVarXiv:1807.10002v12018FASTER: Fast and Safe Trajectory Planner for Navigation in Unknown Environments
Jesus Tordesillas, Brett T. Lopez, Michael Everett +1
cs.ROcs.CVarXiv:2001.04420v22020Rethinking Depthwise Separable Convolutions: How Intra-Kernel Correlations Lead to Improved MobileNets
Daniel Haase, Manuel Amthor
cs.CVarXiv:2003.13549v32020Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration
Daehwan Kim, Haejun Chung, Ikbeom Jang
cs.LGcs.CVarXiv:2609.01072v22026BCNet: Learning Body and Cloth Shape from A Single Image
Boyi Jiang, Juyong Zhang, Yang Hong +3
cs.CVcs.GRarXiv:2004.00214v22020Region Normalization for Image Inpainting
Tao Yu, Zongyu Guo, Xin Jin +5
cs.CVarXiv:1911.10375v22019Highly accurate model for prediction of lung nodule malignancy with CT scans
Jason Causey, Junyu Zhang, Shiqian Ma +6
cs.CVq-bio.QMstat.MLarXiv:1802.01756v12018Summaries:한국어Adversarial Reprogramming of Neural Networks
Gamaleldin F. Elsayed, Ian Goodfellow, Jascha Sohl-Dickstein
cs.LGcs.CRcs.CVarXiv:1806.11146v22018Multi-granularity Generator for Temporal Action Proposal
Yuan Liu, Lin Ma, Yifeng Zhang +2
cs.CVarXiv:1811.11524v22018Video Panoptic Segmentation
Dahun Kim, Sanghyun Woo, Joon-Young Lee +1
cs.CVarXiv:2006.11339v12020Dense Classification and Implanting for Few-Shot Learning
Yann Lifchitz, Yannis Avrithis, Sylvaine Picard +1
cs.CVarXiv:1903.05050v12019Scale-based Approach for Active Wildfire Segmentation on Satellite Imagery
Matheus F. Kovaleski, Cristiano Premebida, João Ruivo Paulo
cs.CVarXiv:2609.01392v12026Beyond Local Search: Tracking Objects Everywhere with Instance-Specific Proposals
Gao Zhu, Fatih Porikli, Hongdong Li
cs.CVarXiv:1605.01839v12016Multimodal RGB-Infrared Combination for UAV-Based Wildfire Segmentation: A Comparative Study on FLAME3
Matheus F. Kovaleski, Luís Garrote, Cristiano Premebida +2
cs.CVarXiv:2609.01390v12026Pixel-wise Anomaly Detection in Complex Driving Scenes
Giancarlo Di Biase, Hermann Blum, Roland Siegwart +1
cs.CVarXiv:2103.05445v12021Learning Context Graph for Person Search
Yichao Yan, Qiang Zhang, Bingbing Ni +3
cs.CVarXiv:1904.01830v12019HomebrewedDB: RGB-D Dataset for 6D Pose Estimation of 3D Objects
Roman Kaskman, Sergey Zakharov, Ivan Shugurov +1
cs.CVcs.ROarXiv:1904.03167v22019Deep Gradient Projection Networks for Pan-sharpening
Shuang Xu, Jiangshe Zhang, Zixiang Zhao +3
cs.CVeess.IVarXiv:2103.04584v12021Deep Unsupervised Saliency Detection: A Multiple Noisy Labeling Perspective
Jing Zhang, Tong Zhang, Yuchao Dai +2
cs.CVarXiv:1803.10910v12018Maximum-Entropy Adversarial Data Augmentation for Improved Generalization and Robustness
Long Zhao, Ting Liu, Xi Peng +1
cs.LGcs.CVarXiv:2010.08001v22020Agentic Multimodal Models for Environmental Hyperspectral Unmixing
Michał Cholewa, Luca Ciampi, Nicola Messina +2
cs.CVarXiv:2609.01289v12026Level Playing Field for Million Scale Face Recognition
Aaron Nech, Ira Kemelmacher-Shlizerman
cs.CVarXiv:1705.00393v12017Fingerprint Spoof Buster
Tarang Chugh, Kai Cao, Anil K. Jain
cs.CVarXiv:1712.04489v12017Instant Volumetric Head Avatars
Wojciech Zielonka, Timo Bolkart, Justus Thies
cs.CVarXiv:2211.12499v22022Something-Else: Compositional Action Recognition with Spatial-Temporal Interaction Networks
Joanna Materzynska, Tete Xiao, Roei Herzig +3
cs.CVarXiv:1912.09930v32019Improved Automatic Target Recognition in Synthetic Aperture Sonar Imagery Using Large Deep Neural Networks
C. J. Moore, Alex Hurt, Jordan Malof
cs.CVarXiv:2609.01800v12026BEVSegFormer: Bird's Eye View Semantic Segmentation From Arbitrary Camera Rigs
Lang Peng, Zhirong Chen, Zhangjie Fu +2
cs.CVarXiv:2203.04050v32022Long-Tailed Recognition via Weight Balancing
Shaden Alshammari, Yu-Xiong Wang, Deva Ramanan +1
cs.CVarXiv:2203.14197v12022A Survey on Long-Tailed Visual Recognition
Lu Yang, He Jiang, Qing Song +1
cs.CVarXiv:2205.13775v12022An End-to-End Transformer Model for Crowd Localization
Dingkang Liang, Wei Xu, Xiang Bai
cs.CVarXiv:2202.13065v22022FeTrIL: Feature Translation for Exemplar-Free Class-Incremental Learning
Grégoire Petit, Adrian Popescu, Hugo Schindler +2
cs.CVcs.AIcs.LGarXiv:2211.13131v22022Remote Sensing Cross-Modal Text-Image Retrieval Based on Global and Local Information
Zhiqiang Yuan, Wenkai Zhang, Changyuan Tian +5
cs.CVcs.IRcs.MMarXiv:2204.09860v12022CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical Flow
Philippe Weinzaepfel, Thomas Lucas, Vincent Leroy +7
cs.CVarXiv:2211.10408v32022Occupancy Anticipation for Efficient Exploration and Navigation
Santhosh K. Ramakrishnan, Ziad Al-Halah, Kristen Grauman
cs.CVarXiv:2008.09285v22020Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental Learning
Kai Zhu, Wei Zhai, Yang Cao +2
cs.CVarXiv:2203.06359v22022Catching Both Gray and Black Swans: Open-set Supervised Anomaly Detection
Choubo Ding, Guansong Pang, Chunhua Shen
cs.CVarXiv:2203.14506v12022Fully Convolutional Networks for Continuous Sign Language Recognition
Ka Leong Cheng, Zhaoyang Yang, Qifeng Chen +1
cs.CVarXiv:2007.12402v12020MonoDETR: Depth-guided Transformer for Monocular 3D Object Detection
Renrui Zhang, Han Qiu, Tai Wang +7
cs.CVcs.AIeess.IVarXiv:2203.13310v52022