Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
14,101 to 14,160 of 18,780
Pixel-Space Diffusion via Observation Operators
Shaojie Guo, Lichen Ma, Haoyang Tong +8
cs.CVarXiv:2608.21885v22026CORe50: a New Dataset and Benchmark for Continuous Object Recognition
Vincenzo Lomonaco, Davide Maltoni
cs.CVcs.AIcs.LGarXiv:1705.03550v12017Pose Invariant Embedding for Deep Person Re-identification
Liang Zheng, Yujia Huang, Huchuan Lu +1
cs.CVarXiv:1701.07732v12017Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks
Bin Xiao, Haiping Wu, Weijian Xu +6
cs.CVarXiv:2311.06242v12023RPM-Net: Robust Point Matching using Learned Features
Zi Jian Yew, Gim Hee Lee
cs.CVarXiv:2003.13479v12020Escaping the Big Data Paradigm with Compact Transformers
Ali Hassani, Steven Walton, Nikhil Shah +3
cs.CVcs.LGarXiv:2104.05704v42021Relation-Aware Global Attention for Person Re-identification
Zhizheng Zhang, Cuiling Lan, Wenjun Zeng +2
cs.CVarXiv:1904.02998v22019SAFE-G: Structure-aware Faithful Evidence-guided Generation for Knowledge-based Visual Question Answering
Long Shu, Shuochen Liu, Wei Chen +4
cs.CVcs.AIarXiv:2608.21796v12026Deep Subspace Clustering Networks
Pan Ji, Tong Zhang, Hongdong Li +2
cs.CVarXiv:1709.02508v12017GhostNetV2: Enhance Cheap Operation with Long-Range Attention
Yehui Tang, Kai Han, Jianyuan Guo +3
cs.CVarXiv:2211.12905v12022Latte: Latent Diffusion Transformer for Video Generation
Xin Ma, Yaohui Wang, Xinyuan Chen +5
cs.CVarXiv:2401.03048v32024Mode Regularized Generative Adversarial Networks
Tong Che, Yanran Li, Athul Paul Jacob +2
cs.LGcs.AIcs.CVarXiv:1612.02136v52016ResShift: Efficient Diffusion Model for Image Super-resolution by Residual Shifting
Zongsheng Yue, Jianyi Wang, Chen Change Loy
cs.CVarXiv:2307.12348v32023The Blending Ratio Is Not Where the Performance Is: Diagnosing Prototype Blending for Few-Shot Adaptation of Vision-Language Models
Liangzhi Li, Bowen Wang, Yiming Qian +3
cs.CVcs.LGarXiv:2608.23634v12026LF-Net: Learning Local Features from Images
Yuki Ono, Eduard Trulls, Pascal Fua +1
cs.CVarXiv:1805.09662v22018Unsupervised Discovery of Mid-Level Discriminative Patches
Saurabh Singh, Abhinav Gupta, Alexei A. Efros
cs.CVcs.AIcs.LGarXiv:1205.3137v22012clDice -- A Novel Topology-Preserving Loss Function for Tubular Structure Segmentation
Suprosanna Shit, Johannes C. Paetzold, Anjany Sekuboyina +6
cs.CVcs.LGeess.IVarXiv:2003.07311v72020OCGAN: One-class Novelty Detection Using GANs with Constrained Latent Representations
Pramuditha Perera, Ramesh Nallapati, Bing Xiang
cs.CVcs.LGarXiv:1903.08550v12019SAM3Dual: A 3rd Place Solution to the MOSEv2 Track, 8th LSVOS Challenge
JeongRae Kim, Chaehyun Kim, Changwon Lim
cs.CVarXiv:2608.22193v12026Image2StyleGAN++: How to Edit the Embedded Images?
Rameen Abdal, Yipeng Qin, Peter Wonka
cs.CVcs.GRarXiv:1911.11544v22019MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images
Yuedong Chen, Haofei Xu, Chuanxia Zheng +5
cs.CVarXiv:2403.14627v22024You Only Learn One Representation: Unified Network for Multiple Tasks
Chien-Yao Wang, I-Hau Yeh, Hong-Yuan Mark Liao
cs.CVarXiv:2105.04206v12021Ternary Weight Networks
Fengfu Li, Bin Liu, Xiaoxing Wang +2
cs.CVarXiv:1605.04711v32016MPIIGaze: Real-World Dataset and Deep Appearance-Based Gaze Estimation
Xucong Zhang, Yusuke Sugano, Mario Fritz +1
cs.CVarXiv:1711.09017v12017BC-IHV: Conditioning the Color Space for Stable Rectified-Flow Low-Light Enhancement
Yi Ai, Zheng Chen, Yuanhao Cai +2
cs.CVarXiv:2608.21847v12026Robust fine-tuning of zero-shot models
Mitchell Wortsman, Gabriel Ilharco, Jong Wook Kim +8
cs.CVcs.LGarXiv:2109.01903v32021Summaries:한국어Real-Time Adaptive Image Compression
Oren Rippel, Lubomir Bourdev
stat.MLcs.CVcs.LGarXiv:1705.05823v12017Progressively Learning Heterogeneous Skills in a Unified Latent Space
Yue-Yi Zhang, Ming Gong, Linpu He +2
cs.CVcs.AIcs.ROarXiv:2608.23258v12026HeatTok: Enhancing Remote Sensing Image Understanding via Thermodiffusion-based Tokenization
Yingying Yan, Jiaqi Tang, Wei Wei +6
cs.CVarXiv:2608.22485v12026Abnormal Event Detection in Videos using Spatiotemporal Autoencoder
Yong Shean Chong, Yong Haur Tay
cs.CVarXiv:1701.01546v12017SAL: Sign Agnostic Learning of Shapes from Raw Data
Matan Atzmon, Yaron Lipman
cs.CVcs.GRcs.LGarXiv:1911.10414v22019Cycle-Dehaze: Enhanced CycleGAN for Single Image Dehazing
Deniz Engin, Anıl Genç, Hazım Kemal Ekenel
cs.CVarXiv:1805.05308v12018Revisiting Local Descriptor based Image-to-Class Measure for Few-shot Learning
Wenbin Li, Lei Wang, Jinglin Xu +3
cs.CVarXiv:1903.12290v22019Curriculum Learning: A Survey
Petru Soviany, Radu Tudor Ionescu, Paolo Rota +1
cs.LGcs.CLcs.CVarXiv:2101.10382v32021Gradient Harmonized Single-stage Detector
Buyu Li, Yu Liu, Xiaogang Wang
cs.CVarXiv:1811.05181v12018Prototypical Pseudo Label Denoising and Target Structure Learning for Domain Adaptive Semantic Segmentation
Pan Zhang, Bo Zhang, Ting Zhang +3
cs.CVarXiv:2101.10979v22021ByteAction: Byte-space Action Recognition Foundation Model
Fangcheng Li, Zhen Yu, Kejun Wu +2
cs.CVarXiv:2608.22760v12026DRAgent: Discriminative Reasoning Agent for Referring Expression Segmentation
Yujie Qi, Luyan Zhang
cs.CVarXiv:2608.22885v12026T-LESS: An RGB-D Dataset for 6D Pose Estimation of Texture-less Objects
Tomas Hodan, Pavel Haluza, Stepan Obdrzalek +3
cs.CVcs.AIcs.ROarXiv:1701.05498v12017RGB-T Object Tracking:Benchmark and Baseline
Chenglong Li, Xinyan Liang, Yijuan Lu +2
cs.CVarXiv:1805.08982v12018Deep Learning for Classification of Hyperspectral Data: A Comparative Review
Nicolas Audebert, Bertrand Saux, Sébastien Lefèvre
cs.LGcs.CVcs.NEarXiv:1904.10674v12019Deep Spatial Autoencoders for Visuomotor Learning
Chelsea Finn, Xin Yu Tan, Yan Duan +3
cs.LGcs.CVcs.ROarXiv:1509.06113v32015Deep Neural Networks Improve Radiologists' Performance in Breast Cancer Screening
Nan Wu, Jason Phang, Jungkyu Park +29
cs.LGcs.CVstat.MLarXiv:1903.08297v12019Deep Learning Ensembles for Melanoma Recognition in Dermoscopy Images
Noel Codella, Quoc-Bao Nguyen, Sharath Pankanti +4
cs.CVarXiv:1610.04662v22016Bayesian CP Factorization of Incomplete Tensors with Automatic Rank Determination
Qibin Zhao, Liqing Zhang, Andrzej Cichocki
cs.LGcs.CVstat.MLarXiv:1401.6497v22014Approximating CNNs with Bag-of-local-Features models works surprisingly well on ImageNet
Wieland Brendel, Matthias Bethge
cs.CVcs.LGstat.MLarXiv:1904.00760v12019The Sound of Pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko +3
cs.CVcs.SDeess.ASarXiv:1804.03160v42018PixelLink: Detecting Scene Text via Instance Segmentation
Dan Deng, Haifeng Liu, Xuelong Li +1
cs.CVarXiv:1801.01315v12018Score-based diffusion models for accelerated MRI
Hyungjin Chung, Jong Chul Ye
eess.IVcs.AIcs.CVarXiv:2110.05243v32021SegAN: Adversarial Network with Multi-scale $L_1$ Loss for Medical Image Segmentation
Yuan Xue, Tao Xu, Han Zhang +2
cs.CVarXiv:1706.01805v22017Dual-Path Convolutional Image-Text Embeddings with Instance Loss
Zhedong Zheng, Liang Zheng, Michael Garrett +3
cs.CVcs.MMarXiv:1711.05535v42017Modelling Uncertainty in Deep Learning for Camera Relocalization
Alex Kendall, Roberto Cipolla
cs.CVcs.ROarXiv:1509.05909v22015PersonLab: Person Pose Estimation and Instance Segmentation with a Bottom-Up, Part-Based, Geometric Embedding Model
George Papandreou, Tyler Zhu, Liang-Chieh Chen +3
cs.CVarXiv:1803.08225v12018More Motion Is Not Always Better Motion: Corpus Composition Governs Whether Augmentation Helps SMPL-Based Parkinsonian Gait Severity Estimation
Michael Caiola, Andrew C. Weitz
cs.CVarXiv:2608.23730v12026M$^3$ISR: A Multi-Modal Multi-View Benchmark for 3D/4D Gaussian Splatting and Feedforward Compression
Xinhui Liu, Lei Liu, Zhenghao Chen +3
cs.CVarXiv:2608.22465v12026Unsupervised Scale-consistent Depth and Ego-motion Learning from Monocular Video
Jia-Wang Bian, Zhichao Li, Naiyan Wang +4
cs.CVarXiv:1908.10553v22019Hyper^2: Unleashing Hyperbolic Geometry's Full Potential via Dual-Space Consistency
Guantian Zheng, Haiyang Xu, Tianyu Gao
cs.CVarXiv:2608.22238v12026DeepMVS: Learning Multi-view Stereopsis
Po-Han Huang, Kevin Matzen, Johannes Kopf +2
cs.CVarXiv:1804.00650v12018Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling
Renrui Zhang, Rongyao Fang, Wei Zhang +5
cs.CVcs.CLarXiv:2111.03930v22021Multiscale Combinatorial Grouping for Image Segmentation and Object Proposal Generation
Jordi Pont-Tuset, Pablo Arbelaez, Jonathan T. Barron +2
cs.CVarXiv:1503.00848v42015