Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
12,481 to 12,540 of 18,857
Graph Attention Tracking
Dongyan Guo, Yanyan Shao, Ying Cui +3
cs.CVarXiv:2011.11204v12020Very Deep VAEs Generalize Autoregressive Models and Can Outperform Them on Images
Rewon Child
cs.LGcs.CVarXiv:2011.10650v22020VINet: Visual-Inertial Odometry as a Sequence-to-Sequence Learning Problem
Ronald Clark, Sen Wang, Hongkai Wen +2
cs.CVarXiv:1701.08376v22017Perception Prioritized Training of Diffusion Models
Jooyoung Choi, Jungbeom Lee, Chaehun Shin +3
cs.CVcs.LGarXiv:2204.00227v12022Diffusion Self-Guidance for Controllable Image Generation
Dave Epstein, Allan Jabri, Ben Poole +2
cs.CVcs.LGstat.MLarXiv:2306.00986v32023Disc-aware Ensemble Network for Glaucoma Screening from Fundus Image
Huazhu Fu, Jun Cheng, Yanwu Xu +4
cs.CVarXiv:1805.07549v12018LMDrive: Closed-Loop End-to-End Driving with Large Language Models
Hao Shao, Yuxuan Hu, Letian Wang +3
cs.CVcs.AIcs.ROarXiv:2312.07488v22023Diverse Part Discovery: Occluded Person Re-identification with Part-Aware Transformer
Yulin Li, Jianfeng He, Tianzhu Zhang +3
cs.CVarXiv:2106.04095v12021Exploring and Distilling Posterior and Prior Knowledge for Radiology Report Generation
Fenglin Liu, Xian Wu, Shen Ge +2
cs.CVcs.CLarXiv:2106.06963v22021Fast Patch-based Style Transfer of Arbitrary Style
Tian Qi Chen, Mark Schmidt
cs.CVcs.GRcs.LGarXiv:1612.04337v12016Style Normalization and Restitution for Generalizable Person Re-identification
Xin Jin, Cuiling Lan, Wenjun Zeng +2
cs.CVarXiv:2005.11037v12020RSVQA: Visual Question Answering for Remote Sensing Data
Sylvain Lobry, Diego Marcos, Jesse Murray +1
cs.CVarXiv:2003.07333v22020Forward and Backward Information Retention for Accurate Binary Neural Networks
Haotong Qin, Ruihao Gong, Xianglong Liu +4
cs.CVarXiv:1909.10788v42019RTNav: Towards Real-Time Zero-Shot Object Navigation
Easop Lee, Lingyu Zhang, Boyuan Chen
cs.ROcs.AIcs.CVarXiv:2608.26496v12026PCL: Proposal Cluster Learning for Weakly Supervised Object Detection
Peng Tang, Xinggang Wang, Song Bai +4
cs.CVarXiv:1807.03342v22018Data Augmentation using Random Image Cropping and Patching for Deep CNNs
Ryo Takahashi, Takashi Matsubara, Kuniaki Uehara
cs.CVcs.LGarXiv:1811.09030v22018A Review Paper: Noise Models in Digital Image Processing
Ajay Kumar Boyat, Brijendra Kumar Joshi
cs.CVarXiv:1505.03489v12015SE-SSD: Self-Ensembling Single-Stage Object Detector From Point Cloud
Wu Zheng, Weiliang Tang, Li Jiang +1
cs.CVarXiv:2104.09804v12021FastViT: A Fast Hybrid Vision Transformer using Structural Reparameterization
Pavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu +2
cs.CVarXiv:2303.14189v22023Unifying Flow, Stereo and Depth Estimation
Haofei Xu, Jing Zhang, Jianfei Cai +4
cs.CVarXiv:2211.05783v32022Parametric Contrastive Learning
Jiequan Cui, Zhisheng Zhong, Shu Liu +2
cs.CVarXiv:2107.12028v22021TransNeXt: Robust Foveal Visual Perception for Vision Transformers
Dai Shi
cs.CVcs.AIarXiv:2311.17132v32023Universal Source-Free Domain Adaptation
Jogendra Nath Kundu, Naveen Venkat, Rahul M +1
cs.CVcs.LGarXiv:2004.04393v12020Estimating Uncertainty and Interpretability in Deep Learning for Coronavirus (COVID-19) Detection
Biraja Ghoshal, Allan Tucker
eess.IVcs.CVcs.LGarXiv:2003.10769v22020FLatten Transformer: Vision Transformer using Focused Linear Attention
Dongchen Han, Xuran Pan, Yizeng Han +2
cs.CVarXiv:2308.00442v22023Optic Disc and Cup Segmentation Methods for Glaucoma Detection with Modification of U-Net Convolutional Neural Network
Artem Sevastopolsky
cs.CVstat.MLarXiv:1704.00979v12017RhythmNet: End-to-end Heart Rate Estimation from Face via Spatial-temporal Representation
Xuesong Niu, Shiguang Shan, Hu Han +1
cs.CVarXiv:1910.11515v22019RoMa: Robust Dense Feature Matching
Johan Edstedt, Qiyu Sun, Georg Bökman +2
cs.CVarXiv:2305.15404v22023xBD: A Dataset for Assessing Building Damage from Satellite Imagery
Ritwik Gupta, Richard Hosfelt, Sandra Sajeev +6
cs.CVarXiv:1911.09296v12019History Aware Multimodal Transformer for Vision-and-Language Navigation
Shizhe Chen, Pierre-Louis Guhur, Cordelia Schmid +1
cs.CVcs.AIarXiv:2110.13309v22021Instance-sensitive Fully Convolutional Networks
Jifeng Dai, Kaiming He, Yi Li +2
cs.CVarXiv:1603.08678v12016Fast and Robust Iterative Closest Point
Juyong Zhang, Yuxin Yao, Bailin Deng
cs.CVcs.GRcs.ROarXiv:2007.07627v320204DSynth: Controllable Procedural World Synthesis for Dynamic Embodied Simulation
Zehao Qi, Haochen Luo, Jia-Wang Bian +2
cs.ROcs.CVarXiv:2608.26947v12026MVC-Bench: Benchmarking Calibration of Medical Vision-Language Models
Ashshak Sharifdeen, Shihab Aaqil Ahamed, Ufaq Khan +5
cs.CVarXiv:2608.27004v12026Recurrence is required to capture the representational dynamics of the human visual system
Tim C Kietzmann, Courtney J Spoerer, Lynn Sörensen +3
q-bio.NCcs.CVcs.LGarXiv:1903.05946v22019A Joint Sequence Fusion Model for Video Question Answering and Retrieval
Youngjae Yu, Jongseok Kim, Gunhee Kim
cs.CVarXiv:1808.02559v12018Resolving 3D Human Pose Ambiguities with 3D Scene Constraints
Mohamed Hassan, Vasileios Choutas, Dimitrios Tzionas +1
cs.CVarXiv:1908.06963v12019Practical Stereo Matching via Cascaded Recurrent Network with Adaptive Correlation
Jiankun Li, Peisen Wang, Pengfei Xiong +6
cs.CVarXiv:2203.11483v12022Highly comparative time-series analysis: The empirical structure of time series and their methods
Ben D. Fulcher, Max A. Little, Nick S. Jones
physics.data-ancs.CVphysics.bio-pharXiv:1304.1209v12013Exposure: A White-Box Photo Post-Processing Framework
Yuanming Hu, Hao He, Chenxi Xu +2
cs.GRcs.CVarXiv:1709.09602v22017Semantic Segmentation of Earth Observation Data Using Multimodal and Multi-scale Deep Networks
Nicolas Audebert, Bertrand Le Saux, Sébastien Lefèvre
cs.CVcs.NEarXiv:1609.06846v12016Temporal Sensitivity Analysis of Tessera Embeddings
Julia Guerrero-Viu, Alex López-Cifuentes, Ignacio Pérez-Villar +1
cs.CVarXiv:2608.27175v12026Domain Adaptation for Image Dehazing
Yuanjie Shao, Lerenhan Li, Wenqi Ren +2
cs.CVarXiv:2005.04668v12020Towards Multimodal Sarcasm Detection (An _Obviously_ Perfect Paper)
Santiago Castro, Devamanyu Hazarika, Verónica Pérez-Rosas +3
cs.CLcs.CVarXiv:1906.01815v12019Reconstructing Hands in 3D with Transformers
Georgios Pavlakos, Dandan Shan, Ilija Radosavovic +3
cs.CVarXiv:2312.05251v12023Span-based Localizing Network for Natural Language Video Localization
Hao Zhang, Aixin Sun, Wei Jing +1
cs.CLcs.CVarXiv:2004.13931v22020Fully-adaptive Feature Sharing in Multi-Task Networks with Applications in Person Attribute Classification
Yongxi Lu, Abhishek Kumar, Shuangfei Zhai +3
cs.CVcs.LGarXiv:1611.05377v12016ShellNet: Efficient Point Cloud Convolutional Neural Networks using Concentric Shells Statistics
Zhiyuan Zhang, Binh-Son Hua, Sai-Kit Yeung
cs.CVarXiv:1908.06295v12019Learning Shape Abstractions by Assembling Volumetric Primitives
Shubham Tulsiani, Hao Su, Leonidas J. Guibas +2
cs.CVarXiv:1612.00404v42016Single-View View Synthesis with Multiplane Images
Richard Tucker, Noah Snavely
cs.CVcs.GRarXiv:2004.11364v12020Knowledge distillation: A good teacher is patient and consistent
Lucas Beyer, Xiaohua Zhai, Amélie Royer +3
cs.CVcs.AIcs.LGarXiv:2106.05237v22021NeRV: Neural Representations for Videos
Hao Chen, Bo He, Hanyu Wang +3
cs.CVeess.IVarXiv:2110.13903v12021Unsupervised Learning of 3D Structure from Images
Danilo Jimenez Rezende, S. M. Ali Eslami, Shakir Mohamed +3
cs.CVcs.LGstat.MLarXiv:1607.00662v22016MMTM: Multimodal Transfer Module for CNN Fusion
Hamid Reza Vaezi Joze, Amirreza Shaban, Michael L. Iuzzolino +1
cs.CVcs.LGarXiv:1911.08670v22019Topology-Preserving Deep Image Segmentation
Xiaoling Hu, Li Fuxin, Dimitris Samaras +1
cs.CVcs.CGarXiv:1906.05404v12019Point-Based Multi-View Stereo Network
Rui Chen, Songfang Han, Jing Xu +1
cs.CVarXiv:1908.04422v12019Neural Unsigned Distance Fields for Implicit Function Learning
Julian Chibane, Aymen Mir, Gerard Pons-Moll
cs.CVcs.LGarXiv:2010.13938v12020RAVEN: A Dataset for Relational and Analogical Visual rEasoNing
Chi Zhang, Feng Gao, Baoxiong Jia +2
cs.CVcs.AIcs.LGarXiv:1903.02741v12019ABCNet: Real-time Scene Text Spotting with Adaptive Bezier-Curve Network
Yuliang Liu, Hao Chen, Chunhua Shen +3
cs.CVarXiv:2002.10200v22020SpinNet: Learning a General Surface Descriptor for 3D Point Cloud Registration
Sheng Ao, Qingyong Hu, Bo Yang +2
cs.CVcs.AIcs.LGarXiv:2011.12149v22020