Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

12,481 to 12,540 of 18,857

  1. Graph Attention Tracking

    Dongyan Guo, Yanyan Shao, Ying Cui +3

    cs.CVarXiv:2011.11204v12020
  2. Very Deep VAEs Generalize Autoregressive Models and Can Outperform Them on Images

    Rewon Child

    cs.LGcs.CVarXiv:2011.10650v22020
  3. VINet: Visual-Inertial Odometry as a Sequence-to-Sequence Learning Problem

    Ronald Clark, Sen Wang, Hongkai Wen +2

    cs.CVarXiv:1701.08376v22017
  4. Perception Prioritized Training of Diffusion Models

    Jooyoung Choi, Jungbeom Lee, Chaehun Shin +3

    cs.CVcs.LGarXiv:2204.00227v12022
  5. Diffusion Self-Guidance for Controllable Image Generation

    Dave Epstein, Allan Jabri, Ben Poole +2

    cs.CVcs.LGstat.MLarXiv:2306.00986v32023
  6. Disc-aware Ensemble Network for Glaucoma Screening from Fundus Image

    Huazhu Fu, Jun Cheng, Yanwu Xu +4

    cs.CVarXiv:1805.07549v12018
  7. LMDrive: Closed-Loop End-to-End Driving with Large Language Models

    Hao Shao, Yuxuan Hu, Letian Wang +3

    cs.CVcs.AIcs.ROarXiv:2312.07488v22023
  8. Diverse Part Discovery: Occluded Person Re-identification with Part-Aware Transformer

    Yulin Li, Jianfeng He, Tianzhu Zhang +3

    cs.CVarXiv:2106.04095v12021
  9. Exploring and Distilling Posterior and Prior Knowledge for Radiology Report Generation

    Fenglin Liu, Xian Wu, Shen Ge +2

    cs.CVcs.CLarXiv:2106.06963v22021
  10. Fast Patch-based Style Transfer of Arbitrary Style

    Tian Qi Chen, Mark Schmidt

    cs.CVcs.GRcs.LGarXiv:1612.04337v12016
  11. Style Normalization and Restitution for Generalizable Person Re-identification

    Xin Jin, Cuiling Lan, Wenjun Zeng +2

    cs.CVarXiv:2005.11037v12020
  12. RSVQA: Visual Question Answering for Remote Sensing Data

    Sylvain Lobry, Diego Marcos, Jesse Murray +1

    cs.CVarXiv:2003.07333v22020
  13. Forward and Backward Information Retention for Accurate Binary Neural Networks

    Haotong Qin, Ruihao Gong, Xianglong Liu +4

    cs.CVarXiv:1909.10788v42019
  14. RTNav: Towards Real-Time Zero-Shot Object Navigation

    Easop Lee, Lingyu Zhang, Boyuan Chen

    cs.ROcs.AIcs.CVarXiv:2608.26496v12026
  15. PCL: Proposal Cluster Learning for Weakly Supervised Object Detection

    Peng Tang, Xinggang Wang, Song Bai +4

    cs.CVarXiv:1807.03342v22018
  16. Data Augmentation using Random Image Cropping and Patching for Deep CNNs

    Ryo Takahashi, Takashi Matsubara, Kuniaki Uehara

    cs.CVcs.LGarXiv:1811.09030v22018
  17. A Review Paper: Noise Models in Digital Image Processing

    Ajay Kumar Boyat, Brijendra Kumar Joshi

    cs.CVarXiv:1505.03489v12015
  18. SE-SSD: Self-Ensembling Single-Stage Object Detector From Point Cloud

    Wu Zheng, Weiliang Tang, Li Jiang +1

    cs.CVarXiv:2104.09804v12021
  19. FastViT: A Fast Hybrid Vision Transformer using Structural Reparameterization

    Pavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu +2

    cs.CVarXiv:2303.14189v22023
  20. Unifying Flow, Stereo and Depth Estimation

    Haofei Xu, Jing Zhang, Jianfei Cai +4

    cs.CVarXiv:2211.05783v32022
  21. Parametric Contrastive Learning

    Jiequan Cui, Zhisheng Zhong, Shu Liu +2

    cs.CVarXiv:2107.12028v22021
  22. TransNeXt: Robust Foveal Visual Perception for Vision Transformers

    Dai Shi

    cs.CVcs.AIarXiv:2311.17132v32023
  23. Universal Source-Free Domain Adaptation

    Jogendra Nath Kundu, Naveen Venkat, Rahul M +1

    cs.CVcs.LGarXiv:2004.04393v12020
  24. Estimating Uncertainty and Interpretability in Deep Learning for Coronavirus (COVID-19) Detection

    Biraja Ghoshal, Allan Tucker

    eess.IVcs.CVcs.LGarXiv:2003.10769v22020
  25. FLatten Transformer: Vision Transformer using Focused Linear Attention

    Dongchen Han, Xuran Pan, Yizeng Han +2

    cs.CVarXiv:2308.00442v22023
  26. Optic Disc and Cup Segmentation Methods for Glaucoma Detection with Modification of U-Net Convolutional Neural Network

    Artem Sevastopolsky

    cs.CVstat.MLarXiv:1704.00979v12017
  27. RhythmNet: End-to-end Heart Rate Estimation from Face via Spatial-temporal Representation

    Xuesong Niu, Shiguang Shan, Hu Han +1

    cs.CVarXiv:1910.11515v22019
  28. RoMa: Robust Dense Feature Matching

    Johan Edstedt, Qiyu Sun, Georg Bökman +2

    cs.CVarXiv:2305.15404v22023
  29. xBD: A Dataset for Assessing Building Damage from Satellite Imagery

    Ritwik Gupta, Richard Hosfelt, Sandra Sajeev +6

    cs.CVarXiv:1911.09296v12019
  30. History Aware Multimodal Transformer for Vision-and-Language Navigation

    Shizhe Chen, Pierre-Louis Guhur, Cordelia Schmid +1

    cs.CVcs.AIarXiv:2110.13309v22021
  31. Instance-sensitive Fully Convolutional Networks

    Jifeng Dai, Kaiming He, Yi Li +2

    cs.CVarXiv:1603.08678v12016
  32. Fast and Robust Iterative Closest Point

    Juyong Zhang, Yuxin Yao, Bailin Deng

    cs.CVcs.GRcs.ROarXiv:2007.07627v32020
  33. 4DSynth: Controllable Procedural World Synthesis for Dynamic Embodied Simulation

    Zehao Qi, Haochen Luo, Jia-Wang Bian +2

    cs.ROcs.CVarXiv:2608.26947v12026
  34. MVC-Bench: Benchmarking Calibration of Medical Vision-Language Models

    Ashshak Sharifdeen, Shihab Aaqil Ahamed, Ufaq Khan +5

    cs.CVarXiv:2608.27004v12026
  35. Recurrence is required to capture the representational dynamics of the human visual system

    Tim C Kietzmann, Courtney J Spoerer, Lynn Sörensen +3

    q-bio.NCcs.CVcs.LGarXiv:1903.05946v22019
  36. A Joint Sequence Fusion Model for Video Question Answering and Retrieval

    Youngjae Yu, Jongseok Kim, Gunhee Kim

    cs.CVarXiv:1808.02559v12018
  37. Resolving 3D Human Pose Ambiguities with 3D Scene Constraints

    Mohamed Hassan, Vasileios Choutas, Dimitrios Tzionas +1

    cs.CVarXiv:1908.06963v12019
  38. Practical Stereo Matching via Cascaded Recurrent Network with Adaptive Correlation

    Jiankun Li, Peisen Wang, Pengfei Xiong +6

    cs.CVarXiv:2203.11483v12022
  39. Highly comparative time-series analysis: The empirical structure of time series and their methods

    Ben D. Fulcher, Max A. Little, Nick S. Jones

    physics.data-ancs.CVphysics.bio-pharXiv:1304.1209v12013
  40. Exposure: A White-Box Photo Post-Processing Framework

    Yuanming Hu, Hao He, Chenxi Xu +2

    cs.GRcs.CVarXiv:1709.09602v22017
  41. Semantic Segmentation of Earth Observation Data Using Multimodal and Multi-scale Deep Networks

    Nicolas Audebert, Bertrand Le Saux, Sébastien Lefèvre

    cs.CVcs.NEarXiv:1609.06846v12016
  42. Temporal Sensitivity Analysis of Tessera Embeddings

    Julia Guerrero-Viu, Alex López-Cifuentes, Ignacio Pérez-Villar +1

    cs.CVarXiv:2608.27175v12026
  43. Domain Adaptation for Image Dehazing

    Yuanjie Shao, Lerenhan Li, Wenqi Ren +2

    cs.CVarXiv:2005.04668v12020
  44. Towards Multimodal Sarcasm Detection (An _Obviously_ Perfect Paper)

    Santiago Castro, Devamanyu Hazarika, Verónica Pérez-Rosas +3

    cs.CLcs.CVarXiv:1906.01815v12019
  45. Reconstructing Hands in 3D with Transformers

    Georgios Pavlakos, Dandan Shan, Ilija Radosavovic +3

    cs.CVarXiv:2312.05251v12023
  46. Span-based Localizing Network for Natural Language Video Localization

    Hao Zhang, Aixin Sun, Wei Jing +1

    cs.CLcs.CVarXiv:2004.13931v22020
  47. Fully-adaptive Feature Sharing in Multi-Task Networks with Applications in Person Attribute Classification

    Yongxi Lu, Abhishek Kumar, Shuangfei Zhai +3

    cs.CVcs.LGarXiv:1611.05377v12016
  48. ShellNet: Efficient Point Cloud Convolutional Neural Networks using Concentric Shells Statistics

    Zhiyuan Zhang, Binh-Son Hua, Sai-Kit Yeung

    cs.CVarXiv:1908.06295v12019
  49. Learning Shape Abstractions by Assembling Volumetric Primitives

    Shubham Tulsiani, Hao Su, Leonidas J. Guibas +2

    cs.CVarXiv:1612.00404v42016
  50. Single-View View Synthesis with Multiplane Images

    Richard Tucker, Noah Snavely

    cs.CVcs.GRarXiv:2004.11364v12020
  51. Knowledge distillation: A good teacher is patient and consistent

    Lucas Beyer, Xiaohua Zhai, Amélie Royer +3

    cs.CVcs.AIcs.LGarXiv:2106.05237v22021
  52. NeRV: Neural Representations for Videos

    Hao Chen, Bo He, Hanyu Wang +3

    cs.CVeess.IVarXiv:2110.13903v12021
  53. Unsupervised Learning of 3D Structure from Images

    Danilo Jimenez Rezende, S. M. Ali Eslami, Shakir Mohamed +3

    cs.CVcs.LGstat.MLarXiv:1607.00662v22016
  54. MMTM: Multimodal Transfer Module for CNN Fusion

    Hamid Reza Vaezi Joze, Amirreza Shaban, Michael L. Iuzzolino +1

    cs.CVcs.LGarXiv:1911.08670v22019
  55. Topology-Preserving Deep Image Segmentation

    Xiaoling Hu, Li Fuxin, Dimitris Samaras +1

    cs.CVcs.CGarXiv:1906.05404v12019
  56. Point-Based Multi-View Stereo Network

    Rui Chen, Songfang Han, Jing Xu +1

    cs.CVarXiv:1908.04422v12019
  57. Neural Unsigned Distance Fields for Implicit Function Learning

    Julian Chibane, Aymen Mir, Gerard Pons-Moll

    cs.CVcs.LGarXiv:2010.13938v12020
  58. RAVEN: A Dataset for Relational and Analogical Visual rEasoNing

    Chi Zhang, Feng Gao, Baoxiong Jia +2

    cs.CVcs.AIcs.LGarXiv:1903.02741v12019
  59. ABCNet: Real-time Scene Text Spotting with Adaptive Bezier-Curve Network

    Yuliang Liu, Hao Chen, Chunhua Shen +3

    cs.CVarXiv:2002.10200v22020
  60. SpinNet: Learning a General Surface Descriptor for 3D Point Cloud Registration

    Sheng Ao, Qingyong Hu, Bo Yang +2

    cs.CVcs.AIcs.LGarXiv:2011.12149v22020