Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1,741 to 1,800 of 18,848

  1. Photorealistic Image Synthesis for Object Instance Detection

    Tomas Hodan, Vibhav Vineet, Ran Gal +6

    cs.CVcs.AIcs.ROarXiv:1902.03334v12019
  2. Temporal Pyramid Pooling Based Convolutional Neural Networks for Action Recognition

    Peng Wang, Yuanzhouhan Cao, Chunhua Shen +2

    cs.CVarXiv:1503.01224v22015
  3. Controlling Vision-Language Models for Multi-Task Image Restoration

    Ziwei Luo, Fredrik K. Gustafsson, Zheng Zhao +2

    cs.CVarXiv:2310.01018v22023
  4. Active Learning for Deep Detection Neural Networks

    Hamed H. Aghdam, Abel Gonzalez-Garcia, Joost van de Weijer +1

    cs.CVcs.LGarXiv:1911.09168v12019
  5. Unsupervised Domain Adaptation via Disentangled Representations: Application to Cross-Modality Liver Segmentation

    Junlin Yang, Nicha C. Dvornek, Fan Zhang +3

    eess.IVcs.CVarXiv:1907.13590v22019
  6. Newtonian Image Understanding: Unfolding the Dynamics of Objects in Static Images

    Roozbeh Mottaghi, Hessam Bagherinezhad, Mohammad Rastegari +1

    cs.CVarXiv:1511.04048v12015
  7. SegICP: Integrated Deep Semantic Segmentation and Pose Estimation

    Jay M. Wong, Vincent Kee, Tiffany Le +10

    cs.ROcs.CVarXiv:1703.01661v22017
  8. Learning Descriptor Networks for 3D Shape Synthesis and Analysis

    Jianwen Xie, Zilong Zheng, Ruiqi Gao +3

    cs.CVarXiv:1804.00586v12018
  9. Learning representations of irregular particle-detector geometry with distance-weighted graph networks

    Shah Rukh Qasim, Jan Kieseler, Yutaro Iiyama +1

    physics.data-ancs.CVcs.LGarXiv:1902.07987v22019
  10. Bridging 2D and 3D Segmentation Networks for Computation Efficient Volumetric Medical Image Segmentation: An Empirical Study of 2.5D Solutions

    Yichi Zhang, Qingcheng Liao, Le Ding +1

    eess.IVcs.CVarXiv:2010.06163v22020
  11. Polyp-SAM: Transfer SAM for Polyp Segmentation

    Yuheng Li, Mingzhe Hu, Xiaofeng Yang

    eess.IVcs.CVarXiv:2305.00293v12023
  12. An Experimental Study of Deep Convolutional Features For Iris Recognition

    Shervin Minaee, Amirali Abdolrashidi, Yao Wang

    cs.CVcs.LGarXiv:1702.01334v12017
  13. Designing BERT for Convolutional Networks: Sparse and Hierarchical Masked Modeling

    Keyu Tian, Yi Jiang, Qishuai Diao +3

    cs.CVcs.AIcs.LGarXiv:2301.03580v22023
  14. DF26: We Cannot Tell Fake From Real Anymore

    Severyn Shykula, Andrii Yermakov, Ivan Samarskyi +3

    cs.CVarXiv:2609.07369v12026
  15. From W-Net to CDGAN: Bi-temporal Change Detection via Deep Learning Techniques

    Bin Hou, Qingjie Liu, Heng Wang +1

    cs.CVeess.IVarXiv:2003.06583v12020
  16. Swapout: Learning an ensemble of deep architectures

    Saurabh Singh, Derek Hoiem, David Forsyth

    cs.CVcs.LGcs.NEarXiv:1605.06465v12016
  17. X-Net: Brain Stroke Lesion Segmentation Based on Depthwise Separable Convolution and Long-range Dependencies

    Kehan Qi, Hao Yang, Cheng Li +4

    eess.IVcs.CVarXiv:1907.07000v22019
  18. Optimized U-Net for Brain Tumor Segmentation

    Michał Futrega, Alexandre Milesi, Michal Marcinkiewicz +1

    eess.IVcs.CVcs.LGarXiv:2110.03352v22021
  19. AdaBatch: Adaptive Batch Sizes for Training Deep Neural Networks

    Aditya Devarakonda, Maxim Naumov, Michael Garland

    cs.LGcs.CVcs.DCarXiv:1712.02029v22017
  20. An automatic multi-tissue human fetal brain segmentation benchmark using the Fetal Tissue Annotation Dataset

    Kelly Payette, Priscille de Dumast, Hamza Kebiri +17

    eess.IVcs.CVarXiv:2010.15526v42020
  21. Detailed 2D-3D Joint Representation for Human-Object Interaction

    Yong-Lu Li, Xinpeng Liu, Han Lu +4

    cs.CVcs.LGarXiv:2004.08154v22020
  22. A Deeper Look into Aleatoric and Epistemic Uncertainty Disentanglement

    Matias Valdenegro-Toro, Daniel Saromo

    cs.LGcs.CVarXiv:2204.09308v12022
  23. i3DMM: Deep Implicit 3D Morphable Model of Human Heads

    Tarun Yenamandra, Ayush Tewari, Florian Bernard +4

    cs.CVcs.GRcs.LGarXiv:2011.14143v12020
  24. Multi-scale Attention Network for Single Image Super-Resolution

    Yan Wang, Yusen Li, Gang Wang +1

    eess.IVcs.CVarXiv:2209.14145v32022
  25. Augmented LiDAR Simulator for Autonomous Driving

    Jin Fang, Dingfu Zhou, Feilong Yan +5

    cs.CVarXiv:1811.07112v22018
  26. MEAL: Multi-Model Ensemble via Adversarial Learning

    Zhiqiang Shen, Zhankui He, Xiangyang Xue

    cs.CVcs.AIcs.LGarXiv:1812.02425v22018
  27. Feature Importance Ranking for Deep Learning

    Maksymilian Wojtas, Ke Chen

    cs.LGcs.AIcs.CVarXiv:2010.08973v12020
  28. 3D Neural Scene Representations for Visuomotor Control

    Yunzhu Li, Shuang Li, Vincent Sitzmann +2

    cs.ROcs.CVcs.LGarXiv:2107.04004v22021
  29. Image-Specific Information Suppression and Implicit Local Alignment for Text-based Person Search

    Shuanglin Yan, Hao Tang, Liyan Zhang +1

    cs.CVarXiv:2208.14365v22022
  30. Depth Pooling Based Large-scale 3D Action Recognition with Convolutional Neural Networks

    Pichao Wang, Wanqing Li, Zhimin Gao +2

    cs.CVarXiv:1804.01194v22018
  31. Efficient Fairness Auditing Across Guidance Scales in Text-to-Image Diffusion Models via Causal Abstraction

    Nabila Tasfiha Rahman, Rajatsubhra Chakraborty, Depeng Xu +1

    cs.LGcs.CVarXiv:2609.09486v12026
  32. LoopReg: Self-supervised Learning of Implicit Surface Correspondences, Pose and Shape for 3D Human Mesh Registration

    Bharat Lal Bhatnagar, Cristian Sminchisescu, Christian Theobalt +1

    cs.CVarXiv:2010.12447v22020
  33. Phase-Shifting Coder: Predicting Accurate Orientation in Oriented Object Detection

    Yi Yu, Feipeng Da

    cs.CVarXiv:2211.06368v22022
  34. Multimodal Attention-based Deep Learning for Alzheimer's Disease Diagnosis

    Michal Golovanevsky, Carsten Eickhoff, Ritambhara Singh

    cs.LGcs.CVeess.IVarXiv:2206.08826v22022
  35. Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

    Zhuoyan Luo, Fengyuan Shi, Yixiao Ge +3

    cs.CVcs.AIarXiv:2409.04410v32024
  36. Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving

    Bo Jiang, Shaoyu Chen, Bencheng Liao +6

    cs.CVcs.ROarXiv:2410.22313v12024
  37. Devil is in the Edges: Learning Semantic Boundaries from Noisy Annotations

    David Acuna, Amlan Kar, Sanja Fidler

    cs.CVcs.AIarXiv:1904.07934v22019
  38. Mining Latent Classes for Few-shot Segmentation

    Lihe Yang, Wei Zhuo, Lei Qi +2

    cs.CVarXiv:2103.15402v32021
  39. Understanding the Latent Space of Diffusion Models through the Lens of Riemannian Geometry

    Yong-Hyun Park, Mingi Kwon, Jaewoong Choi +2

    cs.CVarXiv:2307.12868v22023
  40. A Modulation Module for Multi-task Learning with Applications in Image Retrieval

    Xiangyun Zhao, Haoxiang Li, Xiaohui Shen +2

    cs.CVarXiv:1807.06708v22018
  41. DM-VIO: Delayed Marginalization Visual-Inertial Odometry

    Lukas von Stumberg, Daniel Cremers

    cs.CVcs.ROarXiv:2201.04114v12022
  42. POEM: Out-of-Distribution Detection with Posterior Sampling

    Yifei Ming, Ying Fan, Yixuan Li

    cs.LGcs.AIcs.CVarXiv:2206.13687v12022
  43. Learning Latent Subspaces in Variational Autoencoders

    Jack Klys, Jake Snell, Richard Zemel

    cs.LGcs.CVstat.MLarXiv:1812.06190v12018
  44. MG-GAN: A Multi-Generator Model Preventing Out-of-Distribution Samples in Pedestrian Trajectory Prediction

    Patrick Dendorfer, Sven Elflein, Laura Leal-Taixé

    cs.CVarXiv:2108.09274v12021
  45. MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds

    Zhenggang Tang, Yuchen Fan, Dilin Wang +4

    cs.CVcs.AIarXiv:2412.06974v12024
  46. Change is Everywhere: Single-Temporal Supervised Object Change Detection in Remote Sensing Imagery

    Zhuo Zheng, Ailong Ma, Liangpei Zhang +1

    cs.CVarXiv:2108.07002v32021
  47. MeshAnything: Artist-Created Mesh Generation with Autoregressive Transformers

    Yiwen Chen, Tong He, Di Huang +9

    cs.CVcs.AIarXiv:2406.10163v22024
  48. Variational Transformer Networks for Layout Generation

    Diego Martin Arroyo, Janis Postels, Federico Tombari

    cs.CVcs.LGarXiv:2104.02416v12021
  49. Self-taught Object Localization with Deep Networks

    Loris Bazzani, Alessandro Bergamo, Dragomir Anguelov +1

    cs.CVarXiv:1409.3964v72014
  50. Deep One-Class Classification via Interpolated Gaussian Descriptor

    Yuanhong Chen, Yu Tian, Guansong Pang +1

    cs.CVarXiv:2101.10043v52021
  51. NAG: Network for Adversary Generation

    Konda Reddy Mopuri, Utkarsh Ojha, Utsav Garg +1

    cs.CVcs.AIcs.LGarXiv:1712.03390v22017
  52. Learnable Boundary Guided Adversarial Training

    Jiequan Cui, Shu Liu, Liwei Wang +1

    cs.CVarXiv:2011.11164v22020
  53. On Learning Disentangled Representations for Gait Recognition

    Ziyuan Zhang, Luan Tran, Feng Liu +1

    cs.CVarXiv:1909.03051v12019
  54. Vehicle Re-identification Using Quadruple Directional Deep Learning Features

    Jianqing Zhu, Huanqiang Zeng, Jingchang Huang +4

    cs.CVarXiv:1811.05163v12018
  55. The AVA-Kinetics Localized Human Actions Video Dataset

    Ang Li, Meghana Thotakuri, David A. Ross +3

    cs.CVcs.LGeess.IVarXiv:2005.00214v22020
  56. SG-NN: Sparse Generative Neural Networks for Self-Supervised Scene Completion of RGB-D Scans

    Angela Dai, Christian Diller, Matthias Nießner

    cs.CVarXiv:1912.00036v22019
  57. Neural Head Reenactment with Latent Pose Descriptors

    Egor Burkov, Igor Pasechnik, Artur Grigorev +1

    cs.CVcs.LGarXiv:2004.12000v22020
  58. Embedded real-time stereo estimation via Semi-Global Matching on the GPU

    Daniel Hernandez-Juarez, Alejandro Chacón, Antonio Espinosa +3

    cs.CVarXiv:1610.04121v12016
  59. img2pose: Face Alignment and Detection via 6DoF, Face Pose Estimation

    Vítor Albiero, Xingyu Chen, Xi Yin +2

    cs.CVarXiv:2012.07791v22020
  60. MaskViT: Masked Visual Pre-Training for Video Prediction

    Agrim Gupta, Stephen Tian, Yunzhi Zhang +3

    cs.CVcs.LGcs.ROarXiv:2206.11894v22022