Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1,801 to 1,860 of 18,867

  1. POEM: Out-of-Distribution Detection with Posterior Sampling

    Yifei Ming, Ying Fan, Yixuan Li

    cs.LGcs.AIcs.CVarXiv:2206.13687v12022
  2. Learning Latent Subspaces in Variational Autoencoders

    Jack Klys, Jake Snell, Richard Zemel

    cs.LGcs.CVstat.MLarXiv:1812.06190v12018
  3. MG-GAN: A Multi-Generator Model Preventing Out-of-Distribution Samples in Pedestrian Trajectory Prediction

    Patrick Dendorfer, Sven Elflein, Laura Leal-Taixé

    cs.CVarXiv:2108.09274v12021
  4. MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds

    Zhenggang Tang, Yuchen Fan, Dilin Wang +4

    cs.CVcs.AIarXiv:2412.06974v12024
  5. Change is Everywhere: Single-Temporal Supervised Object Change Detection in Remote Sensing Imagery

    Zhuo Zheng, Ailong Ma, Liangpei Zhang +1

    cs.CVarXiv:2108.07002v32021
  6. MeshAnything: Artist-Created Mesh Generation with Autoregressive Transformers

    Yiwen Chen, Tong He, Di Huang +9

    cs.CVcs.AIarXiv:2406.10163v22024
  7. Variational Transformer Networks for Layout Generation

    Diego Martin Arroyo, Janis Postels, Federico Tombari

    cs.CVcs.LGarXiv:2104.02416v12021
  8. Self-taught Object Localization with Deep Networks

    Loris Bazzani, Alessandro Bergamo, Dragomir Anguelov +1

    cs.CVarXiv:1409.3964v72014
  9. Deep One-Class Classification via Interpolated Gaussian Descriptor

    Yuanhong Chen, Yu Tian, Guansong Pang +1

    cs.CVarXiv:2101.10043v52021
  10. NAG: Network for Adversary Generation

    Konda Reddy Mopuri, Utkarsh Ojha, Utsav Garg +1

    cs.CVcs.AIcs.LGarXiv:1712.03390v22017
  11. Learnable Boundary Guided Adversarial Training

    Jiequan Cui, Shu Liu, Liwei Wang +1

    cs.CVarXiv:2011.11164v22020
  12. On Learning Disentangled Representations for Gait Recognition

    Ziyuan Zhang, Luan Tran, Feng Liu +1

    cs.CVarXiv:1909.03051v12019
  13. Vehicle Re-identification Using Quadruple Directional Deep Learning Features

    Jianqing Zhu, Huanqiang Zeng, Jingchang Huang +4

    cs.CVarXiv:1811.05163v12018
  14. The AVA-Kinetics Localized Human Actions Video Dataset

    Ang Li, Meghana Thotakuri, David A. Ross +3

    cs.CVcs.LGeess.IVarXiv:2005.00214v22020
  15. SG-NN: Sparse Generative Neural Networks for Self-Supervised Scene Completion of RGB-D Scans

    Angela Dai, Christian Diller, Matthias Nießner

    cs.CVarXiv:1912.00036v22019
  16. Neural Head Reenactment with Latent Pose Descriptors

    Egor Burkov, Igor Pasechnik, Artur Grigorev +1

    cs.CVcs.LGarXiv:2004.12000v22020
  17. Embedded real-time stereo estimation via Semi-Global Matching on the GPU

    Daniel Hernandez-Juarez, Alejandro Chacón, Antonio Espinosa +3

    cs.CVarXiv:1610.04121v12016
  18. img2pose: Face Alignment and Detection via 6DoF, Face Pose Estimation

    Vítor Albiero, Xingyu Chen, Xi Yin +2

    cs.CVarXiv:2012.07791v22020
  19. MaskViT: Masked Visual Pre-Training for Video Prediction

    Agrim Gupta, Stephen Tian, Yunzhi Zhang +3

    cs.CVcs.LGcs.ROarXiv:2206.11894v22022
  20. ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accelerator Co-Design

    Haoran You, Zhanyi Sun, Huihong Shi +6

    cs.LGcs.ARcs.CVarXiv:2210.09573v32022
  21. AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers

    Sherwin Bahmani, Ivan Skorokhodov, Guocheng Qian +5

    cs.CVarXiv:2411.18673v42024
  22. 3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark

    Wufei Ma, Haoyu Chen, Guofeng Zhang +4

    cs.CVarXiv:2412.07825v42024
  23. Vague2Detect: Handling Ambiguous Prompts in Knowledge-Based Open-World Detection

    Ibrohimjon Muminov, Jihie Kim

    cs.CVcs.CLcs.LGarXiv:2609.09949v12026
  24. Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

    Duo Zheng, Shijia Huang, Liwei Wang

    cs.CVcs.CLarXiv:2412.00493v22024
  25. MLLMs Hallucinate when Information Distribution Drifts in Synergy Heads

    Meng'en Qin, Junye Chen, Jucheng Liu +3

    cs.CVcs.CLarXiv:2609.09206v12026
  26. Calorimetry with Deep Learning: Particle Simulation and Reconstruction for Collider Physics

    Dawit Belayneh, Federico Carminati, Amir Farbin +13

    physics.ins-detcs.CVcs.LGarXiv:1912.06794v32019
  27. Pre-training without Natural Images

    Hirokatsu Kataoka, Kazushige Okayasu, Asato Matsumoto +5

    cs.CVcs.LGarXiv:2101.08515v12021
  28. Generalizability of Machine Learning Models: Quantitative Evaluation of Three Methodological Pitfalls

    Farhad Maleki, Katie Ovens, Rajiv Gupta +3

    cs.LGcs.CVeess.IVarXiv:2202.01337v22022
  29. SENTRY: Selective Entropy Optimization via Committee Consistency for Unsupervised Domain Adaptation

    Viraj Prabhu, Shivam Khare, Deeksha Kartik +1

    cs.CVcs.LGarXiv:2012.11460v22020
  30. PhraseCut: Language-based Image Segmentation in the Wild

    Chenyun Wu, Zhe Lin, Scott Cohen +2

    cs.CVarXiv:2008.01187v12020
  31. CRN: Camera Radar Net for Accurate, Robust, Efficient 3D Perception

    Youngseok Kim, Juyeb Shin, Sanmin Kim +3

    cs.CVcs.AIcs.ROarXiv:2304.00670v32023
  32. Scaling Up Dynamic Human-Scene Interaction Modeling

    Nan Jiang, Zhiyuan Zhang, Hongjie Li +6

    cs.CVarXiv:2403.08629v22024
  33. ScrabbleGAN: Semi-Supervised Varying Length Handwritten Text Generation

    Sharon Fogel, Hadar Averbuch-Elor, Sarel Cohen +2

    cs.CVcs.CLcs.LGarXiv:2003.10557v12020
  34. Unidentified Video Objects: A Benchmark for Dense, Open-World Segmentation

    Weiyao Wang, Matt Feiszli, Heng Wang +1

    cs.CVarXiv:2104.04691v12021
  35. Mask Guided Matting via Progressive Refinement Network

    Qihang Yu, Jianming Zhang, He Zhang +5

    cs.CVarXiv:2012.06722v22020
  36. Multi-Label Learning from Single Positive Labels

    Elijah Cole, Oisin Mac Aodha, Titouan Lorieul +3

    cs.CVcs.LGarXiv:2106.09708v22021
  37. A Review on Deep Learning in Medical Image Reconstruction

    Haimiao Zhang, Bin Dong

    eess.IVcs.CVcs.LGarXiv:1906.10643v32019
  38. Debiased Self-Training for Semi-Supervised Learning

    Baixu Chen, Junguang Jiang, Ximei Wang +3

    cs.LGcs.CVarXiv:2202.07136v52022
  39. Learning Vision-Guided Quadrupedal Locomotion End-to-End with Cross-Modal Transformers

    Ruihan Yang, Minghao Zhang, Nicklas Hansen +2

    cs.LGcs.CVcs.ROarXiv:2107.03996v32021
  40. BAM! The Behance Artistic Media Dataset for Recognition Beyond Photography

    Michael J. Wilber, Chen Fang, Hailin Jin +3

    cs.CVarXiv:1704.08614v22017
  41. Category Contrast for Unsupervised Domain Adaptation in Visual Tasks

    Jiaxing Huang, Dayan Guan, Aoran Xiao +2

    cs.CVarXiv:2106.02885v32021
  42. Stagewise Unsupervised Domain Adaptation with Adversarial Self-Training for Road Segmentation of Remote Sensing Images

    Lefei Zhang, Meng Lan, Jing Zhang +1

    cs.CVarXiv:2108.12611v12021
  43. Understanding urban landuse from the above and ground perspectives: a deep learning, multimodal solution

    Shivangi Srivastava, John E. Vargas-Muñoz, Devis Tuia

    cs.CVarXiv:1905.01752v12019
  44. Multi-Person Pose Estimation with Local Joint-to-Person Associations

    Umar Iqbal, Juergen Gall

    cs.CVarXiv:1608.08526v22016
  45. NestedFormer: Nested Modality-Aware Transformer for Brain Tumor Segmentation

    Zhaohu Xing, Lequan Yu, Liang Wan +2

    eess.IVcs.CVarXiv:2208.14876v12022
  46. Deep Learning for Iris Recognition: A Survey

    Kien Nguyen, Hugo Proença, Fernando Alonso-Fernandez

    cs.CVcs.AIarXiv:2210.05866v12022
  47. Good Features to Correlate for Visual Tracking

    Erhan Gundogdu, A. Aydin Alatan

    cs.CVarXiv:1704.06326v22017
  48. VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning

    Han Lin, Abhay Zala, Jaemin Cho +1

    cs.CVcs.AIcs.CLarXiv:2309.15091v22023
  49. One Small Step for Generative AI, One Giant Leap for AGI: A Complete Survey on ChatGPT in AIGC Era

    Chaoning Zhang, Chenshuang Zhang, Chenghao Li +13

    cs.CYcs.AIcs.CLarXiv:2304.06488v12023
  50. Artificial-intelligence-based molecular classification of diffuse gliomas using rapid, label-free optical imaging

    Todd C. Hollon, Cheng Jiang, Asadur Chowdury +22

    cs.CVcs.AIcs.LGarXiv:2303.13610v12023
  51. Bunny-VisionPro: Real-Time Bimanual Dexterous Teleoperation for Imitation Learning

    Runyu Ding, Yuzhe Qin, Jiyue Zhu +5

    cs.ROcs.CVcs.LGarXiv:2407.03162v12024
  52. CAM-Convs: Camera-Aware Multi-Scale Convolutions for Single-View Depth

    Jose M. Facil, Benjamin Ummenhofer, Huizhong Zhou +3

    cs.CVarXiv:1904.02028v12019
  53. GiT: Graph Interactive Transformer for Vehicle Re-identification

    Fei Shen, Yi Xie, Jianqing Zhu +2

    cs.CVarXiv:2107.05475v32021
  54. VLX-VR: An Agentic-Aware Video Reasoning Model

    Sheng Li, Peng Liu, Qianqian Zhang +1

    cs.CLcs.CVarXiv:2609.09985v12026
  55. Reconstruction of hidden 3D shapes using diffuse reflections

    Otkrist Gupta, Andreas Velten, Thomas Willwacher +2

    physics.opticscs.CVarXiv:1203.4280v12012
  56. SAR-U-Net: squeeze-and-excitation block and atrous spatial pyramid pooling based residual U-Net for automatic liver segmentation in Computed Tomography

    Jinke Wang, Peiqing Lv, Haiying Wang +1

    eess.IVcs.CVcs.LGarXiv:2103.06419v32021
  57. A test case for application of convolutional neural networks to spatio-temporal climate data: Re-identifying clustered weather patterns

    Ashesh Chattopadhyay, Pedram Hassanzadeh, Saba Pasha

    physics.ao-phcs.CVcs.LGarXiv:1811.04817v12018
  58. From Zero-shot Learning to Conventional Supervised Classification: Unseen Visual Data Synthesis

    Yang Long, Li Liu, Ling Shao +3

    cs.CVarXiv:1705.01782v12017
  59. Deep Learning on Small Datasets without Pre-Training using Cosine Loss

    Björn Barz, Joachim Denzler

    cs.LGcs.CVstat.MLarXiv:1901.09054v22019
  60. Robust Learning Meets Generative Models: Can Proxy Distributions Improve Adversarial Robustness?

    Vikash Sehwag, Saeed Mahloujifar, Tinashe Handina +4

    cs.LGcs.CRcs.CVarXiv:2104.09425v32021