Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

12,601 to 12,660 of 18,866

  1. Beyond Atomic Layouts: Compositional Design Understanding with Vision-Language Models

    Yiyang Huang, Zhaowen Wang, Simon Jenni +4

    cs.CVarXiv:2608.26716v12026
  2. Attention-based Dropout Layer for Weakly Supervised Object Localization

    Junsuk Choe, Hyunjung Shim

    cs.CVarXiv:1908.10028v12019
  3. Deep Attributes Driven Multi-Camera Person Re-identification

    Chi Su, Shiliang Zhang, Junliang Xing +2

    cs.CVarXiv:1605.03259v22016
  4. Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning

    Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo +5

    cs.CVcs.AIcs.CLarXiv:2302.14115v22023
  5. CrossTransformers: spatially-aware few-shot transfer

    Carl Doersch, Ankush Gupta, Andrew Zisserman

    cs.CVarXiv:2007.11498v52020
  6. Multi-Person Human Motion Forecasting in Complex Scenes

    Serdar Ozsoy, Lars Doorenbos, Juergen Gall

    cs.CVcs.AIarXiv:2608.27039v12026
  7. Fixup Initialization: Residual Learning Without Normalization

    Hongyi Zhang, Yann N. Dauphin, Tengyu Ma

    cs.LGcs.CVstat.MLarXiv:1901.09321v22019
  8. Playing with Duality: An Overview of Recent Primal-Dual Approaches for Solving Large-Scale Optimization Problems

    Nikos Komodakis, Jean-Christophe Pesquet

    math.NAcs.CVcs.LGarXiv:1406.5429v22014
  9. Function4D: Real-time Human Volumetric Capture from Very Sparse Consumer RGBD Sensors

    Tao Yu, Zerong Zheng, Kaiwen Guo +3

    cs.CVarXiv:2105.01859v22021
  10. Iterative Geometry Encoding Volume for Stereo Matching

    Gangwei Xu, Xianqi Wang, Xiaohuan Ding +1

    cs.CVarXiv:2303.06615v22023
  11. The Platonic Representation Hypothesis

    Minyoung Huh, Brian Cheung, Tongzhou Wang +1

    cs.LGcs.AIcs.CVarXiv:2405.07987v52024
  12. Expression, Affect, Action Unit Recognition: Aff-Wild2, Multi-Task Learning and ArcFace

    Dimitrios Kollias, Stefanos Zafeiriou

    cs.CVcs.HCcs.LGarXiv:1910.04855v12019
  13. Is Pseudo-Lidar needed for Monocular 3D Object detection?

    Dennis Park, Rares Ambrus, Vitor Guizilini +2

    cs.CVarXiv:2108.06417v12021
  14. Weighted Schatten $p$-Norm Minimization for Image Denoising and Background Subtraction

    Yuan Xie, Shuhang Gu, Yan Liu +3

    cs.CVarXiv:1512.01003v12015
  15. Improving Semantic Segmentation via Video Propagation and Label Relaxation

    Yi Zhu, Karan Sapra, Fitsum A. Reda +4

    cs.CVcs.AIcs.MMarXiv:1812.01593v32018
  16. PointCleanNet: Learning to Denoise and Remove Outliers from Dense Point Clouds

    Marie-Julie Rakotosaona, Vittorio La Barbera, Paul Guerrero +2

    cs.GRcs.CVarXiv:1901.01060v32019
  17. Human Motion Diffusion as a Generative Prior

    Yonatan Shafir, Guy Tevet, Roy Kapon +1

    cs.CVcs.GRarXiv:2303.01418v32023
  18. Skeleton-based Action Recognition via Spatial and Temporal Transformer Networks

    Chiara Plizzari, Marco Cannici, Matteo Matteucci

    cs.CVarXiv:2008.07404v42020
  19. FastComposer: Tuning-Free Multi-Subject Image Generation with Localized Attention

    Guangxuan Xiao, Tianwei Yin, William T. Freeman +2

    cs.CVarXiv:2305.10431v22023
  20. Dataset Condensation with Differentiable Siamese Augmentation

    Bo Zhao, Hakan Bilen

    cs.LGcs.CVarXiv:2102.08259v22021
  21. Image Data Augmentation for Deep Learning: A Survey

    Suorong Yang, Weikang Xiao, Mengchen Zhang +3

    cs.CVarXiv:2204.08610v22022
  22. Models Genesis: Generic Autodidactic Models for 3D Medical Image Analysis

    Zongwei Zhou, Vatsal Sodha, Md Mahfuzur Rahman Siddiquee +4

    eess.IVcs.CVarXiv:1908.06912v12019
  23. Convolutional Networks with Adaptive Inference Graphs

    Andreas Veit, Serge Belongie

    cs.CVcs.LGarXiv:1711.11503v32017
  24. Unsupervised Learning for Fast Probabilistic Diffeomorphic Registration

    Adrian V. Dalca, Guha Balakrishnan, John Guttag +1

    cs.CVcs.GRarXiv:1805.04605v22018
  25. MHFormer: Multi-Hypothesis Transformer for 3D Human Pose Estimation

    Wenhao Li, Hong Liu, Hao Tang +2

    cs.CVcs.AIcs.LGarXiv:2111.12707v42021
  26. 3D Hand Shape and Pose from Images in the Wild

    Adnane Boukhayma, Rodrigo de Bem, Philip H. S. Torr

    cs.CVcs.AIcs.LGarXiv:1902.03451v12019
  27. VIPER: An Expert-Curated Benchmark for Vision-Language Models in Veterinary Pathology

    Luca L. Weishaupt, Simone de Brot, Javier Asin +9

    cs.CVarXiv:2608.26382v12026
  28. CoBEVT: Cooperative Bird's Eye View Semantic Segmentation with Sparse Transformers

    Runsheng Xu, Zhengzhong Tu, Hao Xiang +3

    cs.CVarXiv:2207.02202v22022
  29. Deep ViT Features as Dense Visual Descriptors

    Shir Amir, Yossi Gandelsman, Shai Bagon +1

    cs.CVarXiv:2112.05814v32021
  30. Rethinking Counting and Localization in Crowds:A Purely Point-Based Framework

    Qingyu Song, Changan Wang, Zhengkai Jiang +6

    cs.CVarXiv:2107.12746v32021
  31. Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout

    Hao Tan, Licheng Yu, Mohit Bansal

    cs.CLcs.CVcs.LGarXiv:1904.04195v12019
  32. A Perspective on Deep Imaging

    Ge Wang

    q-bio.QMcs.CVcs.LGarXiv:1609.04375v22016
  33. Image-based 3D Object Reconstruction: State-of-the-Art and Trends in the Deep Learning Era

    Xian-Feng Han, Hamid Laga, Mohammed Bennamoun

    cs.CVcs.CGcs.GRarXiv:1906.06543v32019
  34. TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

    Junjie Wen, Yichen Zhu, Jinming Li +9

    cs.ROcs.CVarXiv:2409.12514v52024
  35. Monocular Human Pose Estimation: A Survey of Deep Learning-based Methods

    Yucheng Chen, Yingli Tian, Mingyi He

    cs.CVarXiv:2006.01423v12020
  36. C2AE: Class Conditioned Auto-Encoder for Open-set Recognition

    Poojan Oza, Vishal M Patel

    cs.CVcs.LGarXiv:1904.01198v12019
  37. CenterFusion: Center-based Radar and Camera Fusion for 3D Object Detection

    Ramin Nabati, Hairong Qi

    cs.CVarXiv:2011.04841v12020
  38. TransAttUnet: Multi-level Attention-guided U-Net with Transformer for Medical Image Segmentation

    Bingzhi Chen, Yishu Liu, Zheng Zhang +2

    eess.IVcs.CVarXiv:2107.05274v22021
  39. How AI Experiences Art: Emergent Aesthetic Structure in a Self-Supervised Multimodal Embedding Space

    Corey D. C. Heath

    cs.MMcs.CVcs.LGarXiv:2608.27121v12026
  40. Multi-stage image denoising with the wavelet transform

    Chunwei Tian, Menghua Zheng, Wangmeng Zuo +3

    eess.IVcs.CVarXiv:2209.12394v32022
  41. Revisiting the Sibling Head in Object Detector

    Guanglu Song, Yu Liu, Xiaogang Wang

    cs.CVarXiv:2003.07540v12020
  42. LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

    Yaohui Wang, Xinyuan Chen, Xin Ma +17

    cs.CVarXiv:2309.15103v22023
  43. PRECOG: PREdiction Conditioned On Goals in Visual Multi-Agent Settings

    Nicholas Rhinehart, Rowan McAllister, Kris Kitani +1

    cs.CVcs.AIcs.LGarXiv:1905.01296v32019
  44. Self supervised contrastive learning for digital histopathology

    Ozan Ciga, Tony Xu, Anne L. Martel

    eess.IVcs.CVarXiv:2011.13971v22020
  45. Mnemonics Training: Multi-Class Incremental Learning without Forgetting

    Yaoyao Liu, Yuting Su, An-An Liu +2

    cs.CVstat.MLarXiv:2002.10211v62020
  46. Unsupervised Person Re-identification via Multi-label Classification

    Dongkai Wang, Shiliang Zhang

    cs.CVarXiv:2004.09228v12020
  47. ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D Understanding

    Le Xue, Mingfei Gao, Chen Xing +6

    cs.CVarXiv:2212.05171v42022
  48. Vision-and-Dialog Navigation

    Jesse Thomason, Michael Murray, Maya Cakmak +1

    cs.CLcs.AIcs.CVarXiv:1907.04957v32019
  49. Cost Volume Pyramid Based Depth Inference for Multi-View Stereo

    Jiayu Yang, Wei Mao, Jose M. Alvarez +1

    cs.CVarXiv:1912.08329v32019
  50. Feature Weighting and Boosting for Few-Shot Segmentation

    Khoi Nguyen, Sinisa Todorovic

    cs.CVarXiv:1909.13140v12019
  51. Generative Adversarial Perturbations

    Omid Poursaeed, Isay Katsman, Bicheng Gao +1

    cs.CVcs.CRcs.LGarXiv:1712.02328v32017
  52. Online Adaptation of Convolutional Neural Networks for Video Object Segmentation

    Paul Voigtlaender, Bastian Leibe

    cs.CVarXiv:1706.09364v22017
  53. Survey on Emotional Body Gesture Recognition

    Fatemeh Noroozi, Ciprian Adrian Corneanu, Dorota Kamińska +3

    cs.CVarXiv:1801.07481v12018
  54. Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

    Hongtao Wu, Ya Jing, Chilam Cheang +6

    cs.ROcs.CVarXiv:2312.13139v22023
  55. Self-Supervised MultiModal Versatile Networks

    Jean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider +6

    cs.CVarXiv:2006.16228v22020
  56. Visual Dynamics: Probabilistic Future Frame Synthesis via Cross Convolutional Networks

    Tianfan Xue, Jiajun Wu, Katherine L. Bouman +1

    cs.CVcs.LGarXiv:1607.02586v12016
  57. Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering

    Gao Peng, Zhengkai Jiang, Haoxuan You +4

    cs.CVeess.IVarXiv:1812.05252v42018
  58. M2TR: Multi-modal Multi-scale Transformers for Deepfake Detection

    Junke Wang, Zuxuan Wu, Wenhao Ouyang +4

    cs.CVarXiv:2104.09770v32021
  59. What's Hidden in a Randomly Weighted Neural Network?

    Vivek Ramanujan, Mitchell Wortsman, Aniruddha Kembhavi +2

    cs.CVcs.LGarXiv:1911.13299v22019
  60. Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models

    Rohit Patel, Dieuwke Hupkes, Sloan Strader

    cs.CVcs.AIcs.MMarXiv:2608.26317v12026