Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
12,601 to 12,660 of 18,866
Beyond Atomic Layouts: Compositional Design Understanding with Vision-Language Models
Yiyang Huang, Zhaowen Wang, Simon Jenni +4
cs.CVarXiv:2608.26716v12026Attention-based Dropout Layer for Weakly Supervised Object Localization
Junsuk Choe, Hyunjung Shim
cs.CVarXiv:1908.10028v12019Deep Attributes Driven Multi-Camera Person Re-identification
Chi Su, Shiliang Zhang, Junliang Xing +2
cs.CVarXiv:1605.03259v22016Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning
Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo +5
cs.CVcs.AIcs.CLarXiv:2302.14115v22023CrossTransformers: spatially-aware few-shot transfer
Carl Doersch, Ankush Gupta, Andrew Zisserman
cs.CVarXiv:2007.11498v52020Multi-Person Human Motion Forecasting in Complex Scenes
Serdar Ozsoy, Lars Doorenbos, Juergen Gall
cs.CVcs.AIarXiv:2608.27039v12026Fixup Initialization: Residual Learning Without Normalization
Hongyi Zhang, Yann N. Dauphin, Tengyu Ma
cs.LGcs.CVstat.MLarXiv:1901.09321v22019Playing with Duality: An Overview of Recent Primal-Dual Approaches for Solving Large-Scale Optimization Problems
Nikos Komodakis, Jean-Christophe Pesquet
math.NAcs.CVcs.LGarXiv:1406.5429v22014Function4D: Real-time Human Volumetric Capture from Very Sparse Consumer RGBD Sensors
Tao Yu, Zerong Zheng, Kaiwen Guo +3
cs.CVarXiv:2105.01859v22021Iterative Geometry Encoding Volume for Stereo Matching
Gangwei Xu, Xianqi Wang, Xiaohuan Ding +1
cs.CVarXiv:2303.06615v22023The Platonic Representation Hypothesis
Minyoung Huh, Brian Cheung, Tongzhou Wang +1
cs.LGcs.AIcs.CVarXiv:2405.07987v52024Expression, Affect, Action Unit Recognition: Aff-Wild2, Multi-Task Learning and ArcFace
Dimitrios Kollias, Stefanos Zafeiriou
cs.CVcs.HCcs.LGarXiv:1910.04855v12019Is Pseudo-Lidar needed for Monocular 3D Object detection?
Dennis Park, Rares Ambrus, Vitor Guizilini +2
cs.CVarXiv:2108.06417v12021Weighted Schatten $p$-Norm Minimization for Image Denoising and Background Subtraction
Yuan Xie, Shuhang Gu, Yan Liu +3
cs.CVarXiv:1512.01003v12015Improving Semantic Segmentation via Video Propagation and Label Relaxation
Yi Zhu, Karan Sapra, Fitsum A. Reda +4
cs.CVcs.AIcs.MMarXiv:1812.01593v32018PointCleanNet: Learning to Denoise and Remove Outliers from Dense Point Clouds
Marie-Julie Rakotosaona, Vittorio La Barbera, Paul Guerrero +2
cs.GRcs.CVarXiv:1901.01060v32019Human Motion Diffusion as a Generative Prior
Yonatan Shafir, Guy Tevet, Roy Kapon +1
cs.CVcs.GRarXiv:2303.01418v32023Skeleton-based Action Recognition via Spatial and Temporal Transformer Networks
Chiara Plizzari, Marco Cannici, Matteo Matteucci
cs.CVarXiv:2008.07404v42020FastComposer: Tuning-Free Multi-Subject Image Generation with Localized Attention
Guangxuan Xiao, Tianwei Yin, William T. Freeman +2
cs.CVarXiv:2305.10431v22023Dataset Condensation with Differentiable Siamese Augmentation
Bo Zhao, Hakan Bilen
cs.LGcs.CVarXiv:2102.08259v22021Image Data Augmentation for Deep Learning: A Survey
Suorong Yang, Weikang Xiao, Mengchen Zhang +3
cs.CVarXiv:2204.08610v22022Models Genesis: Generic Autodidactic Models for 3D Medical Image Analysis
Zongwei Zhou, Vatsal Sodha, Md Mahfuzur Rahman Siddiquee +4
eess.IVcs.CVarXiv:1908.06912v12019Convolutional Networks with Adaptive Inference Graphs
Andreas Veit, Serge Belongie
cs.CVcs.LGarXiv:1711.11503v32017Unsupervised Learning for Fast Probabilistic Diffeomorphic Registration
Adrian V. Dalca, Guha Balakrishnan, John Guttag +1
cs.CVcs.GRarXiv:1805.04605v22018MHFormer: Multi-Hypothesis Transformer for 3D Human Pose Estimation
Wenhao Li, Hong Liu, Hao Tang +2
cs.CVcs.AIcs.LGarXiv:2111.12707v420213D Hand Shape and Pose from Images in the Wild
Adnane Boukhayma, Rodrigo de Bem, Philip H. S. Torr
cs.CVcs.AIcs.LGarXiv:1902.03451v12019VIPER: An Expert-Curated Benchmark for Vision-Language Models in Veterinary Pathology
Luca L. Weishaupt, Simone de Brot, Javier Asin +9
cs.CVarXiv:2608.26382v12026CoBEVT: Cooperative Bird's Eye View Semantic Segmentation with Sparse Transformers
Runsheng Xu, Zhengzhong Tu, Hao Xiang +3
cs.CVarXiv:2207.02202v22022Deep ViT Features as Dense Visual Descriptors
Shir Amir, Yossi Gandelsman, Shai Bagon +1
cs.CVarXiv:2112.05814v32021Rethinking Counting and Localization in Crowds:A Purely Point-Based Framework
Qingyu Song, Changan Wang, Zhengkai Jiang +6
cs.CVarXiv:2107.12746v32021Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout
Hao Tan, Licheng Yu, Mohit Bansal
cs.CLcs.CVcs.LGarXiv:1904.04195v12019A Perspective on Deep Imaging
Ge Wang
q-bio.QMcs.CVcs.LGarXiv:1609.04375v22016Image-based 3D Object Reconstruction: State-of-the-Art and Trends in the Deep Learning Era
Xian-Feng Han, Hamid Laga, Mohammed Bennamoun
cs.CVcs.CGcs.GRarXiv:1906.06543v32019TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
Junjie Wen, Yichen Zhu, Jinming Li +9
cs.ROcs.CVarXiv:2409.12514v52024Monocular Human Pose Estimation: A Survey of Deep Learning-based Methods
Yucheng Chen, Yingli Tian, Mingyi He
cs.CVarXiv:2006.01423v12020C2AE: Class Conditioned Auto-Encoder for Open-set Recognition
Poojan Oza, Vishal M Patel
cs.CVcs.LGarXiv:1904.01198v12019CenterFusion: Center-based Radar and Camera Fusion for 3D Object Detection
Ramin Nabati, Hairong Qi
cs.CVarXiv:2011.04841v12020TransAttUnet: Multi-level Attention-guided U-Net with Transformer for Medical Image Segmentation
Bingzhi Chen, Yishu Liu, Zheng Zhang +2
eess.IVcs.CVarXiv:2107.05274v22021How AI Experiences Art: Emergent Aesthetic Structure in a Self-Supervised Multimodal Embedding Space
Corey D. C. Heath
cs.MMcs.CVcs.LGarXiv:2608.27121v12026Multi-stage image denoising with the wavelet transform
Chunwei Tian, Menghua Zheng, Wangmeng Zuo +3
eess.IVcs.CVarXiv:2209.12394v32022Revisiting the Sibling Head in Object Detector
Guanglu Song, Yu Liu, Xiaogang Wang
cs.CVarXiv:2003.07540v12020LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models
Yaohui Wang, Xinyuan Chen, Xin Ma +17
cs.CVarXiv:2309.15103v22023PRECOG: PREdiction Conditioned On Goals in Visual Multi-Agent Settings
Nicholas Rhinehart, Rowan McAllister, Kris Kitani +1
cs.CVcs.AIcs.LGarXiv:1905.01296v32019Self supervised contrastive learning for digital histopathology
Ozan Ciga, Tony Xu, Anne L. Martel
eess.IVcs.CVarXiv:2011.13971v22020Mnemonics Training: Multi-Class Incremental Learning without Forgetting
Yaoyao Liu, Yuting Su, An-An Liu +2
cs.CVstat.MLarXiv:2002.10211v62020Unsupervised Person Re-identification via Multi-label Classification
Dongkai Wang, Shiliang Zhang
cs.CVarXiv:2004.09228v12020ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D Understanding
Le Xue, Mingfei Gao, Chen Xing +6
cs.CVarXiv:2212.05171v42022Vision-and-Dialog Navigation
Jesse Thomason, Michael Murray, Maya Cakmak +1
cs.CLcs.AIcs.CVarXiv:1907.04957v32019Cost Volume Pyramid Based Depth Inference for Multi-View Stereo
Jiayu Yang, Wei Mao, Jose M. Alvarez +1
cs.CVarXiv:1912.08329v32019Feature Weighting and Boosting for Few-Shot Segmentation
Khoi Nguyen, Sinisa Todorovic
cs.CVarXiv:1909.13140v12019Generative Adversarial Perturbations
Omid Poursaeed, Isay Katsman, Bicheng Gao +1
cs.CVcs.CRcs.LGarXiv:1712.02328v32017Online Adaptation of Convolutional Neural Networks for Video Object Segmentation
Paul Voigtlaender, Bastian Leibe
cs.CVarXiv:1706.09364v22017Survey on Emotional Body Gesture Recognition
Fatemeh Noroozi, Ciprian Adrian Corneanu, Dorota Kamińska +3
cs.CVarXiv:1801.07481v12018Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation
Hongtao Wu, Ya Jing, Chilam Cheang +6
cs.ROcs.CVarXiv:2312.13139v22023Self-Supervised MultiModal Versatile Networks
Jean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider +6
cs.CVarXiv:2006.16228v22020Visual Dynamics: Probabilistic Future Frame Synthesis via Cross Convolutional Networks
Tianfan Xue, Jiajun Wu, Katherine L. Bouman +1
cs.CVcs.LGarXiv:1607.02586v12016Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering
Gao Peng, Zhengkai Jiang, Haoxuan You +4
cs.CVeess.IVarXiv:1812.05252v42018M2TR: Multi-modal Multi-scale Transformers for Deepfake Detection
Junke Wang, Zuxuan Wu, Wenhao Ouyang +4
cs.CVarXiv:2104.09770v32021What's Hidden in a Randomly Weighted Neural Network?
Vivek Ramanujan, Mitchell Wortsman, Aniruddha Kembhavi +2
cs.CVcs.LGarXiv:1911.13299v22019Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models
Rohit Patel, Dieuwke Hupkes, Sloan Strader
cs.CVcs.AIcs.MMarXiv:2608.26317v12026