Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
9,061 to 9,120 of 18,855
VidLoc: A Deep Spatio-Temporal Model for 6-DoF Video-Clip Relocalization
Ronald Clark, Sen Wang, Andrew Markham +2
cs.CVarXiv:1702.06521v22017SentiCap: Generating Image Descriptions with Sentiments
Alexander Mathews, Lexing Xie, Xuming He
cs.CVcs.CLarXiv:1510.01431v22015Multi-Modal Hallucination Control by Visual Information Grounding
Alessandro Favero, Luca Zancato, Matthew Trager +5
cs.CVcs.CLcs.LGarXiv:2403.14003v12024AGIQA-3K: An Open Database for AI-Generated Image Quality Assessment
Chunyi Li, Zicheng Zhang, Haoning Wu +5
cs.CVcs.AIeess.IVarXiv:2306.04717v22023Randomized Smoothing of All Shapes and Sizes
Greg Yang, Tony Duan, J. Edward Hu +3
cs.LGcs.CVcs.NEarXiv:2002.08118v52020Scaling Up Dataset Distillation to ImageNet-1K with Constant Memory
Justin Cui, Ruochen Wang, Si Si +1
cs.CVcs.AIarXiv:2211.10586v42022LEDNet: Joint Low-light Enhancement and Deblurring in the Dark
Shangchen Zhou, Chongyi Li, Chen Change Loy
eess.IVcs.CVarXiv:2202.03373v22022TriDet: Temporal Action Detection with Relative Boundary Modeling
Dingfeng Shi, Yujie Zhong, Qiong Cao +3
cs.CVcs.AIcs.MMarXiv:2303.07347v22023Occluded Prohibited Items Detection: an X-ray Security Inspection Benchmark and De-occlusion Attention Module
Yanlu Wei, Renshuai Tao, Zhangjie Wu +3
cs.CVarXiv:2004.08656v42020Dialog-based Interactive Image Retrieval
Xiaoxiao Guo, Hui Wu, Yu Cheng +3
cs.CVcs.AIarXiv:1805.00145v32018Federated Learning for Computational Pathology on Gigapixel Whole Slide Images
Ming Y. Lu, Dehan Kong, Jana Lipkova +5
eess.IVcs.CVcs.LGarXiv:2009.10190v22020Rearrangement: A Challenge for Embodied AI
Dhruv Batra, Angel X. Chang, Sonia Chernova +9
cs.AIcs.CVcs.LGarXiv:2011.01975v12020Siamese Network for RGB-D Salient Object Detection and Beyond
Keren Fu, Deng-Ping Fan, Ge-Peng Ji +3
cs.CVarXiv:2008.12134v22020On Differentiating Parameterized Argmin and Argmax Problems with Application to Bi-level Optimization
Stephen Gould, Basura Fernando, Anoop Cherian +3
cs.CVmath.OCarXiv:1607.05447v22016Transformer Meets Convolution: A Bilateral Awareness Network for Semantic Segmentation of Very Fine Resolution Urban Scene Images
Libo Wang, Rui Li, Dongzhi Wang +3
cs.CVarXiv:2106.12413v22021Deep Vessel Segmentation By Learning Graphical Connectivity
Seung Yeon Shin, Soochahn Lee, Il Dong Yun +1
cs.CVarXiv:1806.02279v12018From source to target and back: symmetric bi-directional adaptive GAN
Paolo Russo, Fabio Maria Carlucci, Tatiana Tommasi +1
cs.CVarXiv:1705.08824v22017MedMamba: Vision Mamba for Medical Image Classification
Yubiao Yue, Zhenzhang Li
eess.IVcs.CVcs.LGarXiv:2403.03849v52024Tensor Canonical Correlation Analysis for Multi-view Dimension Reduction
Yong Luo, Dacheng Tao, Yonggang Wen +2
stat.MLcs.CVcs.LGarXiv:1502.02330v12015Chat-Edit-3D++: Interactive 3D and 4D Scene Editing via Large Language Models
Shuangkang Fang, Yufeng Wang, Yi-Hsuan Tsai +4
cs.CVarXiv:2608.29137v12026Feature-Guided Black-Box Safety Testing of Deep Neural Networks
Matthew Wicker, Xiaowei Huang, Marta Kwiatkowska
cs.CVarXiv:1710.07859v22017Cycle-Consistent Deep Generative Hashing for Cross-Modal Retrieval
Lin Wu, Yang Wang, Ling Shao
cs.CVarXiv:1804.11013v22018Video-based surgical skill assessment using 3D convolutional neural networks
Isabel Funke, Sören Torge Mees, Jürgen Weitz +1
cs.CVarXiv:1903.02306v32019Machine Learning-Based Prototyping of Graphical User Interfaces for Mobile Apps
Kevin Moran, Carlos Bernal-Cárdenas, Michael Curcio +2
cs.SEcs.CVcs.LGarXiv:1802.02312v22018Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation
Jiaqi Gu, Hyoukjun Kwon, Dilin Wang +6
cs.CVcs.AIcs.LGarXiv:2111.01236v22021Multi-View Spatial-Temporal Graph Convolutional Networks with Domain Generalization for Sleep Stage Classification
Ziyu Jia, Youfang Lin, Jing Wang +5
eess.SPcs.AIcs.CVarXiv:2109.01824v12021DeepInteraction: 3D Object Detection via Modality Interaction
Zeyu Yang, Jiaqi Chen, Zhenwei Miao +3
cs.CVarXiv:2208.11112v42022SCANimate: Weakly Supervised Learning of Skinned Clothed Avatar Networks
Shunsuke Saito, Jinlong Yang, Qianli Ma +1
cs.CVarXiv:2104.03313v22021CLIP2Point: Transfer CLIP to Point Cloud Classification with Image-Depth Pre-training
Tianyu Huang, Bowen Dong, Yunhan Yang +4
cs.CVarXiv:2210.01055v32022A Real-Time Cross-modality Correlation Filtering Method for Referring Expression Comprehension
Yue Liao, Si Liu, Guanbin Li +4
cs.CVarXiv:1909.07072v42019StyleHEAT: One-Shot High-Resolution Editable Talking Face Generation via Pre-trained StyleGAN
Fei Yin, Yong Zhang, Xiaodong Cun +7
cs.CVarXiv:2203.04036v22022Learning to cluster in order to transfer across domains and tasks
Yen-Chang Hsu, Zhaoyang Lv, Zsolt Kira
cs.LGcs.AIcs.CVarXiv:1711.10125v32017Segmenting Transparent Objects in the Wild
Enze Xie, Wenjia Wang, Wenhai Wang +3
cs.CVarXiv:2003.13948v32020Real-time Distracted Driver Posture Classification
Yehya Abouelnaga, Hesham M. Eraqi, Mohamed N. Moustafa
cs.CVarXiv:1706.09498v32017FLamby: Datasets and Benchmarks for Cross-Silo Federated Learning in Realistic Healthcare Settings
Jean Ogier du Terrail, Samy-Safwan Ayed, Edwige Cyffers +21
cs.LGcs.CVarXiv:2210.04620v32022ContextDesc: Local Descriptor Augmentation with Cross-Modality Context
Zixin Luo, Tianwei Shen, Lei Zhou +5
cs.CVarXiv:1904.04084v12019Combined Scaling for Zero-shot Transfer Learning
Hieu Pham, Zihang Dai, Golnaz Ghiasi +9
cs.LGcs.CLcs.CVarXiv:2111.10050v32021ViNT: A Foundation Model for Visual Navigation
Dhruv Shah, Ajay Sridhar, Nitish Dashora +4
cs.ROcs.CVcs.LGarXiv:2306.14846v22023P2T: Pyramid Pooling Transformer for Scene Understanding
Yu-Huan Wu, Yun Liu, Xin Zhan +1
cs.CVarXiv:2106.12011v62021$\mathbf{C}^2$Former: Calibrated and Complementary Transformer for RGB-Infrared Object Detection
Maoxun Yuan, Xingxing Wei
cs.CVcs.MMarXiv:2306.16175v32023TextOCR: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text
Amanpreet Singh, Guan Pang, Mandy Toh +3
cs.CVarXiv:2105.05486v12021CodeNeRF: Disentangled Neural Radiance Fields for Object Categories
Wonbong Jang, Lourdes Agapito
cs.GRcs.CVcs.LGarXiv:2109.01750v12021RVOS: End-to-End Recurrent Network for Video Object Segmentation
Carles Ventura, Miriam Bellver, Andreu Girbau +3
cs.CVarXiv:1903.05612v22019Predicting Head Movement in Panoramic Video: A Deep Reinforcement Learning Approach
Yuhang Song, Mai Xu, Jianyi Wang +3
cs.CVcs.LGarXiv:1710.10755v52017Robust Dynamic Radiance Fields
Yu-Lun Liu, Chen Gao, Andreas Meuleman +6
cs.CVarXiv:2301.02239v22023SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models
Zongrui Wang, Xiangyang Zhu, Sicheng Wang +13
cs.AIcs.CVarXiv:2608.29098v12026SatlasPretrain: A Large-Scale Dataset for Remote Sensing Image Understanding
Favyen Bastani, Piper Wolters, Ritwik Gupta +2
cs.CVarXiv:2211.15660v32022Early Methods for Detecting Adversarial Images
Dan Hendrycks, Kevin Gimpel
cs.LGcs.CRcs.CVarXiv:1608.00530v22016Learning to Track with Object Permanence
Pavel Tokmakov, Jie Li, Wolfram Burgard +1
cs.CVarXiv:2103.14258v22021OpenShape: Scaling Up 3D Shape Representation Towards Open-World Understanding
Minghua Liu, Ruoxi Shi, Kaiming Kuang +6
cs.CVarXiv:2305.10764v22023Infrared Small Target Detection with Scale and Location Sensitivity
Qiankun Liu, Rui Liu, Bolun Zheng +2
cs.CVarXiv:2403.19366v12024A Comparative Review of Recent Kinect-based Action Recognition Algorithms
Lei Wang, Du Q. Huynh, Piotr Koniusz
cs.CVarXiv:1906.09955v12019A Comparative Study of Fingerprint Image-Quality Estimation Methods
Fernando Alonso-Fernandez, Julian Fierrez, Javier Ortega-Garcia +4
cs.CVeess.IVarXiv:2111.07432v12021Exploiting deep residual networks for human action recognition from skeletal data
Huy-Hieu Pham, Louahdi Khoudour, Alain Crouzil +2
cs.CVarXiv:1803.07781v12018Contrastive Masked Autoencoders are Stronger Vision Learners
Zhicheng Huang, Xiaojie Jin, Chengze Lu +5
cs.CVarXiv:2207.13532v32022Constructing the L2-Graph for Robust Subspace Learning and Subspace Clustering
Xi Peng, Zhiding Yu, Huajin Tang +1
cs.CVcs.MMarXiv:1209.0841v72012Learning by Abstraction: The Neural State Machine
Drew A. Hudson, Christopher D. Manning
cs.AIcs.CLcs.CVarXiv:1907.03950v42019Large Scale Visual Food Recognition
Weiqing Min, Zhiling Wang, Yuxin Liu +5
cs.CVarXiv:2103.16107v32021Key-Locked Rank One Editing for Text-to-Image Personalization
Yoad Tewel, Rinon Gal, Gal Chechik +1
cs.CVcs.AIcs.GRarXiv:2305.01644v22023High-Resolution Breast Cancer Screening with Multi-View Deep Convolutional Neural Networks
Krzysztof J. Geras, Stacey Wolfson, Yiqiu Shen +7
cs.CVcs.LGstat.MLarXiv:1703.07047v32017