Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

9,061 to 9,120 of 18,855

  1. VidLoc: A Deep Spatio-Temporal Model for 6-DoF Video-Clip Relocalization

    Ronald Clark, Sen Wang, Andrew Markham +2

    cs.CVarXiv:1702.06521v22017
  2. SentiCap: Generating Image Descriptions with Sentiments

    Alexander Mathews, Lexing Xie, Xuming He

    cs.CVcs.CLarXiv:1510.01431v22015
  3. Multi-Modal Hallucination Control by Visual Information Grounding

    Alessandro Favero, Luca Zancato, Matthew Trager +5

    cs.CVcs.CLcs.LGarXiv:2403.14003v12024
  4. AGIQA-3K: An Open Database for AI-Generated Image Quality Assessment

    Chunyi Li, Zicheng Zhang, Haoning Wu +5

    cs.CVcs.AIeess.IVarXiv:2306.04717v22023
  5. Randomized Smoothing of All Shapes and Sizes

    Greg Yang, Tony Duan, J. Edward Hu +3

    cs.LGcs.CVcs.NEarXiv:2002.08118v52020
  6. Scaling Up Dataset Distillation to ImageNet-1K with Constant Memory

    Justin Cui, Ruochen Wang, Si Si +1

    cs.CVcs.AIarXiv:2211.10586v42022
  7. LEDNet: Joint Low-light Enhancement and Deblurring in the Dark

    Shangchen Zhou, Chongyi Li, Chen Change Loy

    eess.IVcs.CVarXiv:2202.03373v22022
  8. TriDet: Temporal Action Detection with Relative Boundary Modeling

    Dingfeng Shi, Yujie Zhong, Qiong Cao +3

    cs.CVcs.AIcs.MMarXiv:2303.07347v22023
  9. Occluded Prohibited Items Detection: an X-ray Security Inspection Benchmark and De-occlusion Attention Module

    Yanlu Wei, Renshuai Tao, Zhangjie Wu +3

    cs.CVarXiv:2004.08656v42020
  10. Dialog-based Interactive Image Retrieval

    Xiaoxiao Guo, Hui Wu, Yu Cheng +3

    cs.CVcs.AIarXiv:1805.00145v32018
  11. Federated Learning for Computational Pathology on Gigapixel Whole Slide Images

    Ming Y. Lu, Dehan Kong, Jana Lipkova +5

    eess.IVcs.CVcs.LGarXiv:2009.10190v22020
  12. Rearrangement: A Challenge for Embodied AI

    Dhruv Batra, Angel X. Chang, Sonia Chernova +9

    cs.AIcs.CVcs.LGarXiv:2011.01975v12020
  13. Siamese Network for RGB-D Salient Object Detection and Beyond

    Keren Fu, Deng-Ping Fan, Ge-Peng Ji +3

    cs.CVarXiv:2008.12134v22020
  14. On Differentiating Parameterized Argmin and Argmax Problems with Application to Bi-level Optimization

    Stephen Gould, Basura Fernando, Anoop Cherian +3

    cs.CVmath.OCarXiv:1607.05447v22016
  15. Transformer Meets Convolution: A Bilateral Awareness Network for Semantic Segmentation of Very Fine Resolution Urban Scene Images

    Libo Wang, Rui Li, Dongzhi Wang +3

    cs.CVarXiv:2106.12413v22021
  16. Deep Vessel Segmentation By Learning Graphical Connectivity

    Seung Yeon Shin, Soochahn Lee, Il Dong Yun +1

    cs.CVarXiv:1806.02279v12018
  17. From source to target and back: symmetric bi-directional adaptive GAN

    Paolo Russo, Fabio Maria Carlucci, Tatiana Tommasi +1

    cs.CVarXiv:1705.08824v22017
  18. MedMamba: Vision Mamba for Medical Image Classification

    Yubiao Yue, Zhenzhang Li

    eess.IVcs.CVcs.LGarXiv:2403.03849v52024
  19. Tensor Canonical Correlation Analysis for Multi-view Dimension Reduction

    Yong Luo, Dacheng Tao, Yonggang Wen +2

    stat.MLcs.CVcs.LGarXiv:1502.02330v12015
  20. Chat-Edit-3D++: Interactive 3D and 4D Scene Editing via Large Language Models

    Shuangkang Fang, Yufeng Wang, Yi-Hsuan Tsai +4

    cs.CVarXiv:2608.29137v12026
  21. Feature-Guided Black-Box Safety Testing of Deep Neural Networks

    Matthew Wicker, Xiaowei Huang, Marta Kwiatkowska

    cs.CVarXiv:1710.07859v22017
  22. Cycle-Consistent Deep Generative Hashing for Cross-Modal Retrieval

    Lin Wu, Yang Wang, Ling Shao

    cs.CVarXiv:1804.11013v22018
  23. Video-based surgical skill assessment using 3D convolutional neural networks

    Isabel Funke, Sören Torge Mees, Jürgen Weitz +1

    cs.CVarXiv:1903.02306v32019
  24. Machine Learning-Based Prototyping of Graphical User Interfaces for Mobile Apps

    Kevin Moran, Carlos Bernal-Cárdenas, Michael Curcio +2

    cs.SEcs.CVcs.LGarXiv:1802.02312v22018
  25. Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation

    Jiaqi Gu, Hyoukjun Kwon, Dilin Wang +6

    cs.CVcs.AIcs.LGarXiv:2111.01236v22021
  26. Multi-View Spatial-Temporal Graph Convolutional Networks with Domain Generalization for Sleep Stage Classification

    Ziyu Jia, Youfang Lin, Jing Wang +5

    eess.SPcs.AIcs.CVarXiv:2109.01824v12021
  27. DeepInteraction: 3D Object Detection via Modality Interaction

    Zeyu Yang, Jiaqi Chen, Zhenwei Miao +3

    cs.CVarXiv:2208.11112v42022
  28. SCANimate: Weakly Supervised Learning of Skinned Clothed Avatar Networks

    Shunsuke Saito, Jinlong Yang, Qianli Ma +1

    cs.CVarXiv:2104.03313v22021
  29. CLIP2Point: Transfer CLIP to Point Cloud Classification with Image-Depth Pre-training

    Tianyu Huang, Bowen Dong, Yunhan Yang +4

    cs.CVarXiv:2210.01055v32022
  30. A Real-Time Cross-modality Correlation Filtering Method for Referring Expression Comprehension

    Yue Liao, Si Liu, Guanbin Li +4

    cs.CVarXiv:1909.07072v42019
  31. StyleHEAT: One-Shot High-Resolution Editable Talking Face Generation via Pre-trained StyleGAN

    Fei Yin, Yong Zhang, Xiaodong Cun +7

    cs.CVarXiv:2203.04036v22022
  32. Learning to cluster in order to transfer across domains and tasks

    Yen-Chang Hsu, Zhaoyang Lv, Zsolt Kira

    cs.LGcs.AIcs.CVarXiv:1711.10125v32017
  33. Segmenting Transparent Objects in the Wild

    Enze Xie, Wenjia Wang, Wenhai Wang +3

    cs.CVarXiv:2003.13948v32020
  34. Real-time Distracted Driver Posture Classification

    Yehya Abouelnaga, Hesham M. Eraqi, Mohamed N. Moustafa

    cs.CVarXiv:1706.09498v32017
  35. FLamby: Datasets and Benchmarks for Cross-Silo Federated Learning in Realistic Healthcare Settings

    Jean Ogier du Terrail, Samy-Safwan Ayed, Edwige Cyffers +21

    cs.LGcs.CVarXiv:2210.04620v32022
  36. ContextDesc: Local Descriptor Augmentation with Cross-Modality Context

    Zixin Luo, Tianwei Shen, Lei Zhou +5

    cs.CVarXiv:1904.04084v12019
  37. Combined Scaling for Zero-shot Transfer Learning

    Hieu Pham, Zihang Dai, Golnaz Ghiasi +9

    cs.LGcs.CLcs.CVarXiv:2111.10050v32021
  38. ViNT: A Foundation Model for Visual Navigation

    Dhruv Shah, Ajay Sridhar, Nitish Dashora +4

    cs.ROcs.CVcs.LGarXiv:2306.14846v22023
  39. P2T: Pyramid Pooling Transformer for Scene Understanding

    Yu-Huan Wu, Yun Liu, Xin Zhan +1

    cs.CVarXiv:2106.12011v62021
  40. $\mathbf{C}^2$Former: Calibrated and Complementary Transformer for RGB-Infrared Object Detection

    Maoxun Yuan, Xingxing Wei

    cs.CVcs.MMarXiv:2306.16175v32023
  41. TextOCR: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text

    Amanpreet Singh, Guan Pang, Mandy Toh +3

    cs.CVarXiv:2105.05486v12021
  42. CodeNeRF: Disentangled Neural Radiance Fields for Object Categories

    Wonbong Jang, Lourdes Agapito

    cs.GRcs.CVcs.LGarXiv:2109.01750v12021
  43. RVOS: End-to-End Recurrent Network for Video Object Segmentation

    Carles Ventura, Miriam Bellver, Andreu Girbau +3

    cs.CVarXiv:1903.05612v22019
  44. Predicting Head Movement in Panoramic Video: A Deep Reinforcement Learning Approach

    Yuhang Song, Mai Xu, Jianyi Wang +3

    cs.CVcs.LGarXiv:1710.10755v52017
  45. Robust Dynamic Radiance Fields

    Yu-Lun Liu, Chen Gao, Andreas Meuleman +6

    cs.CVarXiv:2301.02239v22023
  46. SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models

    Zongrui Wang, Xiangyang Zhu, Sicheng Wang +13

    cs.AIcs.CVarXiv:2608.29098v12026
  47. SatlasPretrain: A Large-Scale Dataset for Remote Sensing Image Understanding

    Favyen Bastani, Piper Wolters, Ritwik Gupta +2

    cs.CVarXiv:2211.15660v32022
  48. Early Methods for Detecting Adversarial Images

    Dan Hendrycks, Kevin Gimpel

    cs.LGcs.CRcs.CVarXiv:1608.00530v22016
  49. Learning to Track with Object Permanence

    Pavel Tokmakov, Jie Li, Wolfram Burgard +1

    cs.CVarXiv:2103.14258v22021
  50. OpenShape: Scaling Up 3D Shape Representation Towards Open-World Understanding

    Minghua Liu, Ruoxi Shi, Kaiming Kuang +6

    cs.CVarXiv:2305.10764v22023
  51. Infrared Small Target Detection with Scale and Location Sensitivity

    Qiankun Liu, Rui Liu, Bolun Zheng +2

    cs.CVarXiv:2403.19366v12024
  52. A Comparative Review of Recent Kinect-based Action Recognition Algorithms

    Lei Wang, Du Q. Huynh, Piotr Koniusz

    cs.CVarXiv:1906.09955v12019
  53. A Comparative Study of Fingerprint Image-Quality Estimation Methods

    Fernando Alonso-Fernandez, Julian Fierrez, Javier Ortega-Garcia +4

    cs.CVeess.IVarXiv:2111.07432v12021
  54. Exploiting deep residual networks for human action recognition from skeletal data

    Huy-Hieu Pham, Louahdi Khoudour, Alain Crouzil +2

    cs.CVarXiv:1803.07781v12018
  55. Contrastive Masked Autoencoders are Stronger Vision Learners

    Zhicheng Huang, Xiaojie Jin, Chengze Lu +5

    cs.CVarXiv:2207.13532v32022
  56. Constructing the L2-Graph for Robust Subspace Learning and Subspace Clustering

    Xi Peng, Zhiding Yu, Huajin Tang +1

    cs.CVcs.MMarXiv:1209.0841v72012
  57. Learning by Abstraction: The Neural State Machine

    Drew A. Hudson, Christopher D. Manning

    cs.AIcs.CLcs.CVarXiv:1907.03950v42019
  58. Large Scale Visual Food Recognition

    Weiqing Min, Zhiling Wang, Yuxin Liu +5

    cs.CVarXiv:2103.16107v32021
  59. Key-Locked Rank One Editing for Text-to-Image Personalization

    Yoad Tewel, Rinon Gal, Gal Chechik +1

    cs.CVcs.AIcs.GRarXiv:2305.01644v22023
  60. High-Resolution Breast Cancer Screening with Multi-View Deep Convolutional Neural Networks

    Krzysztof J. Geras, Stacey Wolfson, Yiqiu Shen +7

    cs.CVcs.LGstat.MLarXiv:1703.07047v32017