Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
10,321 to 10,380 of 18,866
Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review
Iryna Hartsock, Ghulam Rasool
cs.CVcs.LGarXiv:2403.02469v22024Foundations and Trends in Multimodal Machine Learning: Principles, Challenges, and Open Questions
Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency
cs.LGcs.AIcs.CLarXiv:2209.03430v22022AutoLoc: Weakly-supervised Temporal Action Localization
Zheng Shou, Hang Gao, Lei Zhang +2
cs.CVarXiv:1807.08333v22018ESLAM: Efficient Dense SLAM System Based on Hybrid Representation of Signed Distance Fields
Mohammad Mahdi Johari, Camilla Carta, François Fleuret
cs.CVarXiv:2211.11704v22022Panoptic Segmentation of Satellite Image Time Series with Convolutional Temporal Attention Networks
Vivien Sainte Fare Garnot, Loic Landrieu
cs.CVarXiv:2107.07933v42021PST900: RGB-Thermal Calibration, Dataset and Segmentation Network
Shreyas S. Shivakumar, Neil Rodrigues, Alex Zhou +3
cs.CVcs.ROeess.IVarXiv:1909.10980v12019Associatively Segmenting Instances and Semantics in Point Clouds
Xinlong Wang, Shu Liu, Xiaoyong Shen +2
cs.CVarXiv:1902.09852v22019Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation Learning
Fuying Wang, Yuyin Zhou, Shujun Wang +2
cs.CVcs.AIcs.CLarXiv:2210.06044v12022Automated polyp detection in colon capsule endoscopy
Alexander V. Mamonov, Isabel N. Figueiredo, Pedro N. Figueiredo +1
cs.CVarXiv:1305.1912v42013Boundary-Aware Feature Propagation for Scene Segmentation
Henghui Ding, Xudong Jiang, Ai Qun Liu +2
cs.CVarXiv:1909.00179v12019Conditional Image Generation with Score-Based Diffusion Models
Georgios Batzolis, Jan Stanczuk, Carola-Bibiane Schönlieb +1
cs.LGcs.CVstat.MLarXiv:2111.13606v12021Reachability Analysis of Deep Neural Networks with Provable Guarantees
Wenjie Ruan, Xiaowei Huang, Marta Kwiatkowska
cs.LGcs.CVstat.MLarXiv:1805.02242v12018Unsupervised Semantic Segmentation by Contrasting Object Mask Proposals
Wouter Van Gansbeke, Simon Vandenhende, Stamatios Georgoulis +1
cs.CVcs.LGarXiv:2102.06191v32021TEACh: Task-driven Embodied Agents that Chat
Aishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava +6
cs.CVcs.AIcs.CLarXiv:2110.00534v32021SuperPCA: A Superpixelwise PCA Approach for Unsupervised Feature Extraction of Hyperspectral Imagery
Junjun Jiang, Jiayi Ma, Chen Chen +3
cs.CVarXiv:1806.09807v22018Optimizing Prompts for Text-to-Image Generation
Yaru Hao, Zewen Chi, Li Dong +1
cs.CLcs.CVarXiv:2212.09611v22022AnyLoc: Towards Universal Visual Place Recognition
Nikhil Keetha, Avneesh Mishra, Jay Karhade +4
cs.CVcs.AIcs.ROarXiv:2308.00688v22023Exploring Smoothness and Class-Separation for Semi-supervised Medical Image Segmentation
Yicheng Wu, Zhonghua Wu, Qianyi Wu +2
eess.IVcs.CVarXiv:2203.01324v32022Deblur-NeRF: Neural Radiance Fields from Blurry Images
Li Ma, Xiaoyu Li, Jing Liao +4
cs.CVcs.GRarXiv:2111.14292v22021Text-based Editing of Talking-head Video
Ohad Fried, Ayush Tewari, Michael Zollhöfer +7
cs.CVcs.GRcs.LGarXiv:1906.01524v12019DetNAS: Backbone Search for Object Detection
Yukang Chen, Tong Yang, Xiangyu Zhang +3
cs.CVarXiv:1903.10979v42019AD-Cluster: Augmented Discriminative Clustering for Domain Adaptive Person Re-identification
Yunpeng Zhai, Shijian Lu, Qixiang Ye +4
cs.CVarXiv:2004.08787v22020Diagnose like a Radiologist: Attention Guided Convolutional Neural Network for Thorax Disease Classification
Qingji Guan, Yaping Huang, Zhun Zhong +3
cs.CVarXiv:1801.09927v12018AutoGAN: Neural Architecture Search for Generative Adversarial Networks
Xinyu Gong, Shiyu Chang, Yifan Jiang +1
cs.CVcs.LGeess.IVarXiv:1908.03835v12019AttentionGAN: Unpaired Image-to-Image Translation using Attention-Guided Generative Adversarial Networks
Hao Tang, Hong Liu, Dan Xu +2
cs.CVcs.LGeess.IVarXiv:1911.11897v52019Unsupervised Object Discovery and Localization in the Wild: Part-based Matching with Bottom-up Region Proposals
Minsu Cho, Suha Kwak, Cordelia Schmid +1
cs.CVarXiv:1501.06170v32015Gradually Vanishing Bridge for Adversarial Domain Adaptation
Shuhao Cui, Shuhui Wang, Junbao Zhuo +3
cs.CVarXiv:2003.13183v12020Predicting Ground-Level Scene Layout from Aerial Imagery
Menghua Zhai, Zachary Bessinger, Scott Workman +1
cs.CVarXiv:1612.02709v12016Recurrent Vision Transformers for Object Detection with Event Cameras
Mathias Gehrig, Davide Scaramuzza
cs.CVarXiv:2212.05598v32022End-to-End Robotic Reinforcement Learning without Reward Engineering
Avi Singh, Larry Yang, Kristian Hartikainen +2
cs.LGcs.CVcs.ROarXiv:1904.07854v22019Stereo Radiance Fields (SRF): Learning View Synthesis for Sparse Views of Novel Scenes
Julian Chibane, Aayush Bansal, Verica Lazova +1
cs.CVcs.LGarXiv:2104.06935v12021ThunderNet: Towards Real-time Generic Object Detection
Zheng Qin, Zeming Li, Zhaoning Zhang +4
cs.CVarXiv:1903.11752v32019Sparse and Dense Data with CNNs: Depth Completion and Semantic Segmentation
Maximilian Jaritz, Raoul de Charette, Emilie Wirbel +2
cs.CVarXiv:1808.00769v22018Vision-Language Navigation with Self-Supervised Auxiliary Reasoning Tasks
Fengda Zhu, Yi Zhu, Xiaojun Chang +1
cs.CVarXiv:1911.07883v42019Feature Space Augmentation for Long-Tailed Data
Peng Chu, Xiao Bian, Shaopeng Liu +1
cs.CVarXiv:2008.03673v12020NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation
Jiazhao Zhang, Kunyu Wang, Rongtao Xu +6
cs.CVcs.ROarXiv:2402.15852v72024GRAM: Generative Radiance Manifolds for 3D-Aware Image Generation
Yu Deng, Jiaolong Yang, Jianfeng Xiang +1
cs.CVarXiv:2112.08867v32021Frequency-aware Feature Fusion for Dense Image Prediction
Linwei Chen, Ying Fu, Lin Gu +3
cs.CVcs.AIarXiv:2408.12879v12024Liquid Warping GAN: A Unified Framework for Human Motion Imitation, Appearance Transfer and Novel View Synthesis
Wen Liu, Zhixin Piao, Jie Min +3
cs.CVcs.LGeess.IVarXiv:1909.12224v32019AdaCoF: Adaptive Collaboration of Flows for Video Frame Interpolation
Hyeongmin Lee, Taeoh Kim, Tae-young Chung +3
cs.CVarXiv:1907.10244v32019Diving Deeper into Underwater Image Enhancement: A Survey
Saeed Anwar, Chongyi Li
cs.CVcs.LGeess.IVarXiv:1907.07863v12019Localizing Objects with Self-Supervised Transformers and no Labels
Oriane Siméoni, Gilles Puy, Huy V. Vo +6
cs.CVarXiv:2109.14279v12021Subcategory-aware Convolutional Neural Networks for Object Proposals and Detection
Yu Xiang, Wongun Choi, Yuanqing Lin +1
cs.CVarXiv:1604.04693v32016LayoutGAN: Generating Graphic Layouts with Wireframe Discriminators
Jianan Li, Jimei Yang, Aaron Hertzmann +2
cs.CVarXiv:1901.06767v12019Self-Calibrating Neural Radiance Fields
Yoonwoo Jeong, Seokjun Ahn, Christopher Choy +3
cs.CVarXiv:2108.13826v22021Shallow-UWnet : Compressed Model for Underwater Image Enhancement
Ankita Naik, Apurva Swarnakar, Kartik Mittal
cs.CVeess.IVarXiv:2101.02073v12021A Light CNN for detecting COVID-19 from CT scans of the chest
Matteo Polsinelli, Luigi Cinque, Giuseppe Placidi
eess.IVcs.CVcs.LGarXiv:2004.12837v12020Semi-supervised Medical Image Classification with Relation-driven Self-ensembling Model
Quande Liu, Lequan Yu, Luyang Luo +2
cs.CVarXiv:2005.07377v12020Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters
Jiazuo Yu, Yunzhi Zhuge, Lu Zhang +4
cs.CVarXiv:2403.11549v22024Neural-Guided RANSAC: Learning Where to Sample Model Hypotheses
Eric Brachmann, Carsten Rother
cs.CVarXiv:1905.04132v22019Combining Language and Vision with a Multimodal Skip-gram Model
Angeliki Lazaridou, Nghia The Pham, Marco Baroni
cs.CLcs.CVcs.LGarXiv:1501.02598v32015GaussianPro: 3D Gaussian Splatting with Progressive Propagation
Kai Cheng, Xiaoxiao Long, Kaizhi Yang +5
cs.CVarXiv:2402.14650v12024Channel Pruning via Automatic Structure Search
Mingbao Lin, Rongrong Ji, Yuxin Zhang +3
cs.CVarXiv:2001.08565v32020Pangu-Weather: A 3D High-Resolution Model for Fast and Accurate Global Weather Forecast
Kaifeng Bi, Lingxi Xie, Hengheng Zhang +3
physics.ao-phcs.AIcs.CVarXiv:2211.02556v12022Capsules for Object Segmentation
Rodney LaLonde, Ulas Bagci
stat.MLcs.AIcs.CVarXiv:1804.04241v12018Fractional Calculus In Image Processing: A Review
Qi Yang, Dali Chen, Tiebiao Zhao +1
cs.CVarXiv:1608.03240v12016MSRF-Net: A Multi-Scale Residual Fusion Network for Biomedical Image Segmentation
Abhishek Srivastava, Debesh Jha, Sukalpa Chanda +6
eess.IVcs.CVarXiv:2105.07451v22021NerfingMVS: Guided Optimization of Neural Radiance Fields for Indoor Multi-view Stereo
Yi Wei, Shaohui Liu, Yongming Rao +3
cs.CVarXiv:2109.01129v32021Representative Forgery Mining for Fake Face Detection
Chengrui Wang, Weihong Deng
cs.CVarXiv:2104.06609v12021Single-Path NAS: Designing Hardware-Efficient ConvNets in less than 4 Hours
Dimitrios Stamoulis, Ruizhou Ding, Di Wang +4
cs.LGcs.CVstat.MLarXiv:1904.02877v12019