Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

10,321 to 10,380 of 18,866

  1. Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review

    Iryna Hartsock, Ghulam Rasool

    cs.CVcs.LGarXiv:2403.02469v22024
  2. Foundations and Trends in Multimodal Machine Learning: Principles, Challenges, and Open Questions

    Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency

    cs.LGcs.AIcs.CLarXiv:2209.03430v22022
  3. AutoLoc: Weakly-supervised Temporal Action Localization

    Zheng Shou, Hang Gao, Lei Zhang +2

    cs.CVarXiv:1807.08333v22018
  4. ESLAM: Efficient Dense SLAM System Based on Hybrid Representation of Signed Distance Fields

    Mohammad Mahdi Johari, Camilla Carta, François Fleuret

    cs.CVarXiv:2211.11704v22022
  5. Panoptic Segmentation of Satellite Image Time Series with Convolutional Temporal Attention Networks

    Vivien Sainte Fare Garnot, Loic Landrieu

    cs.CVarXiv:2107.07933v42021
  6. PST900: RGB-Thermal Calibration, Dataset and Segmentation Network

    Shreyas S. Shivakumar, Neil Rodrigues, Alex Zhou +3

    cs.CVcs.ROeess.IVarXiv:1909.10980v12019
  7. Associatively Segmenting Instances and Semantics in Point Clouds

    Xinlong Wang, Shu Liu, Xiaoyong Shen +2

    cs.CVarXiv:1902.09852v22019
  8. Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation Learning

    Fuying Wang, Yuyin Zhou, Shujun Wang +2

    cs.CVcs.AIcs.CLarXiv:2210.06044v12022
  9. Automated polyp detection in colon capsule endoscopy

    Alexander V. Mamonov, Isabel N. Figueiredo, Pedro N. Figueiredo +1

    cs.CVarXiv:1305.1912v42013
  10. Boundary-Aware Feature Propagation for Scene Segmentation

    Henghui Ding, Xudong Jiang, Ai Qun Liu +2

    cs.CVarXiv:1909.00179v12019
  11. Conditional Image Generation with Score-Based Diffusion Models

    Georgios Batzolis, Jan Stanczuk, Carola-Bibiane Schönlieb +1

    cs.LGcs.CVstat.MLarXiv:2111.13606v12021
  12. Reachability Analysis of Deep Neural Networks with Provable Guarantees

    Wenjie Ruan, Xiaowei Huang, Marta Kwiatkowska

    cs.LGcs.CVstat.MLarXiv:1805.02242v12018
  13. Unsupervised Semantic Segmentation by Contrasting Object Mask Proposals

    Wouter Van Gansbeke, Simon Vandenhende, Stamatios Georgoulis +1

    cs.CVcs.LGarXiv:2102.06191v32021
  14. TEACh: Task-driven Embodied Agents that Chat

    Aishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava +6

    cs.CVcs.AIcs.CLarXiv:2110.00534v32021
  15. SuperPCA: A Superpixelwise PCA Approach for Unsupervised Feature Extraction of Hyperspectral Imagery

    Junjun Jiang, Jiayi Ma, Chen Chen +3

    cs.CVarXiv:1806.09807v22018
  16. Optimizing Prompts for Text-to-Image Generation

    Yaru Hao, Zewen Chi, Li Dong +1

    cs.CLcs.CVarXiv:2212.09611v22022
  17. AnyLoc: Towards Universal Visual Place Recognition

    Nikhil Keetha, Avneesh Mishra, Jay Karhade +4

    cs.CVcs.AIcs.ROarXiv:2308.00688v22023
  18. Exploring Smoothness and Class-Separation for Semi-supervised Medical Image Segmentation

    Yicheng Wu, Zhonghua Wu, Qianyi Wu +2

    eess.IVcs.CVarXiv:2203.01324v32022
  19. Deblur-NeRF: Neural Radiance Fields from Blurry Images

    Li Ma, Xiaoyu Li, Jing Liao +4

    cs.CVcs.GRarXiv:2111.14292v22021
  20. Text-based Editing of Talking-head Video

    Ohad Fried, Ayush Tewari, Michael Zollhöfer +7

    cs.CVcs.GRcs.LGarXiv:1906.01524v12019
  21. DetNAS: Backbone Search for Object Detection

    Yukang Chen, Tong Yang, Xiangyu Zhang +3

    cs.CVarXiv:1903.10979v42019
  22. AD-Cluster: Augmented Discriminative Clustering for Domain Adaptive Person Re-identification

    Yunpeng Zhai, Shijian Lu, Qixiang Ye +4

    cs.CVarXiv:2004.08787v22020
  23. Diagnose like a Radiologist: Attention Guided Convolutional Neural Network for Thorax Disease Classification

    Qingji Guan, Yaping Huang, Zhun Zhong +3

    cs.CVarXiv:1801.09927v12018
  24. AutoGAN: Neural Architecture Search for Generative Adversarial Networks

    Xinyu Gong, Shiyu Chang, Yifan Jiang +1

    cs.CVcs.LGeess.IVarXiv:1908.03835v12019
  25. AttentionGAN: Unpaired Image-to-Image Translation using Attention-Guided Generative Adversarial Networks

    Hao Tang, Hong Liu, Dan Xu +2

    cs.CVcs.LGeess.IVarXiv:1911.11897v52019
  26. Unsupervised Object Discovery and Localization in the Wild: Part-based Matching with Bottom-up Region Proposals

    Minsu Cho, Suha Kwak, Cordelia Schmid +1

    cs.CVarXiv:1501.06170v32015
  27. Gradually Vanishing Bridge for Adversarial Domain Adaptation

    Shuhao Cui, Shuhui Wang, Junbao Zhuo +3

    cs.CVarXiv:2003.13183v12020
  28. Predicting Ground-Level Scene Layout from Aerial Imagery

    Menghua Zhai, Zachary Bessinger, Scott Workman +1

    cs.CVarXiv:1612.02709v12016
  29. Recurrent Vision Transformers for Object Detection with Event Cameras

    Mathias Gehrig, Davide Scaramuzza

    cs.CVarXiv:2212.05598v32022
  30. End-to-End Robotic Reinforcement Learning without Reward Engineering

    Avi Singh, Larry Yang, Kristian Hartikainen +2

    cs.LGcs.CVcs.ROarXiv:1904.07854v22019
  31. Stereo Radiance Fields (SRF): Learning View Synthesis for Sparse Views of Novel Scenes

    Julian Chibane, Aayush Bansal, Verica Lazova +1

    cs.CVcs.LGarXiv:2104.06935v12021
  32. ThunderNet: Towards Real-time Generic Object Detection

    Zheng Qin, Zeming Li, Zhaoning Zhang +4

    cs.CVarXiv:1903.11752v32019
  33. Sparse and Dense Data with CNNs: Depth Completion and Semantic Segmentation

    Maximilian Jaritz, Raoul de Charette, Emilie Wirbel +2

    cs.CVarXiv:1808.00769v22018
  34. Vision-Language Navigation with Self-Supervised Auxiliary Reasoning Tasks

    Fengda Zhu, Yi Zhu, Xiaojun Chang +1

    cs.CVarXiv:1911.07883v42019
  35. Feature Space Augmentation for Long-Tailed Data

    Peng Chu, Xiao Bian, Shaopeng Liu +1

    cs.CVarXiv:2008.03673v12020
  36. NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation

    Jiazhao Zhang, Kunyu Wang, Rongtao Xu +6

    cs.CVcs.ROarXiv:2402.15852v72024
  37. GRAM: Generative Radiance Manifolds for 3D-Aware Image Generation

    Yu Deng, Jiaolong Yang, Jianfeng Xiang +1

    cs.CVarXiv:2112.08867v32021
  38. Frequency-aware Feature Fusion for Dense Image Prediction

    Linwei Chen, Ying Fu, Lin Gu +3

    cs.CVcs.AIarXiv:2408.12879v12024
  39. Liquid Warping GAN: A Unified Framework for Human Motion Imitation, Appearance Transfer and Novel View Synthesis

    Wen Liu, Zhixin Piao, Jie Min +3

    cs.CVcs.LGeess.IVarXiv:1909.12224v32019
  40. AdaCoF: Adaptive Collaboration of Flows for Video Frame Interpolation

    Hyeongmin Lee, Taeoh Kim, Tae-young Chung +3

    cs.CVarXiv:1907.10244v32019
  41. Diving Deeper into Underwater Image Enhancement: A Survey

    Saeed Anwar, Chongyi Li

    cs.CVcs.LGeess.IVarXiv:1907.07863v12019
  42. Localizing Objects with Self-Supervised Transformers and no Labels

    Oriane Siméoni, Gilles Puy, Huy V. Vo +6

    cs.CVarXiv:2109.14279v12021
  43. Subcategory-aware Convolutional Neural Networks for Object Proposals and Detection

    Yu Xiang, Wongun Choi, Yuanqing Lin +1

    cs.CVarXiv:1604.04693v32016
  44. LayoutGAN: Generating Graphic Layouts with Wireframe Discriminators

    Jianan Li, Jimei Yang, Aaron Hertzmann +2

    cs.CVarXiv:1901.06767v12019
  45. Self-Calibrating Neural Radiance Fields

    Yoonwoo Jeong, Seokjun Ahn, Christopher Choy +3

    cs.CVarXiv:2108.13826v22021
  46. Shallow-UWnet : Compressed Model for Underwater Image Enhancement

    Ankita Naik, Apurva Swarnakar, Kartik Mittal

    cs.CVeess.IVarXiv:2101.02073v12021
  47. A Light CNN for detecting COVID-19 from CT scans of the chest

    Matteo Polsinelli, Luigi Cinque, Giuseppe Placidi

    eess.IVcs.CVcs.LGarXiv:2004.12837v12020
  48. Semi-supervised Medical Image Classification with Relation-driven Self-ensembling Model

    Quande Liu, Lequan Yu, Luyang Luo +2

    cs.CVarXiv:2005.07377v12020
  49. Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters

    Jiazuo Yu, Yunzhi Zhuge, Lu Zhang +4

    cs.CVarXiv:2403.11549v22024
  50. Neural-Guided RANSAC: Learning Where to Sample Model Hypotheses

    Eric Brachmann, Carsten Rother

    cs.CVarXiv:1905.04132v22019
  51. Combining Language and Vision with a Multimodal Skip-gram Model

    Angeliki Lazaridou, Nghia The Pham, Marco Baroni

    cs.CLcs.CVcs.LGarXiv:1501.02598v32015
  52. GaussianPro: 3D Gaussian Splatting with Progressive Propagation

    Kai Cheng, Xiaoxiao Long, Kaizhi Yang +5

    cs.CVarXiv:2402.14650v12024
  53. Channel Pruning via Automatic Structure Search

    Mingbao Lin, Rongrong Ji, Yuxin Zhang +3

    cs.CVarXiv:2001.08565v32020
  54. Pangu-Weather: A 3D High-Resolution Model for Fast and Accurate Global Weather Forecast

    Kaifeng Bi, Lingxi Xie, Hengheng Zhang +3

    physics.ao-phcs.AIcs.CVarXiv:2211.02556v12022
  55. Capsules for Object Segmentation

    Rodney LaLonde, Ulas Bagci

    stat.MLcs.AIcs.CVarXiv:1804.04241v12018
  56. Fractional Calculus In Image Processing: A Review

    Qi Yang, Dali Chen, Tiebiao Zhao +1

    cs.CVarXiv:1608.03240v12016
  57. MSRF-Net: A Multi-Scale Residual Fusion Network for Biomedical Image Segmentation

    Abhishek Srivastava, Debesh Jha, Sukalpa Chanda +6

    eess.IVcs.CVarXiv:2105.07451v22021
  58. NerfingMVS: Guided Optimization of Neural Radiance Fields for Indoor Multi-view Stereo

    Yi Wei, Shaohui Liu, Yongming Rao +3

    cs.CVarXiv:2109.01129v32021
  59. Representative Forgery Mining for Fake Face Detection

    Chengrui Wang, Weihong Deng

    cs.CVarXiv:2104.06609v12021
  60. Single-Path NAS: Designing Hardware-Efficient ConvNets in less than 4 Hours

    Dimitrios Stamoulis, Ruizhou Ding, Di Wang +4

    cs.LGcs.CVstat.MLarXiv:1904.02877v12019