Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

10,261 to 10,320 of 18,839

  1. Shape-aware Meta-learning for Generalizing Prostate MRI Segmentation to Unseen Domains

    Quande Liu, Qi Dou, Pheng-Ann Heng

    cs.CVarXiv:2007.02035v12020
  2. DynIBaR: Neural Dynamic Image-Based Rendering

    Zhengqi Li, Qianqian Wang, Forrester Cole +2

    cs.CVarXiv:2211.11082v32022
  3. Delving Deep into Label Smoothing

    Chang-Bin Zhang, Peng-Tao Jiang, Qibin Hou +4

    cs.CVarXiv:2011.12562v22020
  4. Adaptive Aggregation Networks for Class-Incremental Learning

    Yaoyao Liu, Bernt Schiele, Qianru Sun

    cs.CVstat.MLarXiv:2010.05063v32020
  5. BirdNet: a 3D Object Detection Framework from LiDAR information

    Jorge Beltran, Carlos Guindel, Francisco Miguel Moreno +3

    cs.CVarXiv:1805.01195v12018
  6. Oriented Response Networks

    Yanzhao Zhou, Qixiang Ye, Qiang Qiu +1

    cs.CVarXiv:1701.01833v22017
  7. PReMVOS: Proposal-generation, Refinement and Merging for Video Object Segmentation

    Jonathon Luiten, Paul Voigtlaender, Bastian Leibe

    cs.CVarXiv:1807.09190v22018
  8. Poisson multi-Bernoulli mixture filter: direct derivation and implementation

    Ángel F. García-Fernández, Jason L. Williams, Karl Granström +1

    cs.CVstat.MEarXiv:1703.04264v42017
  9. MagicDrive: Street View Generation with Diverse 3D Geometry Control

    Ruiyuan Gao, Kai Chen, Enze Xie +4

    cs.CVcs.AIarXiv:2310.02601v72023
  10. Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning

    Hanyang Wang, Yimo Cai, Weiliang Chen +14

    cs.CVarXiv:2608.27549v12026
  11. Learning Parallax Attention for Stereo Image Super-Resolution

    Longguang Wang, Yingqian Wang, Zhengfa Liang +4

    cs.CVarXiv:1903.05784v32019
  12. HarDNet-MSEG: A Simple Encoder-Decoder Polyp Segmentation Neural Network that Achieves over 0.9 Mean Dice and 86 FPS

    Chien-Hsiang Huang, Hung-Yu Wu, Youn-Long Lin

    cs.CVarXiv:2101.07172v22021
  13. OW-DETR: Open-world Detection Transformer

    Akshita Gupta, Sanath Narayan, K J Joseph +3

    cs.CVarXiv:2112.01513v32021
  14. Text2Tex: Text-driven Texture Synthesis via Diffusion Models

    Dave Zhenyu Chen, Yawar Siddiqui, Hsin-Ying Lee +2

    cs.CVarXiv:2303.11396v12023
  15. What Can Neural Networks Reason About?

    Keyulu Xu, Jingling Li, Mozhi Zhang +3

    cs.LGcs.AIcs.CVarXiv:1905.13211v42019
  16. Small Object Detection using Context and Attention

    Jeong-Seon Lim, Marcella Astrid, Hyun-Jin Yoon +1

    cs.CVarXiv:1912.06319v22019
  17. Pixel2Mesh++: Multi-View 3D Mesh Generation via Deformation

    Chao Wen, Yinda Zhang, Zhuwen Li +1

    cs.CVarXiv:1908.01491v22019
  18. Oriented Object Detection in Aerial Images with Box Boundary-Aware Vectors

    Jingru Yi, Pengxiang Wu, Bo Liu +3

    cs.CVarXiv:2008.07043v22020
  19. Adversarial Examples that Fool both Computer Vision and Time-Limited Humans

    Gamaleldin F. Elsayed, Shreya Shankar, Brian Cheung +4

    cs.LGcs.CVq-bio.NCarXiv:1802.08195v32018
  20. H+O: Unified Egocentric Recognition of 3D Hand-Object Poses and Interactions

    Bugra Tekin, Federica Bogo, Marc Pollefeys

    cs.CVarXiv:1904.05349v12019
  21. Fast-MVSNet: Sparse-to-Dense Multi-View Stereo With Learned Propagation and Gauss-Newton Refinement

    Zehao Yu, Shenghua Gao

    cs.CVarXiv:2003.13017v12020
  22. Multi-modal Learning with Missing Modality via Shared-Specific Feature Modelling

    Hu Wang, Yuanhong Chen, Congbo Ma +3

    cs.CVarXiv:2307.14126v22023
  23. Fast, Accurate and Lightweight Super-Resolution with Neural Architecture Search

    Xiangxiang Chu, Bo Zhang, Hailong Ma +2

    cs.CVcs.LGarXiv:1901.07261v32019
  24. What the Constant Velocity Model Can Teach Us About Pedestrian Motion Prediction

    Christoph Schöller, Vincent Aravantinos, Florian Lay +1

    cs.CVcs.LGcs.ROarXiv:1903.07933v32019
  25. LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models

    Peng Xu, Wenqi Shao, Kaipeng Zhang +7

    cs.CVcs.AIarXiv:2306.09265v12023
  26. PeCo: Perceptual Codebook for BERT Pre-training of Vision Transformers

    Xiaoyi Dong, Jianmin Bao, Ting Zhang +7

    cs.CVcs.LGarXiv:2111.12710v32021
  27. Distilling Knowledge from Graph Convolutional Networks

    Yiding Yang, Jiayan Qiu, Mingli Song +2

    cs.CVarXiv:2003.10477v42020
  28. DeepFood: Deep Learning-Based Food Image Recognition for Computer-Aided Dietary Assessment

    Chang Liu, Yu Cao, Yan Luo +3

    cs.CVarXiv:1606.05675v12016
  29. TVQA+: Spatio-Temporal Grounding for Video Question Answering

    Jie Lei, Licheng Yu, Tamara L. Berg +1

    cs.CVcs.AIcs.CLarXiv:1904.11574v22019
  30. SpeedNet: Learning the Speediness in Videos

    Sagie Benaim, Ariel Ephrat, Oran Lang +5

    cs.CVarXiv:2004.06130v22020
  31. Robust Point Cloud Registration Framework Based on Deep Graph Matching

    Kexue Fu, Shaolei Liu, Xiaoyuan Luo +1

    cs.CVarXiv:2103.04256v12021
  32. mmFormer: Multimodal Medical Transformer for Incomplete Multimodal Learning of Brain Tumor Segmentation

    Yao Zhang, Nanjun He, Jiawei Yang +6

    eess.IVcs.CVarXiv:2206.02425v22022
  33. GlobalTrack: A Simple and Strong Baseline for Long-term Tracking

    Lianghua Huang, Xin Zhao, Kaiqi Huang

    cs.CVarXiv:1912.08531v12019
  34. Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review

    Iryna Hartsock, Ghulam Rasool

    cs.CVcs.LGarXiv:2403.02469v22024
  35. Foundations and Trends in Multimodal Machine Learning: Principles, Challenges, and Open Questions

    Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency

    cs.LGcs.AIcs.CLarXiv:2209.03430v22022
  36. AutoLoc: Weakly-supervised Temporal Action Localization

    Zheng Shou, Hang Gao, Lei Zhang +2

    cs.CVarXiv:1807.08333v22018
  37. ESLAM: Efficient Dense SLAM System Based on Hybrid Representation of Signed Distance Fields

    Mohammad Mahdi Johari, Camilla Carta, François Fleuret

    cs.CVarXiv:2211.11704v22022
  38. Panoptic Segmentation of Satellite Image Time Series with Convolutional Temporal Attention Networks

    Vivien Sainte Fare Garnot, Loic Landrieu

    cs.CVarXiv:2107.07933v42021
  39. PST900: RGB-Thermal Calibration, Dataset and Segmentation Network

    Shreyas S. Shivakumar, Neil Rodrigues, Alex Zhou +3

    cs.CVcs.ROeess.IVarXiv:1909.10980v12019
  40. Associatively Segmenting Instances and Semantics in Point Clouds

    Xinlong Wang, Shu Liu, Xiaoyong Shen +2

    cs.CVarXiv:1902.09852v22019
  41. Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation Learning

    Fuying Wang, Yuyin Zhou, Shujun Wang +2

    cs.CVcs.AIcs.CLarXiv:2210.06044v12022
  42. Automated polyp detection in colon capsule endoscopy

    Alexander V. Mamonov, Isabel N. Figueiredo, Pedro N. Figueiredo +1

    cs.CVarXiv:1305.1912v42013
  43. Boundary-Aware Feature Propagation for Scene Segmentation

    Henghui Ding, Xudong Jiang, Ai Qun Liu +2

    cs.CVarXiv:1909.00179v12019
  44. Conditional Image Generation with Score-Based Diffusion Models

    Georgios Batzolis, Jan Stanczuk, Carola-Bibiane Schönlieb +1

    cs.LGcs.CVstat.MLarXiv:2111.13606v12021
  45. Reachability Analysis of Deep Neural Networks with Provable Guarantees

    Wenjie Ruan, Xiaowei Huang, Marta Kwiatkowska

    cs.LGcs.CVstat.MLarXiv:1805.02242v12018
  46. Unsupervised Semantic Segmentation by Contrasting Object Mask Proposals

    Wouter Van Gansbeke, Simon Vandenhende, Stamatios Georgoulis +1

    cs.CVcs.LGarXiv:2102.06191v32021
  47. TEACh: Task-driven Embodied Agents that Chat

    Aishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava +6

    cs.CVcs.AIcs.CLarXiv:2110.00534v32021
  48. SuperPCA: A Superpixelwise PCA Approach for Unsupervised Feature Extraction of Hyperspectral Imagery

    Junjun Jiang, Jiayi Ma, Chen Chen +3

    cs.CVarXiv:1806.09807v22018
  49. Optimizing Prompts for Text-to-Image Generation

    Yaru Hao, Zewen Chi, Li Dong +1

    cs.CLcs.CVarXiv:2212.09611v22022
  50. AnyLoc: Towards Universal Visual Place Recognition

    Nikhil Keetha, Avneesh Mishra, Jay Karhade +4

    cs.CVcs.AIcs.ROarXiv:2308.00688v22023
  51. Exploring Smoothness and Class-Separation for Semi-supervised Medical Image Segmentation

    Yicheng Wu, Zhonghua Wu, Qianyi Wu +2

    eess.IVcs.CVarXiv:2203.01324v32022
  52. Deblur-NeRF: Neural Radiance Fields from Blurry Images

    Li Ma, Xiaoyu Li, Jing Liao +4

    cs.CVcs.GRarXiv:2111.14292v22021
  53. Text-based Editing of Talking-head Video

    Ohad Fried, Ayush Tewari, Michael Zollhöfer +7

    cs.CVcs.GRcs.LGarXiv:1906.01524v12019
  54. DetNAS: Backbone Search for Object Detection

    Yukang Chen, Tong Yang, Xiangyu Zhang +3

    cs.CVarXiv:1903.10979v42019
  55. AD-Cluster: Augmented Discriminative Clustering for Domain Adaptive Person Re-identification

    Yunpeng Zhai, Shijian Lu, Qixiang Ye +4

    cs.CVarXiv:2004.08787v22020
  56. Diagnose like a Radiologist: Attention Guided Convolutional Neural Network for Thorax Disease Classification

    Qingji Guan, Yaping Huang, Zhun Zhong +3

    cs.CVarXiv:1801.09927v12018
  57. AutoGAN: Neural Architecture Search for Generative Adversarial Networks

    Xinyu Gong, Shiyu Chang, Yifan Jiang +1

    cs.CVcs.LGeess.IVarXiv:1908.03835v12019
  58. AttentionGAN: Unpaired Image-to-Image Translation using Attention-Guided Generative Adversarial Networks

    Hao Tang, Hong Liu, Dan Xu +2

    cs.CVcs.LGeess.IVarXiv:1911.11897v52019
  59. Unsupervised Object Discovery and Localization in the Wild: Part-based Matching with Bottom-up Region Proposals

    Minsu Cho, Suha Kwak, Cordelia Schmid +1

    cs.CVarXiv:1501.06170v32015
  60. Gradually Vanishing Bridge for Adversarial Domain Adaptation

    Shuhao Cui, Shuhui Wang, Junbao Zhuo +3

    cs.CVarXiv:2003.13183v12020