Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

10,621 to 10,680 of 18,830

  1. Parallel Multi-Dimensional LSTM, With Application to Fast Biomedical Volumetric Image Segmentation

    Marijn F. Stollenga, Wonmin Byeon, Marcus Liwicki +1

    cs.CVcs.LGarXiv:1506.07452v12015
  2. Learning to Fly by Crashing

    Dhiraj Gandhi, Lerrel Pinto, Abhinav Gupta

    cs.ROcs.CVcs.LGarXiv:1704.05588v22017
  3. ACFNet: Attentional Class Feature Network for Semantic Segmentation

    Fan Zhang, Yanqin Chen, Zhihang Li +5

    cs.CVarXiv:1909.09408v32019
  4. Frozen CLIP Models are Efficient Video Learners

    Ziyi Lin, Shijie Geng, Renrui Zhang +6

    cs.CVarXiv:2208.03550v12022
  5. SVTR: Scene Text Recognition with a Single Visual Model

    Yongkun Du, Zhineng Chen, Caiyan Jia +5

    cs.CVarXiv:2205.00159v22022
  6. Contrastive Learning based Hybrid Networks for Long-Tailed Image Classification

    Peng Wang, Kai Han, Xiu-Shen Wei +2

    cs.CVarXiv:2103.14267v12021
  7. Localizing and Orienting Street Views Using Overhead Imagery

    Nam Vo, James Hays

    cs.CVarXiv:1608.00161v22016
  8. Iteratively Pruned Deep Learning Ensembles for COVID-19 Detection in Chest X-rays

    Sivaramakrishnan Rajaraman, Jen Siegelman, Philip O. Alderson +3

    eess.IVcs.CVcs.LGarXiv:2004.08379v32020
  9. Dark Model Adaptation: Semantic Image Segmentation from Daytime to Nighttime

    Dengxin Dai, Luc Van Gool

    cs.CVarXiv:1810.02575v12018
  10. Vox-Fusion: Dense Tracking and Mapping with Voxel-based Neural Implicit Representation

    Xingrui Yang, Hai Li, Hongjia Zhai +3

    cs.CVcs.GRcs.ROarXiv:2210.15858v32022
  11. DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving

    Xiaofeng Wang, Zheng Zhu, Guan Huang +3

    cs.CVarXiv:2309.09777v22023
  12. VisionZip: Longer is Better but Not Necessary in Vision Language Models

    Senqiao Yang, Yukang Chen, Zhuotao Tian +4

    cs.CVcs.AIcs.CLarXiv:2412.04467v22024
  13. Hierarchical Deep Stereo Matching on High-resolution Images

    Gengshan Yang, Joshua Manela, Michael Happold +1

    cs.CVcs.ROarXiv:1912.06704v12019
  14. Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving

    Xiaosong Jia, Zhenjie Yang, Qifeng Li +2

    cs.ROcs.CVarXiv:2406.03877v32024
  15. Small-Object Detection in Remote Sensing Images with End-to-End Edge-Enhanced GAN and Object Detector Network

    Jakaria Rabbi, Nilanjan Ray, Matthias Schubert +2

    cs.CVcs.LGarXiv:2003.09085v52020
  16. Image Captioning with Deep Bidirectional LSTMs

    Cheng Wang, Haojin Yang, Christian Bartz +1

    cs.CVcs.CLcs.MMarXiv:1604.00790v32016
  17. OccuSeg: Occupancy-aware 3D Instance Segmentation

    Lei Han, Tian Zheng, Lan Xu +1

    cs.CVarXiv:2003.06537v32020
  18. Spatial Aggregation of Holistically-Nested Convolutional Neural Networks for Automated Pancreas Localization and Segmentation

    Holger R. Roth, Le Lu, Nathan Lay +4

    cs.CVarXiv:1702.00045v12017
  19. ImageDream: Image-Prompt Multi-view Diffusion for 3D Generation

    Peng Wang, Yichun Shi

    cs.CVarXiv:2312.02201v12023
  20. Explaining Image Classifiers by Counterfactual Generation

    Chun-Hao Chang, Elliot Creager, Anna Goldenberg +1

    cs.CVarXiv:1807.08024v32018
  21. Detailed, accurate, human shape estimation from clothed 3D scan sequences

    Chao Zhang, Sergi Pujades, Michael Black +1

    cs.CVarXiv:1703.04454v22017
  22. Zero-Shot Video Question Answering via Frozen Bidirectional Language Models

    Antoine Yang, Antoine Miech, Josef Sivic +2

    cs.CVcs.CLcs.LGarXiv:2206.08155v22022
  23. Deep Sketch Hashing: Fast Free-hand Sketch-Based Image Retrieval

    Li Liu, Fumin Shen, Yuming Shen +2

    cs.CVarXiv:1703.05605v12017
  24. Masked Generative Distillation

    Zhendong Yang, Zhe Li, Mingqi Shao +3

    cs.CVarXiv:2205.01529v22022
  25. Beyond Gaussian Pyramid: Multi-skip Feature Stacking for Action Recognition

    Zhenzhong Lan, Ming Lin, Xuanchong Li +2

    cs.CVarXiv:1411.6660v42014
  26. Small Data Challenges in Big Data Era: A Survey of Recent Progress on Unsupervised and Semi-Supervised Methods

    Guo-Jun Qi, Jiebo Luo

    cs.CVarXiv:1903.11260v22019
  27. ROI-10D: Monocular Lifting of 2D Detection to 6D Pose and Metric Shape

    Fabian Manhardt, Wadim Kehl, Adrien Gaidon

    cs.CVarXiv:1812.02781v32018
  28. TextureGAN: Controlling Deep Image Synthesis with Texture Patches

    Wenqi Xian, Patsorn Sangkloy, Varun Agrawal +5

    cs.CVcs.GRarXiv:1706.02823v32017
  29. STAR: A Structure and Texture Aware Retinex Model

    Jun Xu, Yingkun Hou, Dongwei Ren +5

    cs.CVarXiv:1906.06690v52019
  30. Partial success in closing the gap between human and machine vision

    Robert Geirhos, Kantharaju Narayanappa, Benjamin Mitzkus +4

    cs.CVcs.AIcs.LGarXiv:2106.07411v22021
  31. Hierarchical Open-Vocabulary 3D Scene Graphs for Language-Grounded Robot Navigation

    Abdelrhman Werby, Chenguang Huang, Martin Büchner +2

    cs.ROcs.AIcs.CLarXiv:2403.17846v22024
  32. More is Less: A More Complicated Network with Less Inference Complexity

    Xuanyi Dong, Junshi Huang, Yi Yang +1

    cs.CVarXiv:1703.08651v22017
  33. A Bio-Inspired Multi-Exposure Fusion Framework for Low-light Image Enhancement

    Zhenqiang Ying, Ge Li, Wen Gao

    cs.CVarXiv:1711.00591v12017
  34. VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image Captioning

    Jun Chen, Han Guo, Kai Yi +2

    cs.CVcs.AIcs.CLarXiv:2102.10407v52021
  35. Self-Erasing Network for Integral Object Attention

    Qibin Hou, Peng-Tao Jiang, Yunchao Wei +1

    cs.CVarXiv:1810.09821v12018
  36. Investigating Bi-Level Optimization for Learning and Vision from a Unified Perspective: A Survey and Beyond

    Risheng Liu, Jiaxin Gao, Jin Zhang +2

    cs.LGcs.CVmath.DSarXiv:2101.11517v32021
  37. Editing Conditional Radiance Fields

    Steven Liu, Xiuming Zhang, Zhoutong Zhang +3

    cs.CVcs.GRcs.LGarXiv:2105.06466v22021
  38. Learning to Self-Train for Semi-Supervised Few-Shot Classification

    Xinzhe Li, Qianru Sun, Yaoyao Liu +4

    cs.CVcs.LGstat.MLarXiv:1906.00562v22019
  39. MVTN: Multi-View Transformation Network for 3D Shape Recognition

    Abdullah Hamdi, Silvio Giancola, Bernard Ghanem

    cs.CVcs.LGarXiv:2011.13244v32020
  40. Video Object Segmentation with Episodic Graph Memory Networks

    Xiankai Lu, Wenguan Wang, Martin Danelljan +3

    cs.CVcs.LGarXiv:2007.07020v42020
  41. CityGaussian: Real-time High-quality Large-Scale Scene Rendering with Gaussians

    Yang Liu, He Guan, Chuanchen Luo +4

    cs.CVarXiv:2404.01133v32024
  42. Yedrouj-Net: An efficient CNN for spatial steganalysis

    Mehdi Yedroudj, Frederic Comby, Marc Chaumont

    cs.CVcs.CRarXiv:1803.00407v12018
  43. Point-SLAM: Dense Neural Point Cloud-based SLAM

    Erik Sandström, Yue Li, Luc Van Gool +1

    cs.CVarXiv:2304.04278v32023
  44. Neural Architecture Search on ImageNet in Four GPU Hours: A Theoretically Inspired Perspective

    Wuyang Chen, Xinyu Gong, Zhangyang Wang

    cs.CVcs.LGarXiv:2102.11535v42021
  45. NeW CRFs: Neural Window Fully-connected CRFs for Monocular Depth Estimation

    Weihao Yuan, Xiaodong Gu, Zuozhuo Dai +2

    cs.CVarXiv:2203.01502v22022
  46. Rethinking Visual Geo-localization for Large-Scale Applications

    Gabriele Berton, Carlo Masone, Barbara Caputo

    cs.CVarXiv:2204.02287v22022
  47. ParticleNet: Jet Tagging via Particle Clouds

    Huilin Qu, Loukas Gouskos

    hep-phcs.CVhep-exarXiv:1902.08570v32019
  48. AnatomyNet: Deep Learning for Fast and Fully Automated Whole-volume Segmentation of Head and Neck Anatomy

    Wentao Zhu, Yufang Huang, Liang Zeng +6

    cs.CVcs.LGcs.NEarXiv:1808.05238v22018
  49. Reconstruction of three-dimensional porous media using generative adversarial neural networks

    Lukas Mosser, Olivier Dubrule, Martin J. Blunt

    cs.CVcond-mat.mtrl-sciphysics.flu-dynarXiv:1704.03225v12017
  50. FPNN: Field Probing Neural Networks for 3D Data

    Yangyan Li, Soeren Pirk, Hao Su +2

    cs.CVarXiv:1605.06240v32016
  51. Self-supervised Pretraining of Visual Features in the Wild

    Priya Goyal, Mathilde Caron, Benjamin Lefaudeux +8

    cs.CVcs.AIarXiv:2103.01988v22021
  52. Unconstrained Face Verification using Deep CNN Features

    Jun-Cheng Chen, Vishal M. Patel, Rama Chellappa

    cs.CVarXiv:1508.01722v22015
  53. CLIPDraw: Exploring Text-to-Drawing Synthesis through Language-Image Encoders

    Kevin Frans, L. B. Soros, Olaf Witkowski

    cs.CVarXiv:2106.14843v12021
  54. Reviving Iterative Training with Mask Guidance for Interactive Segmentation

    Konstantin Sofiiuk, Ilia A. Petrov, Anton Konushin

    cs.CVarXiv:2102.06583v12021
  55. PIoU Loss: Towards Accurate Oriented Object Detection in Complex Environments

    Zhiming Chen, Kean Chen, Weiyao Lin +4

    cs.CVarXiv:2007.09584v12020
  56. DynaSLAM II: Tightly-Coupled Multi-Object Tracking and SLAM

    Berta Bescos, Carlos Campos, Juan D. Tardós +1

    cs.ROcs.CVarXiv:2010.07820v12020
  57. GLEAN: Generative Latent Bank for Large-Factor Image Super-Resolution

    Kelvin C. K. Chan, Xintao Wang, Xiangyu Xu +2

    cs.CVarXiv:2012.00739v12020
  58. An Efficient Sampling-based Method for Online Informative Path Planning in Unknown Environments

    Lukas Schmid, Michael Pantic, Raghav Khanna +3

    cs.ROcs.CVarXiv:1909.09548v22019
  59. Uncertainty-aware Joint Salient Object and Camouflaged Object Detection

    Aixuan Li, Jing Zhang, Yunqiu Lv +3

    cs.CVarXiv:2104.02628v12021
  60. SAR image despeckling through convolutional neural networks

    G. Chierchia, D. Cozzolino, G. Poggi +1

    cs.CVarXiv:1704.00275v22017