Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1,681 to 1,740 of 18,830

  1. A multilevel thresholding algorithm using Electromagnetism Optimization

    Diego Oliva, Erik Cuevas, Gonzalo Pajares +2

    cs.CVarXiv:1406.6336v12014
  2. Novel Visual Category Discovery with Dual Ranking Statistics and Mutual Knowledge Distillation

    Bingchen Zhao, Kai Han

    cs.CVarXiv:2107.03358v22021
  3. MSeg3D: Multi-modal 3D Semantic Segmentation for Autonomous Driving

    Jiale Li, Hang Dai, Hao Han +1

    cs.CVarXiv:2303.08600v12023
  4. Beyond Physical Connections: Tree Models in Human Pose Estimation

    Fang Wang, Yi Li

    cs.CVarXiv:1305.2269v12013
  5. Multi-Angle Point Cloud-VAE: Unsupervised Feature Learning for 3D Point Clouds from Multiple Angles by Joint Self-Reconstruction and Half-to-Half Prediction

    Zhizhong Han, Xiyang Wang, Yu-Shen Liu +1

    cs.CVarXiv:1907.12704v12019
  6. Planar Prior Assisted PatchMatch Multi-View Stereo

    Qingshan Xu, Wenbing Tao

    cs.CVarXiv:1912.11744v12019
  7. Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

    Homanga Bharadhwaj, Roozbeh Mottaghi, Abhinav Gupta +1

    cs.ROcs.CVarXiv:2405.01527v22024
  8. A Closer Look at the Explainability of Contrastive Language-Image Pre-training

    Yi Li, Hualiang Wang, Yiqun Duan +2

    cs.CVarXiv:2304.05653v22023
  9. Total Denoising: Unsupervised Learning of 3D Point Cloud Cleaning

    Pedro Hermosilla, Tobias Ritschel, Timo Ropinski

    cs.CVcs.GRarXiv:1904.07615v22019
  10. Analyzing and Mitigating the Impact of Permanent Faults on a Systolic Array Based Neural Network Accelerator

    Jeff Zhang, Tianyu Gu, Kanad Basu +1

    cs.LGcs.ARcs.CVarXiv:1802.04657v22018
  11. VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding

    Xiang Li, Jian Ding, Mohamed Elhoseiny

    cs.CVarXiv:2406.12384v22024
  12. Out-of-Distribution Detection for Generalized Zero-Shot Action Recognition

    Devraj Mandal, Sanath Narayan, Saikumar Dwivedi +4

    cs.CVarXiv:1904.08703v22019
  13. Improved Techniques for Training Adaptive Deep Networks

    Hao Li, Hong Zhang, Xiaojuan Qi +2

    cs.CVarXiv:1908.06294v12019
  14. Avatars Grow Legs: Generating Smooth Human Motion from Sparse Tracking Inputs with Diffusion Model

    Yuming Du, Robin Kips, Albert Pumarola +3

    cs.CVarXiv:2304.08577v12023
  15. Low-Resolution Face Recognition

    Zhiyi Cheng, Xiatian Zhu, Shaogang Gong

    cs.CVarXiv:1811.08965v22018
  16. GlyphControl: Glyph Conditional Control for Visual Text Generation

    Yukang Yang, Dongnan Gui, Yuhui Yuan +4

    cs.CVarXiv:2305.18259v22023
  17. CLIP-Count: Towards Text-Guided Zero-Shot Object Counting

    Ruixiang Jiang, Lingbo Liu, Changwen Chen

    cs.CVcs.AIarXiv:2305.07304v22023
  18. Filmy Cloud Removal on Satellite Imagery with Multispectral Conditional Generative Adversarial Nets

    Kenji Enomoto, Ken Sakurada, Weimin Wang +4

    cs.CVarXiv:1710.04835v12017
  19. CovidAID: COVID-19 Detection Using Chest X-Ray

    Arpan Mangal, Surya Kalia, Harish Rajgopal +4

    eess.IVcs.CVcs.LGarXiv:2004.09803v12020
  20. Joint-task Self-supervised Learning for Temporal Correspondence

    Xueting Li, Sifei Liu, Shalini De Mello +3

    cs.CVarXiv:1909.11895v12019
  21. Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching

    Di Hu, Rui Qian, Minyue Jiang +5

    cs.CVcs.LGcs.MMarXiv:2010.05466v12020
  22. FACIAL: Synthesizing Dynamic Talking Face with Implicit Attribute Learning

    Chenxu Zhang, Yifan Zhao, Yifei Huang +4

    cs.CVarXiv:2108.07938v12021
  23. GRiT: A Generative Region-to-text Transformer for Object Understanding

    Jialian Wu, Jianfeng Wang, Zhengyuan Yang +4

    cs.CVarXiv:2212.00280v12022
  24. View-Structured Conformal Prediction for 3D Gaussian Splatting

    Junzheng Chu, Bin Pan, Zhenwei Shi

    cs.LGcs.CVarXiv:2609.10307v12026
  25. Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance

    Phuc D. A. Nguyen, Tuan Duc Ngo, Evangelos Kalogerakis +4

    cs.CVarXiv:2312.10671v32023
  26. Learning A Single Network for Scale-Arbitrary Super-Resolution

    Longguang Wang, Yingqian Wang, Zaiping Lin +3

    cs.CVarXiv:2004.03791v22020
  27. GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

    Drew A. Hudson, Christopher D. Manning

    cs.CLcs.AIcs.CVarXiv:1902.09506v32019
  28. Point-Set Anchors for Object Detection, Instance Segmentation and Pose Estimation

    Fangyun Wei, Xiao Sun, Hongyang Li +2

    cs.CVarXiv:2007.02846v42020
  29. Towards Flexible Blind JPEG Artifacts Removal

    Jiaxi Jiang, Kai Zhang, Radu Timofte

    eess.IVcs.CVarXiv:2109.14573v12021
  30. SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text Recognition

    Mingxin Huang, Yuliang Liu, Zhenghao Peng +6

    cs.CVarXiv:2203.10209v12022
  31. Direct Inversion: Boosting Diffusion-based Editing with 3 Lines of Code

    Xuan Ju, Ailing Zeng, Yuxuan Bian +2

    cs.CVarXiv:2310.01506v22023
  32. Text-Only Training for Image Captioning using Noise-Injected CLIP

    David Nukrai, Ron Mokady, Amir Globerson

    cs.CVcs.AIcs.LGarXiv:2211.00575v12022
  33. Nutrition5k: Towards Automatic Nutritional Understanding of Generic Food

    Quin Thames, Arjun Karpur, Wade Norris +4

    cs.CVcs.LGarXiv:2103.03375v22021
  34. Feature Learning from Incomplete EEG with Denoising Autoencoder

    Junhua Li, Zbigniew Struzik, Liqing Zhang +1

    cs.CVq-bio.NCarXiv:1410.0818v12014
  35. Uncertainty-Informed Deep Learning Models Enable High-Confidence Predictions for Digital Histopathology

    James M Dolezal, Andrew Srisuwananukorn, Dmitry Karpeyev +13

    q-bio.QMcs.CVeess.IVarXiv:2204.04516v12022
  36. Curriculum Model Adaptation with Synthetic and Real Data for Semantic Foggy Scene Understanding

    Dengxin Dai, Christos Sakaridis, Simon Hecker +1

    cs.CVarXiv:1901.01415v22019
  37. Cross-Attention of Disentangled Modalities for 3D Human Mesh Recovery with Transformers

    Junhyeong Cho, Kim Youwang, Tae-Hyun Oh

    cs.CVcs.AIcs.LGarXiv:2207.13820v12022
  38. MMM: Generative Masked Motion Model

    Ekkasit Pinyoanuntapong, Pu Wang, Minwoo Lee +1

    cs.CVcs.AIcs.LGarXiv:2312.03596v22023
  39. DigiFace-1M: 1 Million Digital Face Images for Face Recognition

    Gwangbin Bae, Martin de La Gorce, Tadas Baltrusaitis +5

    cs.CVarXiv:2210.02579v12022
  40. Hashing on Nonlinear Manifolds

    Fumin Shen, Chunhua Shen, Qinfeng Shi +3

    cs.CVarXiv:1412.0826v12014
  41. Manipulation by Feel: Touch-Based Control with Deep Predictive Models

    Stephen Tian, Frederik Ebert, Dinesh Jayaraman +4

    cs.ROcs.AIcs.CVarXiv:1903.04128v12019
  42. Dark, Beyond Deep: A Paradigm Shift to Cognitive AI with Humanlike Common Sense

    Yixin Zhu, Tao Gao, Lifeng Fan +9

    cs.AIcs.CVcs.LGarXiv:2004.09044v12020
  43. Photorealistic Image Synthesis for Object Instance Detection

    Tomas Hodan, Vibhav Vineet, Ran Gal +6

    cs.CVcs.AIcs.ROarXiv:1902.03334v12019
  44. Temporal Pyramid Pooling Based Convolutional Neural Networks for Action Recognition

    Peng Wang, Yuanzhouhan Cao, Chunhua Shen +2

    cs.CVarXiv:1503.01224v22015
  45. Controlling Vision-Language Models for Multi-Task Image Restoration

    Ziwei Luo, Fredrik K. Gustafsson, Zheng Zhao +2

    cs.CVarXiv:2310.01018v22023
  46. Active Learning for Deep Detection Neural Networks

    Hamed H. Aghdam, Abel Gonzalez-Garcia, Joost van de Weijer +1

    cs.CVcs.LGarXiv:1911.09168v12019
  47. Unsupervised Domain Adaptation via Disentangled Representations: Application to Cross-Modality Liver Segmentation

    Junlin Yang, Nicha C. Dvornek, Fan Zhang +3

    eess.IVcs.CVarXiv:1907.13590v22019
  48. Newtonian Image Understanding: Unfolding the Dynamics of Objects in Static Images

    Roozbeh Mottaghi, Hessam Bagherinezhad, Mohammad Rastegari +1

    cs.CVarXiv:1511.04048v12015
  49. SegICP: Integrated Deep Semantic Segmentation and Pose Estimation

    Jay M. Wong, Vincent Kee, Tiffany Le +10

    cs.ROcs.CVarXiv:1703.01661v22017
  50. Learning Descriptor Networks for 3D Shape Synthesis and Analysis

    Jianwen Xie, Zilong Zheng, Ruiqi Gao +3

    cs.CVarXiv:1804.00586v12018
  51. Learning representations of irregular particle-detector geometry with distance-weighted graph networks

    Shah Rukh Qasim, Jan Kieseler, Yutaro Iiyama +1

    physics.data-ancs.CVcs.LGarXiv:1902.07987v22019
  52. Bridging 2D and 3D Segmentation Networks for Computation Efficient Volumetric Medical Image Segmentation: An Empirical Study of 2.5D Solutions

    Yichi Zhang, Qingcheng Liao, Le Ding +1

    eess.IVcs.CVarXiv:2010.06163v22020
  53. Polyp-SAM: Transfer SAM for Polyp Segmentation

    Yuheng Li, Mingzhe Hu, Xiaofeng Yang

    eess.IVcs.CVarXiv:2305.00293v12023
  54. An Experimental Study of Deep Convolutional Features For Iris Recognition

    Shervin Minaee, Amirali Abdolrashidi, Yao Wang

    cs.CVcs.LGarXiv:1702.01334v12017
  55. Designing BERT for Convolutional Networks: Sparse and Hierarchical Masked Modeling

    Keyu Tian, Yi Jiang, Qishuai Diao +3

    cs.CVcs.AIcs.LGarXiv:2301.03580v22023
  56. DF26: We Cannot Tell Fake From Real Anymore

    Severyn Shykula, Andrii Yermakov, Ivan Samarskyi +3

    cs.CVarXiv:2609.07369v12026
  57. From W-Net to CDGAN: Bi-temporal Change Detection via Deep Learning Techniques

    Bin Hou, Qingjie Liu, Heng Wang +1

    cs.CVeess.IVarXiv:2003.06583v12020
  58. Swapout: Learning an ensemble of deep architectures

    Saurabh Singh, Derek Hoiem, David Forsyth

    cs.CVcs.LGcs.NEarXiv:1605.06465v12016
  59. X-Net: Brain Stroke Lesion Segmentation Based on Depthwise Separable Convolution and Long-range Dependencies

    Kehan Qi, Hao Yang, Cheng Li +4

    eess.IVcs.CVarXiv:1907.07000v22019
  60. Optimized U-Net for Brain Tumor Segmentation

    Michał Futrega, Alexandre Milesi, Michal Marcinkiewicz +1

    eess.IVcs.CVcs.LGarXiv:2110.03352v22021