Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

9,961 to 10,020 of 18,822

  1. Do Medical Vision Models Reason About Anatomy? Probing the Spatial Inductive Biases of Learned Visual Representations

    Naren Akash, Neeraja Ramanan

    eess.IVcs.AIcs.CVarXiv:2608.28092v12026
  2. VersaGauss: A Versatile Framework for Generating Multiphase Dynamics with 3D Gaussians

    Ruijie Su, Lingxiao Yang, Xiaohua Xie +1

    cs.CVcs.AIarXiv:2608.28069v12026
  3. Dynamic Alignment Compensation for Hallucination Mitigation in Large Vision-Language Models

    Kairong Yu, Zixin Zhu, Le Yu +1

    cs.CVcs.AIarXiv:2608.28058v12026
  4. Convolutional Neural Network-based Place Recognition

    Zetao Chen, Obadiah Lam, Adam Jacobson +1

    cs.CVcs.LGcs.NEarXiv:1411.1509v12014
  5. Dynamic Curriculum Learning for Imbalanced Data Classification

    Yiru Wang, Weihao Gan, Jie Yang +2

    cs.CVarXiv:1901.06783v22019
  6. DetCLIP: Dictionary-Enriched Visual-Concept Paralleled Pre-training for Open-world Detection

    Lewei Yao, Jianhua Han, Youpeng Wen +6

    cs.CVarXiv:2209.09407v22022
  7. Lightweight Classification of IoT Malware based on Image Recognition

    Jiawei Su, Danilo Vasconcellos Vargas, Sanjiva Prasad +3

    cs.CRcs.CVarXiv:1802.03714v12018
  8. Not All Images are Worth 16x16 Words: Dynamic Transformers for Efficient Image Recognition

    Yulin Wang, Rui Huang, Shiji Song +2

    cs.CVcs.AIcs.LGarXiv:2105.15075v22021
  9. 3D Self-Supervised Methods for Medical Imaging

    Aiham Taleb, Winfried Loetzsch, Noel Danz +4

    cs.CVcs.LGeess.IVarXiv:2006.03829v32020
  10. Adversarial Camouflage: Hiding Physical-World Attacks with Natural Styles

    Ranjie Duan, Xingjun Ma, Yisen Wang +3

    cs.CVarXiv:2003.08757v22020
  11. Knowledge Distillation with the Reused Teacher Classifier

    Defang Chen, Jian-Ping Mei, Hailin Zhang +3

    cs.CVarXiv:2203.14001v12022
  12. Tensor-Train Recurrent Neural Networks for Video Classification

    Yinchong Yang, Denis Krompass, Volker Tresp

    cs.CVarXiv:1707.01786v12017
  13. PCBnet: A Dataset and Automatic Construction of SPICE Netlists from Schematic Images

    Zhen Huang, Yuhao Gao, Yuzhi Liu +15

    cs.CVcs.AIarXiv:2608.27923v12026
  14. Image-to-Markup Generation with Coarse-to-Fine Attention

    Yuntian Deng, Anssi Kanervisto, Jeffrey Ling +1

    cs.CVcs.CLcs.LGarXiv:1609.04938v22016
  15. UAV-Human: A Large Benchmark for Human Behavior Understanding with Unmanned Aerial Vehicles

    Tianjiao Li, Jun Liu, Wei Zhang +3

    cs.CVarXiv:2104.00946v42021
  16. 3D UX-Net: A Large Kernel Volumetric ConvNet Modernizing Hierarchical Transformer for Medical Image Segmentation

    Ho Hin Lee, Shunxing Bao, Yuankai Huo +1

    cs.CVcs.LGarXiv:2209.15076v42022
  17. Deep Kinematic Pose Regression

    Xingyi Zhou, Xiao Sun, Wei Zhang +2

    cs.CVarXiv:1609.05317v12016
  18. Tracking Anything with Decoupled Video Segmentation

    Ho Kei Cheng, Seoung Wug Oh, Brian Price +2

    cs.CVarXiv:2309.03903v12023
  19. Frequency Separation for Real-World Super-Resolution

    Manuel Fritsche, Shuhang Gu, Radu Timofte

    eess.IVcs.CVarXiv:1911.07850v12019
  20. An Enhanced Deep Feature Representation for Person Re-identification

    Shangxuan Wu, Ying-Cong Chen, Xiang Li +3

    cs.CVarXiv:1604.07807v22016
  21. Grid-GCN for Fast and Scalable Point Cloud Learning

    Qiangeng Xu, Xudong Sun, Cho-Ying Wu +2

    cs.CVcs.LGarXiv:1912.02984v52019
  22. SNARF: Differentiable Forward Skinning for Animating Non-Rigid Neural Implicit Shapes

    Xu Chen, Yufeng Zheng, Michael J. Black +2

    cs.CVarXiv:2104.03953v32021
  23. Zoom-in-Net: Deep Mining Lesions for Diabetic Retinopathy Detection

    Zhe Wang, Yanxin Yin, Jianping Shi +3

    cs.CVarXiv:1706.04372v12017
  24. Multi-Scale Geometric Consistency Guided Multi-View Stereo

    Qingshan Xu, Wenbing Tao

    cs.CVarXiv:1904.08103v12019
  25. Uncertainty Estimates and Multi-Hypotheses Networks for Optical Flow

    Eddy Ilg, Özgün Çiçek, Silvio Galesso +4

    cs.CVarXiv:1802.07095v42018
  26. You said that?

    Joon Son Chung, Amir Jamaludin, Andrew Zisserman

    cs.CVarXiv:1705.02966v22017
  27. Neural Head Avatars from Monocular RGB Videos

    Philip-William Grassal, Malte Prinzler, Titus Leistner +3

    cs.CVcs.GRarXiv:2112.01554v22021
  28. Overcoming Language Priors in Visual Question Answering with Adversarial Regularization

    Sainandan Ramakrishnan, Aishwarya Agrawal, Stefan Lee

    cs.CVarXiv:1810.03649v22018
  29. DeepGMR: Learning Latent Gaussian Mixture Models for Registration

    Wentao Yuan, Ben Eckart, Kihwan Kim +3

    cs.CVarXiv:2008.09088v12020
  30. Soft Threshold Weight Reparameterization for Learnable Sparsity

    Aditya Kusupati, Vivek Ramanujan, Raghav Somani +4

    cs.LGcs.CVstat.MLarXiv:2002.03231v92020
  31. From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation

    Rit Gangopadhyay, Alex Wong

    cs.CVcs.AIarXiv:2608.27860v12026
  32. MinkLoc3D: Point Cloud Based Large-Scale Place Recognition

    Jacek Komorowski

    cs.CVarXiv:2011.04530v12020
  33. DexGraspNet: A Large-Scale Robotic Dexterous Grasp Dataset for General Objects Based on Simulation

    Ruicheng Wang, Jialiang Zhang, Jiayi Chen +4

    cs.ROcs.CVarXiv:2210.02697v22022
  34. Sequential Attend, Infer, Repeat: Generative Modelling of Moving Objects

    Adam R. Kosiorek, Hyunjik Kim, Ingmar Posner +1

    cs.LGcs.CVstat.MLarXiv:1806.01794v22018
  35. Semi-Supervised Semantic Segmentation with Pixel-Level Contrastive Learning from a Class-wise Memory Bank

    Inigo Alonso, Alberto Sabater, David Ferstl +2

    cs.CVarXiv:2104.13415v32021
  36. EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions

    Linrui Tian, Qi Wang, Bang Zhang +1

    cs.CVarXiv:2402.17485v32024
  37. Pose-guided Feature Disentangling for Occluded Person Re-identification Based on Transformer

    Tao Wang, Hong Liu, Pinhao Song +2

    cs.CVarXiv:2112.02466v22021
  38. DeepLPF: Deep Local Parametric Filters for Image Enhancement

    Sean Moran, Pierre Marza, Steven McDonagh +2

    cs.CVarXiv:2003.13985v12020
  39. MoFusion: A Framework for Denoising-Diffusion-based Motion Synthesis

    Rishabh Dabral, Muhammad Hamza Mughal, Vladislav Golyanik +1

    cs.CVarXiv:2212.04495v32022
  40. Deep Learning for Unsupervised Anomaly Localization in Industrial Images: A Survey

    Xian Tao, Xinyi Gong, Xin Zhang +2

    cs.CVarXiv:2207.10298v12022
  41. MultiPoseNet: Fast Multi-Person Pose Estimation using Pose Residual Network

    Muhammed Kocabas, Salih Karagoz, Emre Akbas

    cs.CVarXiv:1807.04067v12018
  42. CARDINAL Predicts Cardiovascular Risk From Non-contrast Cardiac CT

    Roy Gabriel, Nattakorn Kittisut, Jamshid Hassanpour +6

    eess.IVcs.AIcs.CVarXiv:2608.27690v12026
  43. Self-supervised Multi-level Face Model Learning for Monocular Reconstruction at over 250 Hz

    Ayush Tewari, Michael Zollhöfer, Pablo Garrido +4

    cs.CVarXiv:1712.02859v22017
  44. VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

    Ziyang Wang, Shoubin Yu, Elias Stengel-Eskin +4

    cs.CVcs.AIcs.CLarXiv:2405.19209v32024
  45. Do Deep Neural Networks Learn Facial Action Units When Doing Expression Recognition?

    Pooya Khorrami, Tom Le Paine, Thomas S. Huang

    cs.CVcs.LGcs.NEarXiv:1510.02969v32015
  46. Railroad is not a Train: Saliency as Pseudo-pixel Supervision for Weakly Supervised Semantic Segmentation

    Seungho Lee, Minhyun Lee, Jongwuk Lee +1

    cs.CVarXiv:2105.08965v12021
  47. USE-Net: incorporating Squeeze-and-Excitation blocks into U-Net for prostate zonal segmentation of multi-institutional MRI datasets

    Leonardo Rundo, Changhee Han, Yudai Nagano +12

    cs.CVcs.LGarXiv:1904.08254v22019
  48. FVeinSyn: Synthetic Finger Vein Image Generator

    Yifan Wang, Jie Gui, Adams Wai Kin Kong +6

    cs.CVcs.AIarXiv:2608.27527v12026
  49. Segment and Track Anything

    Yangming Cheng, Liulei Li, Yuanyou Xu +4

    cs.CVarXiv:2305.06558v12023
  50. Taming 3DGS: High-Quality Radiance Fields with Limited Resources

    Saswat Subhajyoti Mallick, Rahul Goel, Bernhard Kerbl +3

    cs.CVcs.GRarXiv:2406.15643v12024
  51. Training Neural Networks with Local Error Signals

    Arild Nøkland, Lars Hiller Eidnes

    stat.MLcs.CVcs.LGarXiv:1901.06656v22019
  52. TI-POOLING: transformation-invariant pooling for feature learning in Convolutional Neural Networks

    Dmitry Laptev, Nikolay Savinov, Joachim M. Buhmann +1

    cs.CVarXiv:1604.06318v22016
  53. Curvature-Aware Radius Shrinkage for Adaptive Nearest Neighbor Classification

    Alexandre L. M. Levada

    cs.LGcs.AIcs.CVarXiv:2608.27634v12026
  54. Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIP

    Qihang Yu, Ju He, Xueqing Deng +2

    cs.CVarXiv:2308.02487v22023
  55. T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation

    Kaiyi Huang, Chengqi Duan, Kaiyue Sun +3

    cs.CVarXiv:2307.06350v32023
  56. Depth-Aware Pothole Detection Using YOLO and RT-DETR at the Edge

    Md Monjurul Ahsan Prodhan, Md Nour Hossain

    cs.CVcs.AIcs.LGarXiv:2608.27633v12026
  57. Segment Anything Is Not Always Perfect: An Investigation of SAM on Different Real-world Applications

    Wei Ji, Jingjing Li, Qi Bi +3

    cs.CVarXiv:2304.05750v32023
  58. Quanta Perception as Probabilistic Events

    Varun Sundar, Pavan Thodima, Sacha Jungerman +1

    cs.CVcs.AIarXiv:2608.27584v12026
  59. Woodpecker: Hallucination Correction for Multimodal Large Language Models

    Shukang Yin, Chaoyou Fu, Sirui Zhao +7

    cs.CVcs.AIcs.CLarXiv:2310.16045v22023
  60. On Translation Invariance in CNNs: Convolutional Layers can Exploit Absolute Spatial Location

    Osman Semih Kayhan, Jan C. van Gemert

    cs.CVcs.LGeess.IVarXiv:2003.07064v22020