Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
9,961 to 10,020 of 18,822
Do Medical Vision Models Reason About Anatomy? Probing the Spatial Inductive Biases of Learned Visual Representations
Naren Akash, Neeraja Ramanan
eess.IVcs.AIcs.CVarXiv:2608.28092v12026VersaGauss: A Versatile Framework for Generating Multiphase Dynamics with 3D Gaussians
Ruijie Su, Lingxiao Yang, Xiaohua Xie +1
cs.CVcs.AIarXiv:2608.28069v12026Dynamic Alignment Compensation for Hallucination Mitigation in Large Vision-Language Models
Kairong Yu, Zixin Zhu, Le Yu +1
cs.CVcs.AIarXiv:2608.28058v12026Convolutional Neural Network-based Place Recognition
Zetao Chen, Obadiah Lam, Adam Jacobson +1
cs.CVcs.LGcs.NEarXiv:1411.1509v12014Dynamic Curriculum Learning for Imbalanced Data Classification
Yiru Wang, Weihao Gan, Jie Yang +2
cs.CVarXiv:1901.06783v22019DetCLIP: Dictionary-Enriched Visual-Concept Paralleled Pre-training for Open-world Detection
Lewei Yao, Jianhua Han, Youpeng Wen +6
cs.CVarXiv:2209.09407v22022Lightweight Classification of IoT Malware based on Image Recognition
Jiawei Su, Danilo Vasconcellos Vargas, Sanjiva Prasad +3
cs.CRcs.CVarXiv:1802.03714v12018Not All Images are Worth 16x16 Words: Dynamic Transformers for Efficient Image Recognition
Yulin Wang, Rui Huang, Shiji Song +2
cs.CVcs.AIcs.LGarXiv:2105.15075v220213D Self-Supervised Methods for Medical Imaging
Aiham Taleb, Winfried Loetzsch, Noel Danz +4
cs.CVcs.LGeess.IVarXiv:2006.03829v32020Adversarial Camouflage: Hiding Physical-World Attacks with Natural Styles
Ranjie Duan, Xingjun Ma, Yisen Wang +3
cs.CVarXiv:2003.08757v22020Knowledge Distillation with the Reused Teacher Classifier
Defang Chen, Jian-Ping Mei, Hailin Zhang +3
cs.CVarXiv:2203.14001v12022Tensor-Train Recurrent Neural Networks for Video Classification
Yinchong Yang, Denis Krompass, Volker Tresp
cs.CVarXiv:1707.01786v12017PCBnet: A Dataset and Automatic Construction of SPICE Netlists from Schematic Images
Zhen Huang, Yuhao Gao, Yuzhi Liu +15
cs.CVcs.AIarXiv:2608.27923v12026Image-to-Markup Generation with Coarse-to-Fine Attention
Yuntian Deng, Anssi Kanervisto, Jeffrey Ling +1
cs.CVcs.CLcs.LGarXiv:1609.04938v22016UAV-Human: A Large Benchmark for Human Behavior Understanding with Unmanned Aerial Vehicles
Tianjiao Li, Jun Liu, Wei Zhang +3
cs.CVarXiv:2104.00946v420213D UX-Net: A Large Kernel Volumetric ConvNet Modernizing Hierarchical Transformer for Medical Image Segmentation
Ho Hin Lee, Shunxing Bao, Yuankai Huo +1
cs.CVcs.LGarXiv:2209.15076v42022Deep Kinematic Pose Regression
Xingyi Zhou, Xiao Sun, Wei Zhang +2
cs.CVarXiv:1609.05317v12016Tracking Anything with Decoupled Video Segmentation
Ho Kei Cheng, Seoung Wug Oh, Brian Price +2
cs.CVarXiv:2309.03903v12023Frequency Separation for Real-World Super-Resolution
Manuel Fritsche, Shuhang Gu, Radu Timofte
eess.IVcs.CVarXiv:1911.07850v12019An Enhanced Deep Feature Representation for Person Re-identification
Shangxuan Wu, Ying-Cong Chen, Xiang Li +3
cs.CVarXiv:1604.07807v22016Grid-GCN for Fast and Scalable Point Cloud Learning
Qiangeng Xu, Xudong Sun, Cho-Ying Wu +2
cs.CVcs.LGarXiv:1912.02984v52019SNARF: Differentiable Forward Skinning for Animating Non-Rigid Neural Implicit Shapes
Xu Chen, Yufeng Zheng, Michael J. Black +2
cs.CVarXiv:2104.03953v32021Zoom-in-Net: Deep Mining Lesions for Diabetic Retinopathy Detection
Zhe Wang, Yanxin Yin, Jianping Shi +3
cs.CVarXiv:1706.04372v12017Multi-Scale Geometric Consistency Guided Multi-View Stereo
Qingshan Xu, Wenbing Tao
cs.CVarXiv:1904.08103v12019Uncertainty Estimates and Multi-Hypotheses Networks for Optical Flow
Eddy Ilg, Özgün Çiçek, Silvio Galesso +4
cs.CVarXiv:1802.07095v42018You said that?
Joon Son Chung, Amir Jamaludin, Andrew Zisserman
cs.CVarXiv:1705.02966v22017Neural Head Avatars from Monocular RGB Videos
Philip-William Grassal, Malte Prinzler, Titus Leistner +3
cs.CVcs.GRarXiv:2112.01554v22021Overcoming Language Priors in Visual Question Answering with Adversarial Regularization
Sainandan Ramakrishnan, Aishwarya Agrawal, Stefan Lee
cs.CVarXiv:1810.03649v22018DeepGMR: Learning Latent Gaussian Mixture Models for Registration
Wentao Yuan, Ben Eckart, Kihwan Kim +3
cs.CVarXiv:2008.09088v12020Soft Threshold Weight Reparameterization for Learnable Sparsity
Aditya Kusupati, Vivek Ramanujan, Raghav Somani +4
cs.LGcs.CVstat.MLarXiv:2002.03231v92020From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation
Rit Gangopadhyay, Alex Wong
cs.CVcs.AIarXiv:2608.27860v12026MinkLoc3D: Point Cloud Based Large-Scale Place Recognition
Jacek Komorowski
cs.CVarXiv:2011.04530v12020DexGraspNet: A Large-Scale Robotic Dexterous Grasp Dataset for General Objects Based on Simulation
Ruicheng Wang, Jialiang Zhang, Jiayi Chen +4
cs.ROcs.CVarXiv:2210.02697v22022Sequential Attend, Infer, Repeat: Generative Modelling of Moving Objects
Adam R. Kosiorek, Hyunjik Kim, Ingmar Posner +1
cs.LGcs.CVstat.MLarXiv:1806.01794v22018Semi-Supervised Semantic Segmentation with Pixel-Level Contrastive Learning from a Class-wise Memory Bank
Inigo Alonso, Alberto Sabater, David Ferstl +2
cs.CVarXiv:2104.13415v32021EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions
Linrui Tian, Qi Wang, Bang Zhang +1
cs.CVarXiv:2402.17485v32024Pose-guided Feature Disentangling for Occluded Person Re-identification Based on Transformer
Tao Wang, Hong Liu, Pinhao Song +2
cs.CVarXiv:2112.02466v22021DeepLPF: Deep Local Parametric Filters for Image Enhancement
Sean Moran, Pierre Marza, Steven McDonagh +2
cs.CVarXiv:2003.13985v12020MoFusion: A Framework for Denoising-Diffusion-based Motion Synthesis
Rishabh Dabral, Muhammad Hamza Mughal, Vladislav Golyanik +1
cs.CVarXiv:2212.04495v32022Deep Learning for Unsupervised Anomaly Localization in Industrial Images: A Survey
Xian Tao, Xinyi Gong, Xin Zhang +2
cs.CVarXiv:2207.10298v12022MultiPoseNet: Fast Multi-Person Pose Estimation using Pose Residual Network
Muhammed Kocabas, Salih Karagoz, Emre Akbas
cs.CVarXiv:1807.04067v12018CARDINAL Predicts Cardiovascular Risk From Non-contrast Cardiac CT
Roy Gabriel, Nattakorn Kittisut, Jamshid Hassanpour +6
eess.IVcs.AIcs.CVarXiv:2608.27690v12026Self-supervised Multi-level Face Model Learning for Monocular Reconstruction at over 250 Hz
Ayush Tewari, Michael Zollhöfer, Pablo Garrido +4
cs.CVarXiv:1712.02859v22017VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
Ziyang Wang, Shoubin Yu, Elias Stengel-Eskin +4
cs.CVcs.AIcs.CLarXiv:2405.19209v32024Do Deep Neural Networks Learn Facial Action Units When Doing Expression Recognition?
Pooya Khorrami, Tom Le Paine, Thomas S. Huang
cs.CVcs.LGcs.NEarXiv:1510.02969v32015Railroad is not a Train: Saliency as Pseudo-pixel Supervision for Weakly Supervised Semantic Segmentation
Seungho Lee, Minhyun Lee, Jongwuk Lee +1
cs.CVarXiv:2105.08965v12021USE-Net: incorporating Squeeze-and-Excitation blocks into U-Net for prostate zonal segmentation of multi-institutional MRI datasets
Leonardo Rundo, Changhee Han, Yudai Nagano +12
cs.CVcs.LGarXiv:1904.08254v22019FVeinSyn: Synthetic Finger Vein Image Generator
Yifan Wang, Jie Gui, Adams Wai Kin Kong +6
cs.CVcs.AIarXiv:2608.27527v12026Segment and Track Anything
Yangming Cheng, Liulei Li, Yuanyou Xu +4
cs.CVarXiv:2305.06558v12023Taming 3DGS: High-Quality Radiance Fields with Limited Resources
Saswat Subhajyoti Mallick, Rahul Goel, Bernhard Kerbl +3
cs.CVcs.GRarXiv:2406.15643v12024Training Neural Networks with Local Error Signals
Arild Nøkland, Lars Hiller Eidnes
stat.MLcs.CVcs.LGarXiv:1901.06656v22019TI-POOLING: transformation-invariant pooling for feature learning in Convolutional Neural Networks
Dmitry Laptev, Nikolay Savinov, Joachim M. Buhmann +1
cs.CVarXiv:1604.06318v22016Curvature-Aware Radius Shrinkage for Adaptive Nearest Neighbor Classification
Alexandre L. M. Levada
cs.LGcs.AIcs.CVarXiv:2608.27634v12026Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIP
Qihang Yu, Ju He, Xueqing Deng +2
cs.CVarXiv:2308.02487v22023T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation
Kaiyi Huang, Chengqi Duan, Kaiyue Sun +3
cs.CVarXiv:2307.06350v32023Depth-Aware Pothole Detection Using YOLO and RT-DETR at the Edge
Md Monjurul Ahsan Prodhan, Md Nour Hossain
cs.CVcs.AIcs.LGarXiv:2608.27633v12026Segment Anything Is Not Always Perfect: An Investigation of SAM on Different Real-world Applications
Wei Ji, Jingjing Li, Qi Bi +3
cs.CVarXiv:2304.05750v32023Quanta Perception as Probabilistic Events
Varun Sundar, Pavan Thodima, Sacha Jungerman +1
cs.CVcs.AIarXiv:2608.27584v12026Woodpecker: Hallucination Correction for Multimodal Large Language Models
Shukang Yin, Chaoyou Fu, Sirui Zhao +7
cs.CVcs.AIcs.CLarXiv:2310.16045v22023On Translation Invariance in CNNs: Convolutional Layers can Exploit Absolute Spatial Location
Osman Semih Kayhan, Jan C. van Gemert
cs.CVcs.LGeess.IVarXiv:2003.07064v22020