Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
7,741 to 7,800 of 18,841
TransVPR: Transformer-based place recognition with multi-level attention aggregation
Ruotong Wang, Yanqing Shen, Weiliang Zuo +2
cs.CVarXiv:2201.02001v42022Learning Compact Geometric Features
Marc Khoury, Qian-Yi Zhou, Vladlen Koltun
cs.CVcs.GRcs.LGarXiv:1709.05056v12017Occluded Video Instance Segmentation: A Benchmark
Jiyang Qi, Yan Gao, Yao Hu +7
cs.CVarXiv:2102.01558v62021Sparse Tensor-based Multiscale Representation for Point Cloud Geometry Compression
Jianqiang Wang, Dandan Ding, Zhu Li +3
cs.CVeess.IVarXiv:2111.10633v22021Entroformer: A Transformer-based Entropy Model for Learned Image Compression
Yichen Qian, Ming Lin, Xiuyu Sun +2
eess.IVcs.CVarXiv:2202.05492v22022Finetune like you pretrain: Improved finetuning of zero-shot vision models
Sachin Goyal, Ananya Kumar, Sankalp Garg +2
cs.CVcs.LGarXiv:2212.00638v12022SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models
Junchao Huang, Guian Fang, Shengju Qian +15
cs.CVarXiv:2609.02886v12026Cross-Modality Deep Feature Learning for Brain Tumor Segmentation
Dingwen Zhang, Guohai Huang, Qiang Zhang +3
eess.IVcs.CVarXiv:2201.02356v12022VmambaIR: Visual State Space Model for Image Restoration
Yuan Shi, Bin Xia, Xiaoyu Jin +5
cs.CVarXiv:2403.11423v12024Deep Co-Training for Semi-Supervised Image Segmentation
Jizong Peng, Guillermo Estrada, Marco Pedersoli +1
cs.CVarXiv:1903.11233v32019Random Vector Functional Link Neural Network based Ensemble Deep Learning
Rakesh Katuwal, P. N. Suganthan, M. Tanveer
cs.CVarXiv:1907.00350v12019Inconsistency-aware Uncertainty Estimation for Semi-supervised Medical Image Segmentation
Yinghuan Shi, Jian Zhang, Tong Ling +5
cs.CVarXiv:2110.08762v12021Human Skin Detection Using RGB, HSV and YCbCr Color Models
S. Kolkur, D. Kalbande, P. Shimpi +2
cs.CVq-bio.OTarXiv:1708.02694v12017On the Integration of Optical Flow and Action Recognition
Laura Sevilla-Lara, Yiyi Liao, Fatma Guney +3
cs.CVarXiv:1712.08416v12017Using convolutional networks and satellite imagery to identify patterns in urban environments at a large scale
Adrian Albert, Jasleen Kaur, Marta Gonzalez
cs.CVarXiv:1704.02965v22017Can Computers Create Art?
Aaron Hertzmann
cs.AIcs.CVcs.GRarXiv:1801.04486v62018Online Deep Clustering for Unsupervised Representation Learning
Xiaohang Zhan, Jiahao Xie, Ziwei Liu +2
cs.CVcs.LGarXiv:2006.10645v12020CRF Learning with CNN Features for Image Segmentation
Fayao Liu, Guosheng Lin, Chunhua Shen
cs.CVarXiv:1503.08263v12015A Comprehensive Overview of Biometric Fusion
Maneet Singh, Richa Singh, Arun Ross
cs.CVarXiv:1902.02919v12019Semi-Supervised and Unsupervised Deep Visual Learning: A Survey
Yanbei Chen, Massimiliano Mancini, Xiatian Zhu +1
cs.CVcs.AIcs.LGarXiv:2208.11296v12022GeoRay: Gauge-Aware Feed-Forward Satellite 3D Reconstruction in the Geodetic Frame
Zhe Dong, Wanqing Wu, Yuzhe Sun +5
cs.CVarXiv:2608.29680v12026RePair: Turning Retrieval Failures into Counterfactual Hard Pairs
Siyi Liu, Xiaorong Zhu, Enjun Du +6
cs.IRcs.CVarXiv:2608.29604v12026Deep Learning for Sensor-based Activity Recognition: A Survey
Jindong Wang, Yiqiang Chen, Shuji Hao +2
cs.CVcs.AIcs.LGarXiv:1707.03502v22017Bridging Mode Connectivity in Loss Landscapes and Adversarial Robustness
Pu Zhao, Pin-Yu Chen, Payel Das +2
cs.LGcs.CVstat.MLarXiv:2005.00060v22020Multimodal Sentiment Analysis: Addressing Key Issues and Setting up the Baselines
Soujanya Poria, Navonil Majumder, Devamanyu Hazarika +3
cs.CLcs.CVcs.IRarXiv:1803.07427v22018Towards End-to-end Text Spotting with Convolutional Recurrent Neural Networks
Hui Li, Peng Wang, Chunhua Shen
cs.CVarXiv:1707.03985v12017DropConnect Is Effective in Modeling Uncertainty of Bayesian Deep Networks
Aryan Mobiny, Hien V. Nguyen, Supratik Moulik +2
cs.LGcs.AIcs.CVarXiv:1906.04569v12019Glyce: Glyph-vectors for Chinese Character Representations
Yuxian Meng, Wei Wu, Fei Wang +8
cs.CLcs.AIcs.CVarXiv:1901.10125v72019PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System
Chenxia Li, Weiwei Liu, Ruoyu Guo +9
cs.CVarXiv:2206.03001v22022Hierarchical Point-Edge Interaction Network for Point Cloud Semantic Segmentation
Li Jiang, Hengshuang Zhao, Shu Liu +3
cs.CVarXiv:1909.10469v12019SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions
Zirong Chen, Fuda Ye, Kuan Zhang +7
cs.CVcs.IRarXiv:2608.29607v12026RegNet: Self-Regulated Network for Image Classification
Jing Xu, Yu Pan, Xinglin Pan +3
eess.IVcs.CVarXiv:2101.00590v12021MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Haozhe Zhao, Zefan Cai, Shuzheng Si +7
cs.CLcs.AIcs.CVarXiv:2309.07915v32023A MultiPath Network for Object Detection
Sergey Zagoruyko, Adam Lerer, Tsung-Yi Lin +4
cs.CVarXiv:1604.02135v22016Deep Feature Pyramid Reconfiguration for Object Detection
Tao Kong, Fuchun Sun, Wenbing Huang +1
cs.CVarXiv:1808.07993v12018Tensor Ring Decomposition with Rank Minimization on Latent Space: An Efficient Approach for Tensor Completion
Longhao Yuan, Chao Li, Danilo Mandic +2
cs.LGcs.CVstat.MLarXiv:1809.02288v22018Kinematic 3D Object Detection in Monocular Video
Garrick Brazil, Gerard Pons-Moll, Xiaoming Liu +1
cs.CVarXiv:2007.09548v12020Graph Embedded Pose Clustering for Anomaly Detection
Amir Markovitz, Gilad Sharir, Itamar Friedman +2
cs.CVarXiv:1912.11850v22019Towards Adversarial Attack on Vision-Language Pre-training Models
Jiaming Zhang, Qi Yi, Jitao Sang
cs.LGcs.CLcs.CVarXiv:2206.09391v22022Visual Attention Faithfulness in Vision-Language Models is Heterogeneous
Xurui Song, Weishi Wang, Zhongqi Yue +5
cs.CVcs.AIarXiv:2609.00830v12026Coherent Reconstruction of Multiple Humans from a Single Image
Wen Jiang, Nikos Kolotouros, Georgios Pavlakos +2
cs.CVarXiv:2006.08586v12020Capturing and Inferring Dense Full-Body Human-Scene Contact
Chun-Hao P. Huang, Hongwei Yi, Markus Höschle +5
cs.CVarXiv:2206.09553v12022Dense Feature Aggregation and Pruning for RGBT Tracking
Yabin Zhu, Chenglong Li, Bin Luo +2
cs.CVarXiv:1907.10451v12019Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
Yiyang Zhou, Chenhang Cui, Rafael Rafailov +2
cs.LGcs.CLcs.CVarXiv:2402.11411v12024Boosting Monocular Depth Estimation Models to High-Resolution via Content-Adaptive Multi-Resolution Merging
S. Mahdi H. Miangoleh, Sebastian Dille, Long Mai +2
cs.CVarXiv:2105.14021v12021Learning Normal Dynamics in Videos with Meta Prototype Network
Hui Lv, Chen Chen, Zhen Cui +3
cs.CVarXiv:2104.06689v22021GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis
Zhenhui Ye, Ziyue Jiang, Yi Ren +3
cs.CVarXiv:2301.13430v12023RANet: Ranking Attention Network for Fast Video Object Segmentation
Ziqin Wang, Jun Xu, Li Liu +2
cs.CVarXiv:1908.06647v42019Explaining Classifiers with Causal Concept Effect (CaCE)
Yash Goyal, Amir Feder, Uri Shalit +1
cs.LGcs.CVstat.MLarXiv:1907.07165v22019Datamodels: Predicting Predictions from Training Data
Andrew Ilyas, Sung Min Park, Logan Engstrom +2
stat.MLcs.CVcs.LGarXiv:2202.00622v12022Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion
Xunpeng Yi, Han Xu, Hao Zhang +2
cs.CVarXiv:2403.16387v12024CascadeTabNet: An approach for end to end table detection and structure recognition from image-based documents
Devashish Prasad, Ayan Gadpal, Kshitij Kapadni +2
cs.CVarXiv:2004.12629v22020LLM-Driven Autonomous Vehicles Inherit Human Driver Biases in Pedestrian Yielding: Results and Implications From A New Benchmark
Irem Yoldas, Martim Brandão, Jie Zhang +1
cs.AIcs.CLcs.CVarXiv:2609.00192v12026GauHuman: Articulated Gaussian Splatting from Monocular Human Videos
Shoukang Hu, Ziwei Liu
cs.CVarXiv:2312.02973v12023$\mathbf{D^3}$: Deep Dual-Domain Based Fast Restoration of JPEG-Compressed Images
Zhangyang Wang, Ding Liu, Shiyu Chang +3
cs.CVcs.AIcs.LGarXiv:1601.04149v32016EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face Animation
Ziqiao Peng, Haoyu Wu, Zhenbo Song +5
cs.CVcs.SDeess.ASarXiv:2303.11089v22023Actions in the Eye: Dynamic Gaze Datasets and Learnt Saliency Models for Visual Recognition
Stefan Mathe, Cristian Sminchisescu
cs.CVarXiv:1312.7570v12013Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research
Atousa Torabi, Christopher Pal, Hugo Larochelle +1
cs.CVcs.AIarXiv:1503.01070v12015Audio-Visual Segmentation
Jinxing Zhou, Jianyuan Wang, Jiayi Zhang +7
cs.CVcs.MMcs.SDarXiv:2207.05042v32022Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels
Pingchuan Ma, Alexandros Haliassos, Adriana Fernandez-Lopez +3
cs.CVcs.SDeess.ASarXiv:2303.14307v32023