Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

7,741 to 7,800 of 18,841

  1. TransVPR: Transformer-based place recognition with multi-level attention aggregation

    Ruotong Wang, Yanqing Shen, Weiliang Zuo +2

    cs.CVarXiv:2201.02001v42022
  2. Learning Compact Geometric Features

    Marc Khoury, Qian-Yi Zhou, Vladlen Koltun

    cs.CVcs.GRcs.LGarXiv:1709.05056v12017
  3. Occluded Video Instance Segmentation: A Benchmark

    Jiyang Qi, Yan Gao, Yao Hu +7

    cs.CVarXiv:2102.01558v62021
  4. Sparse Tensor-based Multiscale Representation for Point Cloud Geometry Compression

    Jianqiang Wang, Dandan Ding, Zhu Li +3

    cs.CVeess.IVarXiv:2111.10633v22021
  5. Entroformer: A Transformer-based Entropy Model for Learned Image Compression

    Yichen Qian, Ming Lin, Xiuyu Sun +2

    eess.IVcs.CVarXiv:2202.05492v22022
  6. Finetune like you pretrain: Improved finetuning of zero-shot vision models

    Sachin Goyal, Ananya Kumar, Sankalp Garg +2

    cs.CVcs.LGarXiv:2212.00638v12022
  7. SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

    Junchao Huang, Guian Fang, Shengju Qian +15

    cs.CVarXiv:2609.02886v12026
  8. Cross-Modality Deep Feature Learning for Brain Tumor Segmentation

    Dingwen Zhang, Guohai Huang, Qiang Zhang +3

    eess.IVcs.CVarXiv:2201.02356v12022
  9. VmambaIR: Visual State Space Model for Image Restoration

    Yuan Shi, Bin Xia, Xiaoyu Jin +5

    cs.CVarXiv:2403.11423v12024
  10. Deep Co-Training for Semi-Supervised Image Segmentation

    Jizong Peng, Guillermo Estrada, Marco Pedersoli +1

    cs.CVarXiv:1903.11233v32019
  11. Random Vector Functional Link Neural Network based Ensemble Deep Learning

    Rakesh Katuwal, P. N. Suganthan, M. Tanveer

    cs.CVarXiv:1907.00350v12019
  12. Inconsistency-aware Uncertainty Estimation for Semi-supervised Medical Image Segmentation

    Yinghuan Shi, Jian Zhang, Tong Ling +5

    cs.CVarXiv:2110.08762v12021
  13. Human Skin Detection Using RGB, HSV and YCbCr Color Models

    S. Kolkur, D. Kalbande, P. Shimpi +2

    cs.CVq-bio.OTarXiv:1708.02694v12017
  14. On the Integration of Optical Flow and Action Recognition

    Laura Sevilla-Lara, Yiyi Liao, Fatma Guney +3

    cs.CVarXiv:1712.08416v12017
  15. Using convolutional networks and satellite imagery to identify patterns in urban environments at a large scale

    Adrian Albert, Jasleen Kaur, Marta Gonzalez

    cs.CVarXiv:1704.02965v22017
  16. Can Computers Create Art?

    Aaron Hertzmann

    cs.AIcs.CVcs.GRarXiv:1801.04486v62018
  17. Online Deep Clustering for Unsupervised Representation Learning

    Xiaohang Zhan, Jiahao Xie, Ziwei Liu +2

    cs.CVcs.LGarXiv:2006.10645v12020
  18. CRF Learning with CNN Features for Image Segmentation

    Fayao Liu, Guosheng Lin, Chunhua Shen

    cs.CVarXiv:1503.08263v12015
  19. A Comprehensive Overview of Biometric Fusion

    Maneet Singh, Richa Singh, Arun Ross

    cs.CVarXiv:1902.02919v12019
  20. Semi-Supervised and Unsupervised Deep Visual Learning: A Survey

    Yanbei Chen, Massimiliano Mancini, Xiatian Zhu +1

    cs.CVcs.AIcs.LGarXiv:2208.11296v12022
  21. GeoRay: Gauge-Aware Feed-Forward Satellite 3D Reconstruction in the Geodetic Frame

    Zhe Dong, Wanqing Wu, Yuzhe Sun +5

    cs.CVarXiv:2608.29680v12026
  22. RePair: Turning Retrieval Failures into Counterfactual Hard Pairs

    Siyi Liu, Xiaorong Zhu, Enjun Du +6

    cs.IRcs.CVarXiv:2608.29604v12026
  23. Deep Learning for Sensor-based Activity Recognition: A Survey

    Jindong Wang, Yiqiang Chen, Shuji Hao +2

    cs.CVcs.AIcs.LGarXiv:1707.03502v22017
  24. Bridging Mode Connectivity in Loss Landscapes and Adversarial Robustness

    Pu Zhao, Pin-Yu Chen, Payel Das +2

    cs.LGcs.CVstat.MLarXiv:2005.00060v22020
  25. Multimodal Sentiment Analysis: Addressing Key Issues and Setting up the Baselines

    Soujanya Poria, Navonil Majumder, Devamanyu Hazarika +3

    cs.CLcs.CVcs.IRarXiv:1803.07427v22018
  26. Towards End-to-end Text Spotting with Convolutional Recurrent Neural Networks

    Hui Li, Peng Wang, Chunhua Shen

    cs.CVarXiv:1707.03985v12017
  27. DropConnect Is Effective in Modeling Uncertainty of Bayesian Deep Networks

    Aryan Mobiny, Hien V. Nguyen, Supratik Moulik +2

    cs.LGcs.AIcs.CVarXiv:1906.04569v12019
  28. Glyce: Glyph-vectors for Chinese Character Representations

    Yuxian Meng, Wei Wu, Fei Wang +8

    cs.CLcs.AIcs.CVarXiv:1901.10125v72019
  29. PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System

    Chenxia Li, Weiwei Liu, Ruoyu Guo +9

    cs.CVarXiv:2206.03001v22022
  30. Hierarchical Point-Edge Interaction Network for Point Cloud Semantic Segmentation

    Li Jiang, Hengshuang Zhao, Shu Liu +3

    cs.CVarXiv:1909.10469v12019
  31. SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions

    Zirong Chen, Fuda Ye, Kuan Zhang +7

    cs.CVcs.IRarXiv:2608.29607v12026
  32. RegNet: Self-Regulated Network for Image Classification

    Jing Xu, Yu Pan, Xinglin Pan +3

    eess.IVcs.CVarXiv:2101.00590v12021
  33. MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

    Haozhe Zhao, Zefan Cai, Shuzheng Si +7

    cs.CLcs.AIcs.CVarXiv:2309.07915v32023
  34. A MultiPath Network for Object Detection

    Sergey Zagoruyko, Adam Lerer, Tsung-Yi Lin +4

    cs.CVarXiv:1604.02135v22016
  35. Deep Feature Pyramid Reconfiguration for Object Detection

    Tao Kong, Fuchun Sun, Wenbing Huang +1

    cs.CVarXiv:1808.07993v12018
  36. Tensor Ring Decomposition with Rank Minimization on Latent Space: An Efficient Approach for Tensor Completion

    Longhao Yuan, Chao Li, Danilo Mandic +2

    cs.LGcs.CVstat.MLarXiv:1809.02288v22018
  37. Kinematic 3D Object Detection in Monocular Video

    Garrick Brazil, Gerard Pons-Moll, Xiaoming Liu +1

    cs.CVarXiv:2007.09548v12020
  38. Graph Embedded Pose Clustering for Anomaly Detection

    Amir Markovitz, Gilad Sharir, Itamar Friedman +2

    cs.CVarXiv:1912.11850v22019
  39. Towards Adversarial Attack on Vision-Language Pre-training Models

    Jiaming Zhang, Qi Yi, Jitao Sang

    cs.LGcs.CLcs.CVarXiv:2206.09391v22022
  40. Visual Attention Faithfulness in Vision-Language Models is Heterogeneous

    Xurui Song, Weishi Wang, Zhongqi Yue +5

    cs.CVcs.AIarXiv:2609.00830v12026
  41. Coherent Reconstruction of Multiple Humans from a Single Image

    Wen Jiang, Nikos Kolotouros, Georgios Pavlakos +2

    cs.CVarXiv:2006.08586v12020
  42. Capturing and Inferring Dense Full-Body Human-Scene Contact

    Chun-Hao P. Huang, Hongwei Yi, Markus Höschle +5

    cs.CVarXiv:2206.09553v12022
  43. Dense Feature Aggregation and Pruning for RGBT Tracking

    Yabin Zhu, Chenglong Li, Bin Luo +2

    cs.CVarXiv:1907.10451v12019
  44. Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

    Yiyang Zhou, Chenhang Cui, Rafael Rafailov +2

    cs.LGcs.CLcs.CVarXiv:2402.11411v12024
  45. Boosting Monocular Depth Estimation Models to High-Resolution via Content-Adaptive Multi-Resolution Merging

    S. Mahdi H. Miangoleh, Sebastian Dille, Long Mai +2

    cs.CVarXiv:2105.14021v12021
  46. Learning Normal Dynamics in Videos with Meta Prototype Network

    Hui Lv, Chen Chen, Zhen Cui +3

    cs.CVarXiv:2104.06689v22021
  47. GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis

    Zhenhui Ye, Ziyue Jiang, Yi Ren +3

    cs.CVarXiv:2301.13430v12023
  48. RANet: Ranking Attention Network for Fast Video Object Segmentation

    Ziqin Wang, Jun Xu, Li Liu +2

    cs.CVarXiv:1908.06647v42019
  49. Explaining Classifiers with Causal Concept Effect (CaCE)

    Yash Goyal, Amir Feder, Uri Shalit +1

    cs.LGcs.CVstat.MLarXiv:1907.07165v22019
  50. Datamodels: Predicting Predictions from Training Data

    Andrew Ilyas, Sung Min Park, Logan Engstrom +2

    stat.MLcs.CVcs.LGarXiv:2202.00622v12022
  51. Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion

    Xunpeng Yi, Han Xu, Hao Zhang +2

    cs.CVarXiv:2403.16387v12024
  52. CascadeTabNet: An approach for end to end table detection and structure recognition from image-based documents

    Devashish Prasad, Ayan Gadpal, Kshitij Kapadni +2

    cs.CVarXiv:2004.12629v22020
  53. LLM-Driven Autonomous Vehicles Inherit Human Driver Biases in Pedestrian Yielding: Results and Implications From A New Benchmark

    Irem Yoldas, Martim Brandão, Jie Zhang +1

    cs.AIcs.CLcs.CVarXiv:2609.00192v12026
  54. GauHuman: Articulated Gaussian Splatting from Monocular Human Videos

    Shoukang Hu, Ziwei Liu

    cs.CVarXiv:2312.02973v12023
  55. $\mathbf{D^3}$: Deep Dual-Domain Based Fast Restoration of JPEG-Compressed Images

    Zhangyang Wang, Ding Liu, Shiyu Chang +3

    cs.CVcs.AIcs.LGarXiv:1601.04149v32016
  56. EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face Animation

    Ziqiao Peng, Haoyu Wu, Zhenbo Song +5

    cs.CVcs.SDeess.ASarXiv:2303.11089v22023
  57. Actions in the Eye: Dynamic Gaze Datasets and Learnt Saliency Models for Visual Recognition

    Stefan Mathe, Cristian Sminchisescu

    cs.CVarXiv:1312.7570v12013
  58. Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research

    Atousa Torabi, Christopher Pal, Hugo Larochelle +1

    cs.CVcs.AIarXiv:1503.01070v12015
  59. Audio-Visual Segmentation

    Jinxing Zhou, Jianyuan Wang, Jiayi Zhang +7

    cs.CVcs.MMcs.SDarXiv:2207.05042v32022
  60. Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels

    Pingchuan Ma, Alexandros Haliassos, Adriana Fernandez-Lopez +3

    cs.CVcs.SDeess.ASarXiv:2303.14307v32023