Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

7,801 to 7,860 of 18,848

  1. GauHuman: Articulated Gaussian Splatting from Monocular Human Videos

    Shoukang Hu, Ziwei Liu

    cs.CVarXiv:2312.02973v12023
  2. $\mathbf{D^3}$: Deep Dual-Domain Based Fast Restoration of JPEG-Compressed Images

    Zhangyang Wang, Ding Liu, Shiyu Chang +3

    cs.CVcs.AIcs.LGarXiv:1601.04149v32016
  3. EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face Animation

    Ziqiao Peng, Haoyu Wu, Zhenbo Song +5

    cs.CVcs.SDeess.ASarXiv:2303.11089v22023
  4. Actions in the Eye: Dynamic Gaze Datasets and Learnt Saliency Models for Visual Recognition

    Stefan Mathe, Cristian Sminchisescu

    cs.CVarXiv:1312.7570v12013
  5. Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research

    Atousa Torabi, Christopher Pal, Hugo Larochelle +1

    cs.CVcs.AIarXiv:1503.01070v12015
  6. Audio-Visual Segmentation

    Jinxing Zhou, Jianyuan Wang, Jiayi Zhang +7

    cs.CVcs.MMcs.SDarXiv:2207.05042v32022
  7. Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels

    Pingchuan Ma, Alexandros Haliassos, Adriana Fernandez-Lopez +3

    cs.CVcs.SDeess.ASarXiv:2303.14307v32023
  8. Lesion Border Detection in Dermoscopy Images Using Ensembles of Thresholding Methods

    M. Emre Celebi, Quan Wen, Sae Hwang +2

    cs.CVarXiv:1312.7345v12013
  9. Event-based, 6-DOF Camera Tracking from Photometric Depth Maps

    Guillermo Gallego, Jon E. A. Lund, Elias Mueggler +3

    cs.CVcs.ROarXiv:1607.03468v22016
  10. Fast Fourier Color Constancy

    Jonathan T. Barron, Yun-Ta Tsai

    cs.CVarXiv:1611.07596v32016
  11. A Dataset for Improved RGBD-based Object Detection and Pose Estimation for Warehouse Pick-and-Place

    Colin Rennie, Rahul Shome, Kostas E. Bekris +1

    cs.CVcs.ROarXiv:1509.01277v22015
  12. GluonCV and GluonNLP: Deep Learning in Computer Vision and Natural Language Processing

    Jian Guo, He He, Tong He +13

    cs.LGcs.CLcs.CVarXiv:1907.04433v22019
  13. Deep Inertial Poser: Learning to Reconstruct Human Pose from Sparse Inertial Measurements in Real Time

    Yinghao Huang, Manuel Kaufmann, Emre Aksan +3

    cs.GRcs.CVarXiv:1810.04703v12018
  14. A Deep Metric for Multimodal Registration

    Martin Simonovsky, Benjamín Gutiérrez-Becker, Diana Mateus +2

    cs.CVcs.LGcs.NEarXiv:1609.05396v12016
  15. MotionLM: Multi-Agent Motion Forecasting as Language Modeling

    Ari Seff, Brian Cera, Dian Chen +6

    cs.CVcs.AIcs.LGarXiv:2309.16534v12023
  16. Single-Shot Object Detection with Enriched Semantics

    Zhishuai Zhang, Siyuan Qiao, Cihang Xie +3

    cs.CVarXiv:1712.00433v22017
  17. A Survey on 3D Skeleton-Based Action Recognition Using Learning Method

    Bin Ren, Mengyuan Liu, Runwei Ding +1

    cs.CVarXiv:2002.05907v22020
  18. Areas of Attention for Image Captioning

    Marco Pedersoli, Thomas Lucas, Cordelia Schmid +1

    cs.CVarXiv:1612.01033v22016
  19. Learning Dual Convolutional Neural Networks for Low-Level Vision

    Jinshan Pan, Sifei Liu, Deqing Sun +8

    cs.CVarXiv:1805.05020v12018
  20. Dr. Claw: An AI Scientist Workspace for Vibe Research

    Dingjie Song, Hanrong Zhang, Dawei Liu +10

    cs.AIcs.CLcs.CVarXiv:2609.00365v12026
  21. Attention Clusters: Purely Attention Based Local Feature Integration for Video Classification

    Xiang Long, Chuang Gan, Gerard de Melo +3

    cs.CVcs.LGarXiv:1711.09550v12017
  22. Audio-Driven Adversarial Defense for 3D Talking Face Generation with totally Visual Fidelity Preservation

    Rui-Qing Sun, Chen-Hao Cui, Hui-Yang Zhao +3

    cs.CVcs.MMarXiv:2608.30951v12026
  23. Unpaired Multi-modal Segmentation via Knowledge Distillation

    Qi Dou, Quande Liu, Pheng Ann Heng +1

    cs.CVeess.IVarXiv:2001.03111v12020
  24. FTU-Seek: Foundation Model-Guided Hard-Negative Learning for Sparse Functional Tissue Unit Segmentation

    Zonghao Liu, Lei Su, Jiguang Yu +4

    cs.CVmath.NAarXiv:2609.00704v12026
  25. Causal Attention for Vision-Language Tasks

    Xu Yang, Hanwang Zhang, Guojun Qi +1

    cs.CVarXiv:2103.03493v12021
  26. Cross-Domain Self-supervised Multi-task Feature Learning using Synthetic Imagery

    Zhongzheng Ren, Yong Jae Lee

    cs.CVarXiv:1711.09082v12017
  27. AET vs. AED: Unsupervised Representation Learning by Auto-Encoding Transformations rather than Data

    Liheng Zhang, Guo-Jun Qi, Liqiang Wang +1

    cs.CVarXiv:1901.04596v22019
  28. LSKNet: A Foundation Lightweight Backbone for Remote Sensing

    Yuxuan Li, Xiang Li, Yimian Dai +5

    cs.CVcs.LGarXiv:2403.11735v62024
  29. Deep Spatial Gradient and Temporal Depth Learning for Face Anti-spoofing

    Zezheng Wang, Zitong Yu, Chenxu Zhao +5

    cs.CVarXiv:2003.08061v12020
  30. Diversity in Faces

    Michele Merler, Nalini Ratha, Rogerio S. Feris +1

    cs.CVarXiv:1901.10436v62019
  31. Fully Automatic Wound Segmentation with Deep Convolutional Neural Networks

    Chuanbo Wang, DM Anisuzzaman, Victor Williamson +5

    eess.IVcs.CVarXiv:2010.05855v12020
  32. Visual object tracking performance measures revisited

    Luka Čehovin, Aleš Leonardis, Matej Kristan

    cs.CVarXiv:1502.05803v32015
  33. Towards Automatic Face-to-Face Translation

    Prajwal K R, Rudrabha Mukhopadhyay, Jerin Philip +3

    cs.CVcs.AIcs.LGarXiv:2003.00418v12020
  34. EPro-PnP: Generalized End-to-End Probabilistic Perspective-n-Points for Monocular Object Pose Estimation

    Hansheng Chen, Wei Tian, Pichao Wang +3

    cs.CVarXiv:2303.12787v32023
  35. FCN-rLSTM: Deep Spatio-Temporal Neural Networks for Vehicle Counting in City Cameras

    Shanghang Zhang, Guanhang Wu, João P. Costeira +1

    cs.CVarXiv:1707.09476v22017
  36. Interpreting CLIP's Image Representation via Text-Based Decomposition

    Yossi Gandelsman, Alexei A. Efros, Jacob Steinhardt

    cs.CVcs.AIarXiv:2310.05916v42023
  37. Equalization Loss v2: A New Gradient Balance Approach for Long-tailed Object Detection

    Jingru Tan, Xin Lu, Gang Zhang +2

    cs.CVcs.LGarXiv:2012.08548v22020
  38. Generalized Singular Value Thresholding

    Canyi Lu, Changbo Zhu, Chunyan Xu +2

    cs.CVcs.LGmath.NAarXiv:1412.2231v22014
  39. Deep Geometric Prior for Surface Reconstruction

    Francis Williams, Teseo Schneider, Claudio Silva +3

    cs.CVcs.GRcs.LGarXiv:1811.10943v22018
  40. PVANET: Deep but Lightweight Neural Networks for Real-time Object Detection

    Kye-Hyeon Kim, Sanghoon Hong, Byungseok Roh +2

    cs.CVarXiv:1608.08021v32016
  41. Mask-Guided Attention Network for Occluded Pedestrian Detection

    Yanwei Pang, Jin Xie, Muhammad Haris Khan +3

    cs.CVarXiv:1910.06160v22019
  42. OctSqueeze: Octree-Structured Entropy Model for LiDAR Compression

    Lila Huang, Shenlong Wang, Kelvin Wong +2

    eess.IVcs.CVarXiv:2005.07178v22020
  43. Unsupervised Learning of Object Keypoints for Perception and Control

    Tejas Kulkarni, Ankush Gupta, Catalin Ionescu +4

    cs.CVcs.LGarXiv:1906.11883v22019
  44. Blind2Unblind: Self-Supervised Image Denoising with Visible Blind Spots

    Zejin Wang, Jiazheng Liu, Guoqing Li +1

    eess.IVcs.CVarXiv:2203.06967v32022
  45. Co-Mixup: Saliency Guided Joint Mixup with Supermodular Diversity

    Jang-Hyun Kim, Wonho Choo, Hosan Jeong +1

    cs.LGcs.AIcs.CVarXiv:2102.03065v12021
  46. Unsupervised 3D Pose Estimation with Geometric Self-Supervision

    Ching-Hang Chen, Ambrish Tyagi, Amit Agrawal +4

    cs.CVarXiv:1904.04812v12019
  47. Making LLaMA SEE and Draw with SEED Tokenizer

    Yuying Ge, Sijie Zhao, Ziyun Zeng +4

    cs.CVarXiv:2310.01218v12023
  48. A Survey on Content-Aware Video Analysis for Sports

    Huang-Chia Shih

    cs.CVcs.MMarXiv:1703.01170v12017
  49. Confidence-Aware Ensemble and Long-Word Refinement for Artistic Text Recognition

    Lucas A. Dias, Henrique A. Schulz, Rafaela de Miranda +3

    cs.CVarXiv:2608.29970v12026
  50. On the Role of MRI Sequences in Cross-Dataset Generalization for Brain Tumor Segmentation

    Henrique Zan Grande, João G. Pitol, Lucas B. Schuck +3

    cs.CVarXiv:2608.29944v12026
  51. Dynamic Instance Normalization for Arbitrary Style Transfer

    Yongcheng Jing, Xiao Liu, Yukang Ding +4

    cs.CVarXiv:1911.06953v12019
  52. An end-to-end generative framework for video segmentation and recognition

    Hilde Kuehne, Juergen Gall, Thomas Serre

    cs.CVarXiv:1509.01947v22015
  53. AMT: All-Pairs Multi-Field Transforms for Efficient Frame Interpolation

    Zhen Li, Zuo-Liang Zhu, Ling-Hao Han +3

    cs.CVarXiv:2304.09790v12023
  54. CRAFT: Concept Recursive Activation FacTorization for Explainability

    Thomas Fel, Agustin Picard, Louis Bethune +5

    cs.CVcs.AIarXiv:2211.10154v22022
  55. RED: Reinforced Encoder-Decoder Networks for Action Anticipation

    Jiyang Gao, Zhenheng Yang, Ram Nevatia

    cs.CVarXiv:1707.04818v12017
  56. FuncRoom-Agent: Sequential Feed-Forward 3D Functional Indoor Scene Generation

    Hao Feng, Zhi Zuo, MingJian Liang +6

    cs.CVarXiv:2608.29519v12026
  57. When will you do what? - Anticipating Temporal Occurrences of Activities

    Yazan Abu Farha, Alexander Richard, Juergen Gall

    cs.CVarXiv:1804.00892v12018
  58. Seeing Through Extreme Visual Sparsity: Surface Understanding from a Single Random Visual Patch

    Sindhuja Penchala, Sudip Mittal, Noorbakhsh Amiri Golilarz

    cs.CVarXiv:2608.29475v12026
  59. CERF: Communication-Efficient and Retraining-Free Collaborative Perception

    Jiuwu Hao, Ziyi Ni, Liguo Sun +5

    cs.CVarXiv:2609.00951v12026
  60. SG-Net: Spatial Granularity Network for One-Stage Video Instance Segmentation

    Dongfang Liu, Yiming Cui, Wenbo Tan +1

    cs.CVarXiv:2103.10284v22021