Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
7,801 to 7,860 of 18,848
GauHuman: Articulated Gaussian Splatting from Monocular Human Videos
Shoukang Hu, Ziwei Liu
cs.CVarXiv:2312.02973v12023$\mathbf{D^3}$: Deep Dual-Domain Based Fast Restoration of JPEG-Compressed Images
Zhangyang Wang, Ding Liu, Shiyu Chang +3
cs.CVcs.AIcs.LGarXiv:1601.04149v32016EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face Animation
Ziqiao Peng, Haoyu Wu, Zhenbo Song +5
cs.CVcs.SDeess.ASarXiv:2303.11089v22023Actions in the Eye: Dynamic Gaze Datasets and Learnt Saliency Models for Visual Recognition
Stefan Mathe, Cristian Sminchisescu
cs.CVarXiv:1312.7570v12013Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research
Atousa Torabi, Christopher Pal, Hugo Larochelle +1
cs.CVcs.AIarXiv:1503.01070v12015Audio-Visual Segmentation
Jinxing Zhou, Jianyuan Wang, Jiayi Zhang +7
cs.CVcs.MMcs.SDarXiv:2207.05042v32022Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels
Pingchuan Ma, Alexandros Haliassos, Adriana Fernandez-Lopez +3
cs.CVcs.SDeess.ASarXiv:2303.14307v32023Lesion Border Detection in Dermoscopy Images Using Ensembles of Thresholding Methods
M. Emre Celebi, Quan Wen, Sae Hwang +2
cs.CVarXiv:1312.7345v12013Event-based, 6-DOF Camera Tracking from Photometric Depth Maps
Guillermo Gallego, Jon E. A. Lund, Elias Mueggler +3
cs.CVcs.ROarXiv:1607.03468v22016Fast Fourier Color Constancy
Jonathan T. Barron, Yun-Ta Tsai
cs.CVarXiv:1611.07596v32016A Dataset for Improved RGBD-based Object Detection and Pose Estimation for Warehouse Pick-and-Place
Colin Rennie, Rahul Shome, Kostas E. Bekris +1
cs.CVcs.ROarXiv:1509.01277v22015GluonCV and GluonNLP: Deep Learning in Computer Vision and Natural Language Processing
Jian Guo, He He, Tong He +13
cs.LGcs.CLcs.CVarXiv:1907.04433v22019Deep Inertial Poser: Learning to Reconstruct Human Pose from Sparse Inertial Measurements in Real Time
Yinghao Huang, Manuel Kaufmann, Emre Aksan +3
cs.GRcs.CVarXiv:1810.04703v12018A Deep Metric for Multimodal Registration
Martin Simonovsky, Benjamín Gutiérrez-Becker, Diana Mateus +2
cs.CVcs.LGcs.NEarXiv:1609.05396v12016MotionLM: Multi-Agent Motion Forecasting as Language Modeling
Ari Seff, Brian Cera, Dian Chen +6
cs.CVcs.AIcs.LGarXiv:2309.16534v12023Single-Shot Object Detection with Enriched Semantics
Zhishuai Zhang, Siyuan Qiao, Cihang Xie +3
cs.CVarXiv:1712.00433v22017A Survey on 3D Skeleton-Based Action Recognition Using Learning Method
Bin Ren, Mengyuan Liu, Runwei Ding +1
cs.CVarXiv:2002.05907v22020Areas of Attention for Image Captioning
Marco Pedersoli, Thomas Lucas, Cordelia Schmid +1
cs.CVarXiv:1612.01033v22016Learning Dual Convolutional Neural Networks for Low-Level Vision
Jinshan Pan, Sifei Liu, Deqing Sun +8
cs.CVarXiv:1805.05020v12018Dr. Claw: An AI Scientist Workspace for Vibe Research
Dingjie Song, Hanrong Zhang, Dawei Liu +10
cs.AIcs.CLcs.CVarXiv:2609.00365v12026Attention Clusters: Purely Attention Based Local Feature Integration for Video Classification
Xiang Long, Chuang Gan, Gerard de Melo +3
cs.CVcs.LGarXiv:1711.09550v12017Audio-Driven Adversarial Defense for 3D Talking Face Generation with totally Visual Fidelity Preservation
Rui-Qing Sun, Chen-Hao Cui, Hui-Yang Zhao +3
cs.CVcs.MMarXiv:2608.30951v12026Unpaired Multi-modal Segmentation via Knowledge Distillation
Qi Dou, Quande Liu, Pheng Ann Heng +1
cs.CVeess.IVarXiv:2001.03111v12020FTU-Seek: Foundation Model-Guided Hard-Negative Learning for Sparse Functional Tissue Unit Segmentation
Zonghao Liu, Lei Su, Jiguang Yu +4
cs.CVmath.NAarXiv:2609.00704v12026Causal Attention for Vision-Language Tasks
Xu Yang, Hanwang Zhang, Guojun Qi +1
cs.CVarXiv:2103.03493v12021Cross-Domain Self-supervised Multi-task Feature Learning using Synthetic Imagery
Zhongzheng Ren, Yong Jae Lee
cs.CVarXiv:1711.09082v12017AET vs. AED: Unsupervised Representation Learning by Auto-Encoding Transformations rather than Data
Liheng Zhang, Guo-Jun Qi, Liqiang Wang +1
cs.CVarXiv:1901.04596v22019LSKNet: A Foundation Lightweight Backbone for Remote Sensing
Yuxuan Li, Xiang Li, Yimian Dai +5
cs.CVcs.LGarXiv:2403.11735v62024Deep Spatial Gradient and Temporal Depth Learning for Face Anti-spoofing
Zezheng Wang, Zitong Yu, Chenxu Zhao +5
cs.CVarXiv:2003.08061v12020Diversity in Faces
Michele Merler, Nalini Ratha, Rogerio S. Feris +1
cs.CVarXiv:1901.10436v62019Fully Automatic Wound Segmentation with Deep Convolutional Neural Networks
Chuanbo Wang, DM Anisuzzaman, Victor Williamson +5
eess.IVcs.CVarXiv:2010.05855v12020Visual object tracking performance measures revisited
Luka Čehovin, Aleš Leonardis, Matej Kristan
cs.CVarXiv:1502.05803v32015Towards Automatic Face-to-Face Translation
Prajwal K R, Rudrabha Mukhopadhyay, Jerin Philip +3
cs.CVcs.AIcs.LGarXiv:2003.00418v12020EPro-PnP: Generalized End-to-End Probabilistic Perspective-n-Points for Monocular Object Pose Estimation
Hansheng Chen, Wei Tian, Pichao Wang +3
cs.CVarXiv:2303.12787v32023FCN-rLSTM: Deep Spatio-Temporal Neural Networks for Vehicle Counting in City Cameras
Shanghang Zhang, Guanhang Wu, João P. Costeira +1
cs.CVarXiv:1707.09476v22017Interpreting CLIP's Image Representation via Text-Based Decomposition
Yossi Gandelsman, Alexei A. Efros, Jacob Steinhardt
cs.CVcs.AIarXiv:2310.05916v42023Equalization Loss v2: A New Gradient Balance Approach for Long-tailed Object Detection
Jingru Tan, Xin Lu, Gang Zhang +2
cs.CVcs.LGarXiv:2012.08548v22020Generalized Singular Value Thresholding
Canyi Lu, Changbo Zhu, Chunyan Xu +2
cs.CVcs.LGmath.NAarXiv:1412.2231v22014Deep Geometric Prior for Surface Reconstruction
Francis Williams, Teseo Schneider, Claudio Silva +3
cs.CVcs.GRcs.LGarXiv:1811.10943v22018PVANET: Deep but Lightweight Neural Networks for Real-time Object Detection
Kye-Hyeon Kim, Sanghoon Hong, Byungseok Roh +2
cs.CVarXiv:1608.08021v32016Mask-Guided Attention Network for Occluded Pedestrian Detection
Yanwei Pang, Jin Xie, Muhammad Haris Khan +3
cs.CVarXiv:1910.06160v22019OctSqueeze: Octree-Structured Entropy Model for LiDAR Compression
Lila Huang, Shenlong Wang, Kelvin Wong +2
eess.IVcs.CVarXiv:2005.07178v22020Unsupervised Learning of Object Keypoints for Perception and Control
Tejas Kulkarni, Ankush Gupta, Catalin Ionescu +4
cs.CVcs.LGarXiv:1906.11883v22019Blind2Unblind: Self-Supervised Image Denoising with Visible Blind Spots
Zejin Wang, Jiazheng Liu, Guoqing Li +1
eess.IVcs.CVarXiv:2203.06967v32022Co-Mixup: Saliency Guided Joint Mixup with Supermodular Diversity
Jang-Hyun Kim, Wonho Choo, Hosan Jeong +1
cs.LGcs.AIcs.CVarXiv:2102.03065v12021Unsupervised 3D Pose Estimation with Geometric Self-Supervision
Ching-Hang Chen, Ambrish Tyagi, Amit Agrawal +4
cs.CVarXiv:1904.04812v12019Making LLaMA SEE and Draw with SEED Tokenizer
Yuying Ge, Sijie Zhao, Ziyun Zeng +4
cs.CVarXiv:2310.01218v12023A Survey on Content-Aware Video Analysis for Sports
Huang-Chia Shih
cs.CVcs.MMarXiv:1703.01170v12017Confidence-Aware Ensemble and Long-Word Refinement for Artistic Text Recognition
Lucas A. Dias, Henrique A. Schulz, Rafaela de Miranda +3
cs.CVarXiv:2608.29970v12026On the Role of MRI Sequences in Cross-Dataset Generalization for Brain Tumor Segmentation
Henrique Zan Grande, João G. Pitol, Lucas B. Schuck +3
cs.CVarXiv:2608.29944v12026Dynamic Instance Normalization for Arbitrary Style Transfer
Yongcheng Jing, Xiao Liu, Yukang Ding +4
cs.CVarXiv:1911.06953v12019An end-to-end generative framework for video segmentation and recognition
Hilde Kuehne, Juergen Gall, Thomas Serre
cs.CVarXiv:1509.01947v22015AMT: All-Pairs Multi-Field Transforms for Efficient Frame Interpolation
Zhen Li, Zuo-Liang Zhu, Ling-Hao Han +3
cs.CVarXiv:2304.09790v12023CRAFT: Concept Recursive Activation FacTorization for Explainability
Thomas Fel, Agustin Picard, Louis Bethune +5
cs.CVcs.AIarXiv:2211.10154v22022RED: Reinforced Encoder-Decoder Networks for Action Anticipation
Jiyang Gao, Zhenheng Yang, Ram Nevatia
cs.CVarXiv:1707.04818v12017FuncRoom-Agent: Sequential Feed-Forward 3D Functional Indoor Scene Generation
Hao Feng, Zhi Zuo, MingJian Liang +6
cs.CVarXiv:2608.29519v12026When will you do what? - Anticipating Temporal Occurrences of Activities
Yazan Abu Farha, Alexander Richard, Juergen Gall
cs.CVarXiv:1804.00892v12018Seeing Through Extreme Visual Sparsity: Surface Understanding from a Single Random Visual Patch
Sindhuja Penchala, Sudip Mittal, Noorbakhsh Amiri Golilarz
cs.CVarXiv:2608.29475v12026CERF: Communication-Efficient and Retraining-Free Collaborative Perception
Jiuwu Hao, Ziyi Ni, Liguo Sun +5
cs.CVarXiv:2609.00951v12026SG-Net: Spatial Granularity Network for One-Stage Video Instance Segmentation
Dongfang Liu, Yiming Cui, Wenbo Tan +1
cs.CVarXiv:2103.10284v22021