Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
12,301 to 12,360 of 18,848
Classification of Time-Series Images Using Deep Convolutional Neural Networks
Nima Hatami, Yann Gavet, Johan Debayle
cs.CVarXiv:1710.00886v22017Identity-Guided Human Semantic Parsing for Person Re-Identification
Kuan Zhu, Haiyun Guo, Zhiwei Liu +2
cs.CVarXiv:2007.13467v12020Automatic Description Generation from Images: A Survey of Models, Datasets, and Evaluation Measures
Raffaella Bernardi, Ruket Cakici, Desmond Elliott +6
cs.CLcs.CVarXiv:1601.03896v22016Fast YOLO: A Fast You Only Look Once System for Real-time Embedded Object Detection in Video
Mohammad Javad Shafiee, Brendan Chywl, Francis Li +1
cs.CVcs.AIcs.NEarXiv:1709.05943v12017TADP: Task-Aware Deformable Prediction for Single-Stage 3D Object Detection
Su Wang, Yaochen Li, Min Yang +3
cs.CVcs.AIcs.ROarXiv:2608.27282v12026Pix2Video: Video Editing using Image Diffusion
Duygu Ceylan, Chun-Hao Paul Huang, Niloy J. Mitra
cs.CVarXiv:2303.12688v12023Interaction-and-Aggregation Network for Person Re-identification
Ruibing Hou, Bingpeng Ma, Hong Chang +3
cs.CVarXiv:1907.08435v12019EdgeNeXt: Efficiently Amalgamated CNN-Transformer Architecture for Mobile Vision Applications
Muhammad Maaz, Abdelrahman Shaker, Hisham Cholakkal +4
cs.CVarXiv:2206.10589v32022CoGeo-GS: Concept-Driven and Geometry-Aware Multi-Object Removal in 3D Scenes
Yuanxiang Ni, Xianliang Huang, Chenhang Ma +4
cs.CVcs.AIarXiv:2608.26656v12026LSTD: A Low-Shot Transfer Detector for Object Detection
Hao Chen, Yali Wang, Guoyou Wang +1
cs.CVarXiv:1803.01529v12018Automatic Instrument Segmentation in Robot-Assisted Surgery Using Deep Learning
Alexey Shvets, Alexander Rakhlin, Alexandr A. Kalinin +1
cs.CVarXiv:1803.01207v22018Disentangled and Controllable Face Image Generation via 3D Imitative-Contrastive Learning
Yu Deng, Jiaolong Yang, Dong Chen +2
cs.CVarXiv:2004.11660v22020In Defense of Pre-trained ImageNet Architectures for Real-time Semantic Segmentation of Road-driving Images
Marin Oršić, Ivan Krešo, Petra Bevandić +1
cs.CVarXiv:1903.08469v22019Learning to Discover Novel Visual Categories via Deep Transfer Clustering
Kai Han, Andrea Vedaldi, Andrew Zisserman
cs.CVarXiv:1908.09884v12019OpenOOD: Benchmarking Generalized Out-of-Distribution Detection
Jingkang Yang, Pengyun Wang, Dejian Zou +13
cs.CVcs.AIcs.LGarXiv:2210.07242v12022Tensor-based formulation and nuclear norm regularization for multi-energy computed tomography
Oguz Semerci, Ning Hao, Misha E. Kilmer +1
cs.CVphysics.med-pharXiv:1307.5348v12013Deep Multi-task Learning for Railway Track Inspection
Xavier Gibert, Vishal M. Patel, Rama Chellappa
cs.CVarXiv:1509.05267v12015Transferrable Prototypical Networks for Unsupervised Domain Adaptation
Yingwei Pan, Ting Yao, Yehao Li +3
cs.CVarXiv:1904.11227v12019When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations
Xiangning Chen, Cho-Jui Hsieh, Boqing Gong
cs.CVcs.LGarXiv:2106.01548v32021Relation-Aware Graph Attention Network for Visual Question Answering
Linjie Li, Zhe Gan, Yu Cheng +1
cs.CVcs.AIarXiv:1903.12314v32019SonoNet: Real-Time Detection and Localisation of Fetal Standard Scan Planes in Freehand Ultrasound
Christian F. Baumgartner, Konstantinos Kamnitsas, Jacqueline Matthew +5
cs.CVarXiv:1612.05601v22016PyMAF: 3D Human Pose and Shape Regression with Pyramidal Mesh Alignment Feedback Loop
Hongwen Zhang, Yating Tian, Xinchi Zhou +4
cs.CVarXiv:2103.16507v42021Classification-Reconstruction Learning for Open-Set Recognition
Ryota Yoshihashi, Wen Shao, Rei Kawakami +3
cs.CVarXiv:1812.04246v32018NeuralRecon: Real-Time Coherent 3D Reconstruction from Monocular Video
Jiaming Sun, Yiming Xie, Linghao Chen +2
cs.CVcs.ROarXiv:2104.00681v12021Learning Self-Consistency for Deepfake Detection
Tianchen Zhao, Xiang Xu, Mingze Xu +3
cs.CVarXiv:2012.09311v22020Animating Arbitrary Objects via Deep Motion Transfer
Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov +2
cs.GRcs.CVcs.LGarXiv:1812.08861v32018Kinect Range Sensing: Structured-Light versus Time-of-Flight Kinect
Hamed Sarbolandi, Damien Lefloch, Andreas Kolb
cs.CVarXiv:1505.05459v12015Blind Backdoors in Deep Learning Models
Eugene Bagdasaryan, Vitaly Shmatikov
cs.CRcs.CVcs.LGarXiv:2005.03823v420203D Human Pose Estimation in the Wild by Adversarial Learning
Wei Yang, Wanli Ouyang, Xiaolong Wang +3
cs.CVarXiv:1803.09722v22018R3LIVE: A Robust, Real-time, RGB-colored, LiDAR-Inertial-Visual tightly-coupled state Estimation and mapping package
Jiarong Lin, Fu Zhang
cs.ROcs.CVarXiv:2109.07982v12021Improving automated multiple sclerosis lesion segmentation with a cascaded 3D convolutional neural network approach
Sergi Valverde, Mariano Cabezas, Eloy Roura +7
cs.CVarXiv:1702.04869v12017Mixed supervision for surface-defect detection: from weakly to fully supervised learning
Jakob Božič, Domen Tabernik, Danijel Skočaj
cs.CVarXiv:2104.06064v32021Graph-Cut RANSAC
Daniel Barath, Jiri Matas
cs.CVarXiv:1706.00984v22017Multi-Attention Multi-Class Constraint for Fine-grained Image Recognition
Ming Sun, Yuchen Yuan, Feng Zhou +1
cs.CVarXiv:1806.05372v12018Unsolved Problems in ML Safety
Dan Hendrycks, Nicholas Carlini, John Schulman +1
cs.LGcs.AIcs.CLarXiv:2109.13916v52021Summaries:한국어Component Divide-and-Conquer for Real-World Image Super-Resolution
Pengxu Wei, Ziwei Xie, Hannan Lu +4
cs.CVarXiv:2008.01928v12020PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding
Zhen Li, Mingdeng Cao, Xintao Wang +3
cs.CVcs.AIcs.LGarXiv:2312.04461v12023Facial Landmark Detection: a Literature Survey
Yue Wu, Qiang Ji
cs.CVarXiv:1805.05563v12018HiFormer: Hierarchical Multi-scale Representations Using Transformers for Medical Image Segmentation
Moein Heidari, Amirhossein Kazerouni, Milad Soltany +4
cs.CVcs.AIarXiv:2207.08518v22022Diffusion Models already have a Semantic Latent Space
Mingi Kwon, Jaeseok Jeong, Youngjung Uh
cs.CVarXiv:2210.10960v22022ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
Wenlong Huang, Chen Wang, Yunzhu Li +2
cs.ROcs.AIcs.CVarXiv:2409.01652v22024Fully Convolutional Networks for Dense Semantic Labelling of High-Resolution Aerial Imagery
Jamie Sherrah
cs.CVarXiv:1606.02585v12016Universal Domain Adaptation through Self Supervision
Kuniaki Saito, Donghyun Kim, Stan Sclaroff +1
cs.CVarXiv:2002.07953v32020Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities
Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov +4
cs.CVarXiv:2203.14712v22022CFNet: Cascade and Fused Cost Volume for Robust Stereo Matching
Zhelun Shen, Yuchao Dai, Zhibo Rao
cs.CVarXiv:2104.04314v12021SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Yuan Zhang, Chun-Kai Fan, Junpeng Ma +8
cs.CVarXiv:2410.04417v42024ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis
Wangbo Yu, Jinbo Xing, Li Yuan +7
cs.CVarXiv:2409.02048v12024Neural Aggregation Network for Video Face Recognition
Jiaolong Yang, Peiran Ren, Dongqing Zhang +4
cs.CVcs.AIarXiv:1603.05474v42016Attention-Based Multimodal Fusion for Video Description
Chiori Hori, Takaaki Hori, Teng-Yok Lee +3
cs.CVcs.CLcs.MMarXiv:1701.03126v22017MetaIQA: Deep Meta-learning for No-Reference Image Quality Assessment
Hancheng Zhu, Leida Li, Jinjian Wu +2
eess.IVcs.CVarXiv:2004.05508v12020Neural Scene Graphs for Dynamic Scenes
Julian Ost, Fahim Mannan, Nils Thuerey +2
cs.CVcs.GRarXiv:2011.10379v32020StyleGAN-V: A Continuous Video Generator with the Price, Image Quality and Perks of StyleGAN2
Ivan Skorokhodov, Sergey Tulyakov, Mohamed Elhoseiny
cs.CVcs.AIcs.LGarXiv:2112.14683v42021Human uncertainty makes classification more robust
Joshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths +1
cs.CVarXiv:1908.07086v12019Lightweight Pyramid Networks for Image Deraining
Xueyang Fu, Borong Liang, Yue Huang +2
cs.CVarXiv:1805.06173v12018Single Image Reflection Separation with Perceptual Losses
Xuaner Zhang, Ren Ng, Qifeng Chen
cs.CVarXiv:1806.05376v12018Rodin: A Generative Model for Sculpting 3D Digital Avatars Using Diffusion
Tengfei Wang, Bo Zhang, Ting Zhang +8
cs.CVarXiv:2212.06135v12022Attention-Guided Reliability Scaling for Contrastive Decoding in Robust Audio-Visual Speech Recognition
YoungChae Kim, Da-Hee Yang, Joon-Hyuk Chang
cs.SDcs.CVeess.ASarXiv:2608.26213v12026Video Representation Learning by Dense Predictive Coding
Tengda Han, Weidi Xie, Andrew Zisserman
cs.CVarXiv:1909.04656v32019EfficientAD: Accurate Visual Anomaly Detection at Millisecond-Level Latencies
Kilian Batzner, Lars Heckler, Rebecca König
cs.CVarXiv:2303.14535v32023A Stable Multi-Scale Kernel for Topological Machine Learning
Jan Reininghaus, Stefan Huber, Ulrich Bauer +1
stat.MLcs.CVcs.LGarXiv:1412.6821v12014