Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

12,301 to 12,360 of 18,848

  1. Classification of Time-Series Images Using Deep Convolutional Neural Networks

    Nima Hatami, Yann Gavet, Johan Debayle

    cs.CVarXiv:1710.00886v22017
  2. Identity-Guided Human Semantic Parsing for Person Re-Identification

    Kuan Zhu, Haiyun Guo, Zhiwei Liu +2

    cs.CVarXiv:2007.13467v12020
  3. Automatic Description Generation from Images: A Survey of Models, Datasets, and Evaluation Measures

    Raffaella Bernardi, Ruket Cakici, Desmond Elliott +6

    cs.CLcs.CVarXiv:1601.03896v22016
  4. Fast YOLO: A Fast You Only Look Once System for Real-time Embedded Object Detection in Video

    Mohammad Javad Shafiee, Brendan Chywl, Francis Li +1

    cs.CVcs.AIcs.NEarXiv:1709.05943v12017
  5. TADP: Task-Aware Deformable Prediction for Single-Stage 3D Object Detection

    Su Wang, Yaochen Li, Min Yang +3

    cs.CVcs.AIcs.ROarXiv:2608.27282v12026
  6. Pix2Video: Video Editing using Image Diffusion

    Duygu Ceylan, Chun-Hao Paul Huang, Niloy J. Mitra

    cs.CVarXiv:2303.12688v12023
  7. Interaction-and-Aggregation Network for Person Re-identification

    Ruibing Hou, Bingpeng Ma, Hong Chang +3

    cs.CVarXiv:1907.08435v12019
  8. EdgeNeXt: Efficiently Amalgamated CNN-Transformer Architecture for Mobile Vision Applications

    Muhammad Maaz, Abdelrahman Shaker, Hisham Cholakkal +4

    cs.CVarXiv:2206.10589v32022
  9. CoGeo-GS: Concept-Driven and Geometry-Aware Multi-Object Removal in 3D Scenes

    Yuanxiang Ni, Xianliang Huang, Chenhang Ma +4

    cs.CVcs.AIarXiv:2608.26656v12026
  10. LSTD: A Low-Shot Transfer Detector for Object Detection

    Hao Chen, Yali Wang, Guoyou Wang +1

    cs.CVarXiv:1803.01529v12018
  11. Automatic Instrument Segmentation in Robot-Assisted Surgery Using Deep Learning

    Alexey Shvets, Alexander Rakhlin, Alexandr A. Kalinin +1

    cs.CVarXiv:1803.01207v22018
  12. Disentangled and Controllable Face Image Generation via 3D Imitative-Contrastive Learning

    Yu Deng, Jiaolong Yang, Dong Chen +2

    cs.CVarXiv:2004.11660v22020
  13. In Defense of Pre-trained ImageNet Architectures for Real-time Semantic Segmentation of Road-driving Images

    Marin Oršić, Ivan Krešo, Petra Bevandić +1

    cs.CVarXiv:1903.08469v22019
  14. Learning to Discover Novel Visual Categories via Deep Transfer Clustering

    Kai Han, Andrea Vedaldi, Andrew Zisserman

    cs.CVarXiv:1908.09884v12019
  15. OpenOOD: Benchmarking Generalized Out-of-Distribution Detection

    Jingkang Yang, Pengyun Wang, Dejian Zou +13

    cs.CVcs.AIcs.LGarXiv:2210.07242v12022
  16. Tensor-based formulation and nuclear norm regularization for multi-energy computed tomography

    Oguz Semerci, Ning Hao, Misha E. Kilmer +1

    cs.CVphysics.med-pharXiv:1307.5348v12013
  17. Deep Multi-task Learning for Railway Track Inspection

    Xavier Gibert, Vishal M. Patel, Rama Chellappa

    cs.CVarXiv:1509.05267v12015
  18. Transferrable Prototypical Networks for Unsupervised Domain Adaptation

    Yingwei Pan, Ting Yao, Yehao Li +3

    cs.CVarXiv:1904.11227v12019
  19. When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations

    Xiangning Chen, Cho-Jui Hsieh, Boqing Gong

    cs.CVcs.LGarXiv:2106.01548v32021
  20. Relation-Aware Graph Attention Network for Visual Question Answering

    Linjie Li, Zhe Gan, Yu Cheng +1

    cs.CVcs.AIarXiv:1903.12314v32019
  21. SonoNet: Real-Time Detection and Localisation of Fetal Standard Scan Planes in Freehand Ultrasound

    Christian F. Baumgartner, Konstantinos Kamnitsas, Jacqueline Matthew +5

    cs.CVarXiv:1612.05601v22016
  22. PyMAF: 3D Human Pose and Shape Regression with Pyramidal Mesh Alignment Feedback Loop

    Hongwen Zhang, Yating Tian, Xinchi Zhou +4

    cs.CVarXiv:2103.16507v42021
  23. Classification-Reconstruction Learning for Open-Set Recognition

    Ryota Yoshihashi, Wen Shao, Rei Kawakami +3

    cs.CVarXiv:1812.04246v32018
  24. NeuralRecon: Real-Time Coherent 3D Reconstruction from Monocular Video

    Jiaming Sun, Yiming Xie, Linghao Chen +2

    cs.CVcs.ROarXiv:2104.00681v12021
  25. Learning Self-Consistency for Deepfake Detection

    Tianchen Zhao, Xiang Xu, Mingze Xu +3

    cs.CVarXiv:2012.09311v22020
  26. Animating Arbitrary Objects via Deep Motion Transfer

    Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov +2

    cs.GRcs.CVcs.LGarXiv:1812.08861v32018
  27. Kinect Range Sensing: Structured-Light versus Time-of-Flight Kinect

    Hamed Sarbolandi, Damien Lefloch, Andreas Kolb

    cs.CVarXiv:1505.05459v12015
  28. Blind Backdoors in Deep Learning Models

    Eugene Bagdasaryan, Vitaly Shmatikov

    cs.CRcs.CVcs.LGarXiv:2005.03823v42020
  29. 3D Human Pose Estimation in the Wild by Adversarial Learning

    Wei Yang, Wanli Ouyang, Xiaolong Wang +3

    cs.CVarXiv:1803.09722v22018
  30. R3LIVE: A Robust, Real-time, RGB-colored, LiDAR-Inertial-Visual tightly-coupled state Estimation and mapping package

    Jiarong Lin, Fu Zhang

    cs.ROcs.CVarXiv:2109.07982v12021
  31. Improving automated multiple sclerosis lesion segmentation with a cascaded 3D convolutional neural network approach

    Sergi Valverde, Mariano Cabezas, Eloy Roura +7

    cs.CVarXiv:1702.04869v12017
  32. Mixed supervision for surface-defect detection: from weakly to fully supervised learning

    Jakob Božič, Domen Tabernik, Danijel Skočaj

    cs.CVarXiv:2104.06064v32021
  33. Graph-Cut RANSAC

    Daniel Barath, Jiri Matas

    cs.CVarXiv:1706.00984v22017
  34. Multi-Attention Multi-Class Constraint for Fine-grained Image Recognition

    Ming Sun, Yuchen Yuan, Feng Zhou +1

    cs.CVarXiv:1806.05372v12018
  35. Unsolved Problems in ML Safety

    Dan Hendrycks, Nicholas Carlini, John Schulman +1

    cs.LGcs.AIcs.CLarXiv:2109.13916v52021
    Summaries:한국어
  36. Component Divide-and-Conquer for Real-World Image Super-Resolution

    Pengxu Wei, Ziwei Xie, Hannan Lu +4

    cs.CVarXiv:2008.01928v12020
  37. PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding

    Zhen Li, Mingdeng Cao, Xintao Wang +3

    cs.CVcs.AIcs.LGarXiv:2312.04461v12023
  38. Facial Landmark Detection: a Literature Survey

    Yue Wu, Qiang Ji

    cs.CVarXiv:1805.05563v12018
  39. HiFormer: Hierarchical Multi-scale Representations Using Transformers for Medical Image Segmentation

    Moein Heidari, Amirhossein Kazerouni, Milad Soltany +4

    cs.CVcs.AIarXiv:2207.08518v22022
  40. Diffusion Models already have a Semantic Latent Space

    Mingi Kwon, Jaeseok Jeong, Youngjung Uh

    cs.CVarXiv:2210.10960v22022
  41. ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

    Wenlong Huang, Chen Wang, Yunzhu Li +2

    cs.ROcs.AIcs.CVarXiv:2409.01652v22024
  42. Fully Convolutional Networks for Dense Semantic Labelling of High-Resolution Aerial Imagery

    Jamie Sherrah

    cs.CVarXiv:1606.02585v12016
  43. Universal Domain Adaptation through Self Supervision

    Kuniaki Saito, Donghyun Kim, Stan Sclaroff +1

    cs.CVarXiv:2002.07953v32020
  44. Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities

    Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov +4

    cs.CVarXiv:2203.14712v22022
  45. CFNet: Cascade and Fused Cost Volume for Robust Stereo Matching

    Zhelun Shen, Yuchao Dai, Zhibo Rao

    cs.CVarXiv:2104.04314v12021
  46. SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

    Yuan Zhang, Chun-Kai Fan, Junpeng Ma +8

    cs.CVarXiv:2410.04417v42024
  47. ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis

    Wangbo Yu, Jinbo Xing, Li Yuan +7

    cs.CVarXiv:2409.02048v12024
  48. Neural Aggregation Network for Video Face Recognition

    Jiaolong Yang, Peiran Ren, Dongqing Zhang +4

    cs.CVcs.AIarXiv:1603.05474v42016
  49. Attention-Based Multimodal Fusion for Video Description

    Chiori Hori, Takaaki Hori, Teng-Yok Lee +3

    cs.CVcs.CLcs.MMarXiv:1701.03126v22017
  50. MetaIQA: Deep Meta-learning for No-Reference Image Quality Assessment

    Hancheng Zhu, Leida Li, Jinjian Wu +2

    eess.IVcs.CVarXiv:2004.05508v12020
  51. Neural Scene Graphs for Dynamic Scenes

    Julian Ost, Fahim Mannan, Nils Thuerey +2

    cs.CVcs.GRarXiv:2011.10379v32020
  52. StyleGAN-V: A Continuous Video Generator with the Price, Image Quality and Perks of StyleGAN2

    Ivan Skorokhodov, Sergey Tulyakov, Mohamed Elhoseiny

    cs.CVcs.AIcs.LGarXiv:2112.14683v42021
  53. Human uncertainty makes classification more robust

    Joshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths +1

    cs.CVarXiv:1908.07086v12019
  54. Lightweight Pyramid Networks for Image Deraining

    Xueyang Fu, Borong Liang, Yue Huang +2

    cs.CVarXiv:1805.06173v12018
  55. Single Image Reflection Separation with Perceptual Losses

    Xuaner Zhang, Ren Ng, Qifeng Chen

    cs.CVarXiv:1806.05376v12018
  56. Rodin: A Generative Model for Sculpting 3D Digital Avatars Using Diffusion

    Tengfei Wang, Bo Zhang, Ting Zhang +8

    cs.CVarXiv:2212.06135v12022
  57. Attention-Guided Reliability Scaling for Contrastive Decoding in Robust Audio-Visual Speech Recognition

    YoungChae Kim, Da-Hee Yang, Joon-Hyuk Chang

    cs.SDcs.CVeess.ASarXiv:2608.26213v12026
  58. Video Representation Learning by Dense Predictive Coding

    Tengda Han, Weidi Xie, Andrew Zisserman

    cs.CVarXiv:1909.04656v32019
  59. EfficientAD: Accurate Visual Anomaly Detection at Millisecond-Level Latencies

    Kilian Batzner, Lars Heckler, Rebecca König

    cs.CVarXiv:2303.14535v32023
  60. A Stable Multi-Scale Kernel for Topological Machine Learning

    Jan Reininghaus, Stefan Huber, Ulrich Bauer +1

    stat.MLcs.CVcs.LGarXiv:1412.6821v12014