Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,181 to 3,240 of 18,839

  1. Deep Burst Super-Resolution

    Goutam Bhat, Martin Danelljan, Luc Van Gool +1

    cs.CVarXiv:2101.10997v22021
  2. HandOccNet: Occlusion-Robust 3D Hand Mesh Estimation Network

    JoonKyu Park, Yeonguk Oh, Gyeongsik Moon +2

    cs.CVarXiv:2203.14564v12022
  3. World-Grounded Human Motion Recovery via Gravity-View Coordinates

    Zehong Shen, Huaijin Pi, Yan Xia +6

    cs.CVcs.AIarXiv:2409.06662v12024
  4. Effective Aesthetics Prediction with Multi-level Spatially Pooled Features

    Vlad Hosu, Bastian Goldlucke, Dietmar Saupe

    cs.CVarXiv:1904.01382v12019
  5. Connections Between Nuclear Norm and Frobenius Norm Based Representations

    Xi Peng, Canyi Lu, Zhang Yi +1

    cs.CVarXiv:1502.07423v22015
  6. 3D Object Detection from Images for Autonomous Driving: A Survey

    Xinzhu Ma, Wanli Ouyang, Andrea Simonelli +1

    cs.CVarXiv:2202.02980v62022
  7. LowKey: Leveraging Adversarial Attacks to Protect Social Media Users from Facial Recognition

    Valeriia Cherepanova, Micah Goldblum, Harrison Foley +4

    cs.CVcs.CRcs.LGarXiv:2101.07922v22021
  8. DwNet: Dense warp-based network for pose-guided human video generation

    Polina Zablotskaia, Aliaksandr Siarohin, Bo Zhao +1

    cs.CVcs.LGarXiv:1910.09139v12019
  9. Weakly-Supervised Action Localization by Generative Attention Modeling

    Baifeng Shi, Qi Dai, Yadong Mu +1

    cs.CVarXiv:2003.12424v22020
  10. EditGuard: Versatile Image Watermarking for Tamper Localization and Copyright Protection

    Xuanyu Zhang, Runyi Li, Jiwen Yu +3

    cs.CVarXiv:2312.08883v12023
  11. EmotiCon: Context-Aware Multimodal Emotion Recognition using Frege's Principle

    Trisha Mittal, Pooja Guhan, Uttaran Bhattacharya +3

    cs.CVcs.HCeess.IVarXiv:2003.06692v12020
  12. Learning Monocular Dense Depth from Events

    Javier Hidalgo-Carrió, Daniel Gehrig, Davide Scaramuzza

    cs.CVcs.LGarXiv:2010.08350v22020
  13. DaST: Data-free Substitute Training for Adversarial Attacks

    Mingyi Zhou, Jing Wu, Yipeng Liu +2

    cs.CRcs.CVcs.LGarXiv:2003.12703v22020
  14. Visibility-aware Multi-view Stereo Network

    Jingyang Zhang, Yao Yao, Shiwei Li +2

    cs.CVarXiv:2008.07928v22020
  15. 3C-Net: Category Count and Center Loss for Weakly-Supervised Action Localization

    Sanath Narayan, Hisham Cholakkal, Fahad Shahbaz Khan +1

    cs.CVarXiv:1908.08216v22019
  16. Deep learning enhanced mobile-phone microscopy

    Yair Rivenson, Hatice Ceylan Koydemir, Hongda Wang +8

    cs.LGcs.CVphysics.med-pharXiv:1712.04139v12017
  17. MC-Blur: A Comprehensive Benchmark for Image Deblurring

    Kaihao Zhang, Tao Wang, Wenhan Luo +6

    cs.CVarXiv:2112.00234v32021
  18. Wan-Image: Pushing the Boundaries of Generative Visual Intelligence

    Chaojie Mao, Chen-Wei Xie, Chongyang Zhong +55

    cs.CVarXiv:2604.19858v22026
  19. VISD: Enhancing Video Reasoning via Structured Self-Distillation

    Hao Lin, Kunyang Lv, Xu Jiang +5

    cs.CVcs.AIarXiv:2605.06094v52026
  20. Medical Graph RAG: Towards Safe Medical Large Language Model via Graph Retrieval-Augmented Generation

    Junde Wu, Jiayuan Zhu, Yunli Qi +4

    cs.CVarXiv:2408.04187v22024
  21. Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing

    Bingyu Li, Haocheng Dong, Da Zhang +3

    cs.CVcs.AIarXiv:2509.12040v32025
  22. PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop

    Chenyu Li, Oscar Michel, Xichen Pan +3

    cs.CVarXiv:2503.09595v12025
  23. Good Practice in CNN Feature Transfer

    Liang Zheng, Yali Zhao, Shengjin Wang +2

    cs.CVarXiv:1604.00133v12016
  24. Person Re-identification with Correspondence Structure Learning

    Yang Shen, Weiyao Lin, Junchi Yan +3

    cs.CVarXiv:1504.06243v12015
  25. Hierarchical Graph Representations in Digital Pathology

    Pushpak Pati, Guillaume Jaume, Antonio Foncubierta +14

    cs.CVarXiv:2102.11057v22021
  26. CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

    Haoyu Song, Li Dong, Wei-Nan Zhang +2

    cs.CVcs.CLarXiv:2203.07190v12022
  27. On the Robustness of the CVPR 2018 White-Box Adversarial Example Defenses

    Anish Athalye, Nicholas Carlini

    cs.CVcs.CRcs.LGarXiv:1804.03286v12018
  28. Temporal Residual Neural Radiance Fields for Monocular Video Dynamic Human Body Reconstruction

    Tianle Du, Jie Wang, Xiaolong Xie +3

    cs.CVarXiv:2609.04984v12026
  29. Explainable COVID-19 Detection Using Chest CT Scans and Deep Learning

    Hammam Alshazly, Christoph Linse, Erhardt Barth +1

    eess.IVcs.CVarXiv:2011.05317v12020
  30. SPG: Unsupervised Domain Adaptation for 3D Object Detection via Semantic Point Generation

    Qiangeng Xu, Yin Zhou, Weiyue Wang +2

    cs.CVarXiv:2108.06709v12021
  31. Neural Proximal Gradient Descent for Compressive Imaging

    Morteza Mardani, Qingyun Sun, Shreyas Vasawanala +4

    cs.CVcs.LGarXiv:1806.03963v12018
  32. Invertible generative models for inverse problems: mitigating representation error and dataset bias

    Muhammad Asim, Mara Daniels, Oscar Leong +2

    cs.CVarXiv:1905.11672v52019
  33. CollaGAN : Collaborative GAN for Missing Image Data Imputation

    Dongwook Lee, Junyoung Kim, Won-Jin Moon +1

    cs.CVcs.LGstat.MLarXiv:1901.09764v32019
  34. Synthesize then Compare: Detecting Failures and Anomalies for Semantic Segmentation

    Yingda Xia, Yi Zhang, Fengze Liu +2

    cs.CVarXiv:2003.08440v22020
  35. VinVL: Revisiting Visual Representations in Vision-Language Models

    Pengchuan Zhang, Xiujun Li, Xiaowei Hu +5

    cs.CVcs.AIcs.CLarXiv:2101.00529v22021
  36. Image Classification of Melanoma, Nevus and Seborrheic Keratosis by Deep Neural Network Ensemble

    Kazuhisa Matsunaga, Akira Hamada, Akane Minagawa +1

    cs.CVarXiv:1703.03108v12017
  37. Py-Feat: Python Facial Expression Analysis Toolbox

    Jin Hyun Cheong, Eshin Jolly, Tiankang Xie +3

    cs.CVcs.LGeess.IVarXiv:2104.03509v42021
  38. Automatic Hip Fracture Identification and Functional Subclassification with Deep Learning

    Justin D Krogue, Kaiyang V Cheng, Kevin M Hwang +13

    q-bio.QMcs.CVcs.LGarXiv:1909.06326v12019
  39. CamoFormer: Masked Separable Attention for Camouflaged Object Detection

    Bowen Yin, Xuying Zhang, Qibin Hou +3

    cs.CVarXiv:2212.06570v12022
  40. Restoring Images in Adverse Weather Conditions via Histogram Transformer

    Shangquan Sun, Wenqi Ren, Xinwei Gao +2

    cs.CVarXiv:2407.10172v22024
  41. Anomalib: A Deep Learning Library for Anomaly Detection

    Samet Akcay, Dick Ameln, Ashwin Vaidya +3

    cs.CVcs.LGarXiv:2202.08341v12022
  42. Associative Alignment for Few-shot Image Classification

    Arman Afrasiyabi, Jean-François Lalonde, Christian Gagné

    cs.CVcs.LGarXiv:1912.05094v32019
  43. MIN2Net: End-to-End Multi-Task Learning for Subject-Independent Motor Imagery EEG Classification

    Phairot Autthasan, Rattanaphon Chaisaen, Thapanun Sudhawiyangkul +7

    eess.SPcs.AIcs.CVarXiv:2102.03814v42021
  44. Object Level Visual Reasoning in Videos

    Fabien Baradel, Natalia Neverova, Christian Wolf +2

    cs.CVarXiv:1806.06157v32018
  45. Frequency-driven Imperceptible Adversarial Attack on Semantic Similarity

    Cheng Luo, Qinliang Lin, Weicheng Xie +3

    cs.CVarXiv:2203.05151v42022
  46. Help, Anna! Visual Navigation with Natural Multimodal Assistance via Retrospective Curiosity-Encouraging Imitation Learning

    Khanh Nguyen, Hal Daumé

    cs.HCcs.AIcs.CLarXiv:1909.01871v62019
  47. When AWGN-based Denoiser Meets Real Noises

    Yuqian Zhou, Jianbo Jiao, Haibin Huang +4

    cs.CVarXiv:1904.03485v22019
  48. CholecSeg8k: A Semantic Segmentation Dataset for Laparoscopic Cholecystectomy Based on Cholec80

    W. -Y. Hong, C. -L. Kao, Y. -H. Kuo +3

    cs.CVarXiv:2012.12453v12020
  49. FREE: Feature Refinement for Generalized Zero-Shot Learning

    Shiming Chen, Wenjie Wang, Beihao Xia +4

    cs.CVcs.AIarXiv:2107.13807v12021
  50. LENS: Localization enhanced by NeRF synthesis

    Arthur Moreau, Nathan Piasco, Dzmitry Tsishkou +2

    cs.CVcs.AIcs.LGarXiv:2110.06558v12021
  51. Hessian Schatten-Norm Regularization for Linear Inverse Problems

    Stamatios Lefkimmiatis, John Paul Ward, Michael Unser

    math.OCcs.CVmath.NAarXiv:1209.3318v32012
  52. MLPACK: A Scalable C++ Machine Learning Library

    Ryan R. Curtin, James R. Cline, N. P. Slagle +4

    cs.MScs.CVcs.LGarXiv:1210.6293v12012
  53. Text to Image Generation with Semantic-Spatial Aware GAN

    Kai Hu, Wentong Liao, Michael Ying Yang +1

    cs.CVcs.LGarXiv:2104.00567v62021
  54. Instance Segmentation in 3D Scenes using Semantic Superpoint Tree Networks

    Zhihao Liang, Zhihao Li, Songcen Xu +2

    cs.CVarXiv:2108.07478v12021
  55. GODS: Generalized One-class Discriminative Subspaces for Anomaly Detection

    Jue Wang, Anoop Cherian

    cs.CVarXiv:1908.05884v12019
  56. PaStaNet: Toward Human Activity Knowledge Engine

    Yong-Lu Li, Liang Xu, Xinpeng Liu +7

    cs.CVcs.AIcs.LGarXiv:2004.00945v22020
  57. Cache Me if You Can: Accelerating Diffusion Models through Block Caching

    Felix Wimbauer, Bichen Wu, Edgar Schoenfeld +11

    cs.CVarXiv:2312.03209v22023
  58. Learning Warped Guidance for Blind Face Restoration

    Xiaoming Li, Ming Liu, Yuting Ye +3

    cs.CVarXiv:1804.04829v22018
  59. Augmenting Physical Models with Deep Networks for Complex Dynamics Forecasting

    Yuan Yin, Vincent Le Guen, Jérémie Dona +4

    stat.MLcs.AIcs.CVarXiv:2010.04456v62020
  60. Supervised Transformer Network for Efficient Face Detection

    Dong Chen, Gang Hua, Fang Wen +1

    cs.CVarXiv:1607.05477v12016