Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

9,181 to 9,240 of 18,830

  1. Explainable Medical Imaging AI Needs Human-Centered Design: Guidelines and Evidence from a Systematic Review

    Haomin Chen, Catalina Gomez, Chien-Ming Huang +1

    cs.HCcs.CVcs.LGarXiv:2112.12596v42021
  2. Practical Blind Image Denoising via Swin-Conv-UNet and Data Synthesis

    Kai Zhang, Yawei Li, Jingyun Liang +6

    cs.CVcs.GReess.IVarXiv:2203.13278v42022
  3. Parametric Multimodal User Memory: Storing What Captions Cannot Carry

    Bojie Li, Noah Shi

    cs.CLcs.AIcs.CVarXiv:2608.28609v12026
  4. Batch-Instance Normalization for Adaptively Style-Invariant Neural Networks

    Hyeonseob Nam, Hyo-Eun Kim

    cs.CVarXiv:1805.07925v32018
  5. DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

    Jiashu Zhu, Yanhao Zheng, Ruitian Tian +7

    cs.CVcs.SDarXiv:2608.31106v12026
  6. Fast Inference in Sparse Coding Algorithms with Applications to Object Recognition

    Koray Kavukcuoglu, Marc'Aurelio Ranzato, Yann LeCun

    cs.CVcs.LGarXiv:1010.3467v12010
  7. Defense Against Adversarial Attacks Using Feature Scattering-based Adversarial Training

    Haichao Zhang, Jianyu Wang

    cs.CVcs.CRcs.LGarXiv:1907.10764v42019
  8. Learning Task-Oriented Grasping for Tool Manipulation from Simulated Self-Supervision

    Kuan Fang, Yuke Zhu, Animesh Garg +4

    cs.ROcs.CVcs.LGarXiv:1806.09266v12018
  9. Image Deformation Meta-Networks for One-Shot Learning

    Zitian Chen, Yanwei Fu, Yu-Xiong Wang +3

    cs.CVarXiv:1905.11641v22019
  10. Shield: Fast, Practical Defense and Vaccination for Deep Learning using JPEG Compression

    Nilaksh Das, Madhuri Shanbhogue, Shang-Tse Chen +5

    cs.CVcs.AIcs.CRarXiv:1802.06816v12018
  11. Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained Models

    Guillermo Ortiz-Jimenez, Alessandro Favero, Pascal Frossard

    cs.LGcs.CVarXiv:2305.12827v32023
  12. Unsupervised Domain Adaptation through Self-Supervision

    Yu Sun, Eric Tzeng, Trevor Darrell +1

    cs.LGcs.CVstat.MLarXiv:1909.11825v22019
  13. HyperSeg: Patch-wise Hypernetwork for Real-time Semantic Segmentation

    Yuval Nirkin, Lior Wolf, Tal Hassner

    cs.CVarXiv:2012.11582v22020
  14. AdvSim: Generating Safety-Critical Scenarios for Self-Driving Vehicles

    Jingkang Wang, Ava Pun, James Tu +5

    cs.ROcs.AIcs.CVarXiv:2101.06549v42021
  15. Multispectral and Hyperspectral Image Fusion by MS/HS Fusion Net

    Qi Xie, Minghao Zhou, Qian Zhao +3

    cs.CVarXiv:1901.03281v12019
  16. 3D Gaussian Splatting as Markov Chain Monte Carlo

    Shakiba Kheradmand, Daniel Rebain, Gopal Sharma +6

    cs.CVarXiv:2404.09591v32024
  17. FedRolex: Model-Heterogeneous Federated Learning with Rolling Sub-Model Extraction

    Samiul Alam, Luyang Liu, Ming Yan +1

    cs.LGcs.CRcs.CVarXiv:2212.01548v22022
  18. BLARM: Animating 3D Objects from Video via Blending Latent Rigid Motion Primitives

    Pradyumn Goyal, Yizhak Ben-Shabat, Hsueh-Ti Derek Liu +6

    cs.CVarXiv:2608.31113v12026
  19. Adversarial Unlearning of Backdoors via Implicit Hypergradient

    Yi Zeng, Si Chen, Won Park +3

    cs.LGcs.CRcs.CVarXiv:2110.03735v42021
  20. DeepRhythm: Exposing DeepFakes with Attentional Visual Heartbeat Rhythms

    Hua Qi, Qing Guo, Felix Juefei-Xu +5

    cs.CVarXiv:2006.07634v22020
  21. ICDAR2017 Competition on Reading Chinese Text in the Wild (RCTW-17)

    Baoguang Shi, Cong Yao, Minghui Liao +6

    cs.CVarXiv:1708.09585v32017
  22. Unsupervised Person Re-identification by Deep Learning Tracklet Association

    Minxian Li, Xiatian Zhu, Shaogang Gong

    cs.CVarXiv:1809.02874v12018
  23. Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation

    Zhenxin Li, Kailin Li, Shihao Wang +9

    cs.CVarXiv:2406.06978v42024
  24. Cross View Fusion for 3D Human Pose Estimation

    Haibo Qiu, Chunyu Wang, Jingdong Wang +2

    cs.CVarXiv:1909.01203v12019
  25. Implicit Identity Leakage: The Stumbling Block to Improving Deepfake Detection Generalization

    Shichao Dong, Jin Wang, Renhe Ji +3

    cs.CVarXiv:2210.14457v22022
  26. Practical Full Resolution Learned Lossless Image Compression

    Fabian Mentzer, Eirikur Agustsson, Michael Tschannen +2

    eess.IVcs.CVcs.LGarXiv:1811.12817v32018
  27. BioCLIP: A Vision Foundation Model for the Tree of Life

    Samuel Stevens, Jiaman Wu, Matthew J Thompson +9

    cs.CVcs.CLcs.LGarXiv:2311.18803v32023
  28. PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback

    Xueqing Wu, Ashwin Balasubramanian, Bingxuan Li +7

    cs.CLcs.CVarXiv:2608.30241v12026
  29. Masked Discrimination for Self-Supervised Learning on Point Clouds

    Haotian Liu, Mu Cai, Yong Jae Lee

    cs.CVarXiv:2203.11183v22022
  30. PCAN: 3D Attention Map Learning Using Contextual Information for Point Cloud Based Retrieval

    Wenxiao Zhang, Chunxia Xiao

    cs.CVcs.ROarXiv:1904.09793v12019
  31. MNIST-C: A Robustness Benchmark for Computer Vision

    Norman Mu, Justin Gilmer

    cs.CVcs.LGarXiv:1906.02337v12019
  32. Social Scene Understanding: End-to-End Multi-Person Action Localization and Collective Activity Recognition

    Timur Bagautdinov, Alexandre Alahi, François Fleuret +2

    cs.CVarXiv:1611.09078v12016
  33. Dress Code: High-Resolution Multi-Category Virtual Try-On

    Davide Morelli, Matteo Fincato, Marcella Cornia +3

    cs.CVcs.AIcs.GRarXiv:2204.08532v22022
  34. KRISP: Integrating Implicit and Symbolic Knowledge for Open-Domain Knowledge-Based VQA

    Kenneth Marino, Xinlei Chen, Devi Parikh +2

    cs.CVcs.CLarXiv:2012.11014v12020
  35. Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling

    Minghan Qin, Yuang Wang, Xiuyu Yang +6

    cs.CVcs.AIarXiv:2608.30821v12026
  36. Accurate Optical Flow via Direct Cost Volume Processing

    Jia Xu, René Ranftl, Vladlen Koltun

    cs.CVarXiv:1704.07325v12017
  37. Transformer-Based Attention Networks for Continuous Pixel-Wise Prediction

    Guanglei Yang, Hao Tang, Mingli Ding +2

    cs.CVarXiv:2103.12091v22021
  38. Box-driven Class-wise Region Masking and Filling Rate Guided Loss for Weakly Supervised Semantic Segmentation

    Chunfeng Song, Yan Huang, Wanli Ouyang +1

    cs.CVarXiv:1904.11693v12019
  39. RegNet: Multimodal Sensor Registration Using Deep Neural Networks

    Nick Schneider, Florian Piewak, Christoph Stiller +1

    cs.CVcs.AIcs.LGarXiv:1707.03167v12017
  40. LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

    Shilong Liu, Hao Cheng, Haotian Liu +10

    cs.CVcs.AIcs.CLarXiv:2311.05437v12023
  41. Video Captioning via Hierarchical Reinforcement Learning

    Xin Wang, Wenhu Chen, Jiawei Wu +2

    cs.CVcs.AIcs.CLarXiv:1711.11135v32017
  42. Learning Human-Object Interaction Detection using Interaction Points

    Tiancai Wang, Tong Yang, Martin Danelljan +3

    cs.CVarXiv:2003.14023v12020
  43. BlockGAN: Learning 3D Object-aware Scene Representations from Unlabelled Images

    Thu Nguyen-Phuoc, Christian Richardt, Long Mai +2

    cs.CVarXiv:2002.08988v42020
  44. PatDNN: Achieving Real-Time DNN Execution on Mobile Devices with Pattern-based Weight Pruning

    Wei Niu, Xiaolong Ma, Sheng Lin +5

    cs.LGcs.CVcs.DCarXiv:2001.00138v42020
  45. Low-light Image Enhancement via Breaking Down the Darkness

    Qiming Hu, Xiaojie Guo

    cs.CVarXiv:2111.15557v12021
  46. iGibson 1.0: a Simulation Environment for Interactive Tasks in Large Realistic Scenes

    Bokui Shen, Fei Xia, Chengshu Li +12

    cs.AIcs.CVcs.ROarXiv:2012.02924v62020
  47. Learning image representations tied to ego-motion

    Dinesh Jayaraman, Kristen Grauman

    cs.CVcs.AIstat.MLarXiv:1505.02206v22015
  48. An Experimental-based Review of Image Enhancement and Image Restoration Methods for Underwater Imaging

    Yan Wang, Wei Song, Giancarlo Fortino +3

    eess.IVcs.CVcs.MMarXiv:1907.03246v12019
  49. Pillar-based Object Detection for Autonomous Driving

    Yue Wang, Alireza Fathi, Abhijit Kundu +4

    cs.CVcs.LGcs.ROarXiv:2007.10323v22020
  50. Variational Relational Point Completion Network

    Liang Pan, Xinyi Chen, Zhongang Cai +4

    cs.CVcs.LGarXiv:2104.10154v12021
  51. Ear Recognition: More Than a Survey

    Žiga Emeršič, Vitomir Štruc, Peter Peer

    cs.CVarXiv:1611.06203v22016
  52. Seeing Voices and Hearing Faces: Cross-modal biometric matching

    Arsha Nagrani, Samuel Albanie, Andrew Zisserman

    cs.CVarXiv:1804.00326v22018
  53. The Devil is in the Details: Delving into Unbiased Data Processing for Human Pose Estimation

    Junjie Huang, Zheng Zhu, Feng Guo +2

    cs.CVarXiv:1911.07524v22019
  54. Trustworthy clinical AI solutions: a unified review of uncertainty quantification in deep learning models for medical image analysis

    Benjamin Lambert, Florence Forbes, Alan Tucholka +3

    eess.IVcs.AIcs.CVarXiv:2210.03736v12022
  55. Self-trained Deep Ordinal Regression for End-to-End Video Anomaly Detection

    Guansong Pang, Cheng Yan, Chunhua Shen +2

    cs.CVarXiv:2003.06780v12020
  56. Single-Stage Multi-Person Pose Machines

    Xuecheng Nie, Jianfeng Zhang, Shuicheng Yan +1

    cs.CVarXiv:1908.09220v12019
  57. Differential Treatment for Stuff and Things: A Simple Unsupervised Domain Adaptation Method for Semantic Segmentation

    Zhonghao Wang, Mo Yu, Yunchao Wei +5

    cs.CVcs.LGeess.IVarXiv:2003.08040v32020
  58. Are GAN generated images easy to detect? A critical analysis of the state-of-the-art

    Diego Gragnaniello, Davide Cozzolino, Francesco Marra +2

    cs.CVcs.AIarXiv:2104.02617v12021
  59. AdaCos: Adaptively Scaling Cosine Logits for Effectively Learning Deep Face Representations

    Xiao Zhang, Rui Zhao, Yu Qiao +2

    cs.CVarXiv:1905.00292v22019
  60. Automatic Breast Ultrasound Image Segmentation: A Survey

    Min Xian, Yingtao Zhang, H. D. Cheng +3

    cs.CVcs.LGarXiv:1704.01472v22017