Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,001 to 3,060 of 18,988

  1. Dense Cross-Query-and-Support Attention Weighted Mask Aggregation for Few-Shot Segmentation

    Xinyu Shi, Dong Wei, Yu Zhang +5

    cs.CVarXiv:2207.08549v12022
  2. Feature Erasing and Diffusion Network for Occluded Person Re-Identification

    Zhikang Wang, Feng Zhu, Shixiang Tang +3

    cs.CVarXiv:2112.08740v22021
  3. In-Context LoRA for Diffusion Transformers

    Lianghua Huang, Wei Wang, Zhi-Fan Wu +6

    cs.CVcs.GRarXiv:2410.23775v32024
  4. Fine-grained Visual Textual Alignment for Cross-Modal Retrieval using Transformer Encoders

    Nicola Messina, Giuseppe Amato, Andrea Esuli +3

    cs.CVarXiv:2008.05231v22020
  5. Read, Watch, and Move: Reinforcement Learning for Temporally Grounding Natural Language Descriptions in Videos

    Dongliang He, Xiang Zhao, Jizhou Huang +3

    cs.CVarXiv:1901.06829v12019
  6. CLIP-Dissect: Automatic Description of Neuron Representations in Deep Vision Networks

    Tuomas Oikarinen, Tsui-Wei Weng

    cs.CVcs.AIcs.LGarXiv:2204.10965v52022
  7. Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation

    Jiawei Mao, Haoqin Tu, Hardy Chen +8

    cs.CVarXiv:2609.06373v12026
  8. Hyper-SD: Trajectory Segmented Consistency Model for Efficient Image Synthesis

    Yuxi Ren, Xin Xia, Yanzuo Lu +5

    cs.CVarXiv:2404.13686v32024
  9. CODA: A Real-World Road Corner Case Dataset for Object Detection in Autonomous Driving

    Kaican Li, Kai Chen, Haoyu Wang +10

    cs.CVcs.LGcs.ROarXiv:2203.07724v32022
  10. Learnable PINs: Cross-Modal Embeddings for Person Identity

    Arsha Nagrani, Samuel Albanie, Andrew Zisserman

    cs.CVarXiv:1805.00833v22018
  11. SO-Pose: Exploiting Self-Occlusion for Direct 6D Pose Estimation

    Yan Di, Fabian Manhardt, Gu Wang +3

    cs.CVarXiv:2108.08367v12021
  12. Unsupervised Real-world Image Super Resolution via Domain-distance Aware Training

    Yunxuan Wei, Shuhang Gu, Yawei Li +1

    cs.CVarXiv:2004.01178v12020
  13. Learning Discriminative Representations for Skeleton Based Action Recognition

    Huanyu Zhou, Qingjie Liu, Yunhong Wang

    cs.CVarXiv:2303.03729v32023
  14. Region-based Quality Estimation Network for Large-scale Person Re-identification

    Guanglu Song, Biao Leng, Yu Liu +2

    cs.CVarXiv:1711.08766v22017
  15. NeuRAD: Neural Rendering for Autonomous Driving

    Adam Tonderski, Carl Lindström, Georg Hess +3

    cs.CVarXiv:2311.15260v32023
  16. Black-box Explanation of Object Detectors via Saliency Maps

    Vitali Petsiuk, Rajiv Jain, Varun Manjunatha +4

    cs.CVcs.AIcs.LGarXiv:2006.03204v22020
  17. 3D Human Motion Estimation via Motion Compression and Refinement

    Zhengyi Luo, S. Alireza Golestaneh, Kris M. Kitani

    cs.CVarXiv:2008.03789v22020
  18. Weakly- and Semi-Supervised Panoptic Segmentation

    Qizhu Li, Anurag Arnab, Philip H. S. Torr

    cs.CVarXiv:1808.03575v32018
  19. GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

    AgiBot Research Team, Renhang Liu, Wenzhi Zhao +42

    cs.ROcs.CVarXiv:2609.05588v12026
  20. Edit Probability for Scene Text Recognition

    Fan Bai, Zhanzhan Cheng, Yi Niu +2

    cs.CVarXiv:1805.03384v12018
  21. TARN: Temporal Attentive Relation Network for Few-Shot and Zero-Shot Action Recognition

    Mina Bishay, Georgios Zoumpourlis, Ioannis Patras

    cs.CVarXiv:1907.09021v12019
  22. Semantic-Guided Zero-Shot Learning for Low-Light Image/Video Enhancement

    Shen Zheng, Gaurav Gupta

    cs.CVarXiv:2110.00970v42021
  23. How to Exploit Hyperspherical Embeddings for Out-of-Distribution Detection?

    Yifei Ming, Yiyou Sun, Ousmane Dia +1

    cs.CVcs.LGarXiv:2203.04450v32022
  24. Epithelium segmentation using deep learning in H&E-stained prostate specimens with immunohistochemistry as reference standard

    Wouter Bulten, Péter Bándi, Jeffrey Hoven +7

    cs.CVarXiv:1808.05883v22018
  25. Gradient Step Denoiser for convergent Plug-and-Play

    Samuel Hurault, Arthur Leclaire, Nicolas Papadakis

    cs.CVeess.IVmath.OCarXiv:2110.03220v22021
  26. Anti-DreamBooth: Protecting users from personalized text-to-image synthesis

    Thanh Van Le, Hao Phung, Thuan Hoang Nguyen +3

    cs.CVcs.CRcs.LGarXiv:2303.15433v22023
  27. Unsupervised Depth Completion from Visual Inertial Odometry

    Alex Wong, Xiaohan Fei, Stephanie Tsuei +1

    cs.CVcs.AIcs.LGarXiv:1905.08616v42019
  28. PSGAN: Pose and Expression Robust Spatial-Aware GAN for Customizable Makeup Transfer

    Wentao Jiang, Si Liu, Chen Gao +4

    cs.CVarXiv:1909.06956v22019
  29. Improving Diffusion Inverse Problem Solving with Decoupled Noise Annealing

    Bingliang Zhang, Wenda Chu, Julius Berner +3

    cs.LGcs.AIcs.CVarXiv:2407.01521v32024
  30. RGB-D Object Detection and Semantic Segmentation for Autonomous Manipulation in Clutter

    Max Schwarz, Anton Milan, Arul Selvam Periyasamy +1

    cs.CVarXiv:1810.00818v12018
  31. SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution

    Xingjian Ran, Xiaoye Mo, Sihao Liu +3

    cs.CVarXiv:2609.05594v12026
  32. Everybody's Talkin': Let Me Talk as You Want

    Linsen Song, Wayne Wu, Chen Qian +2

    cs.CVcs.GRcs.MMarXiv:2001.05201v12020
  33. VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes

    Yan Ma, Jiadi Su, Zhulin Hu +4

    cs.CVarXiv:2609.06652v12026
  34. Joint Image Filtering with Deep Convolutional Networks

    Yijun Li, Jia-Bin Huang, Narendra Ahuja +1

    cs.CVarXiv:1710.04200v52017
  35. Self-Supervised Monocular Depth and Ego-Motion Estimation in Endoscopy: Appearance Flow to the Rescue

    Shuwei Shao, Zhongcai Pei, Weihai Chen +4

    cs.CVarXiv:2112.08122v12021
  36. An Autonomous Drone for Search and Rescue in Forests using Airborne Optical Sectioning

    D. C. Schedl, I. Kurmi, O. Bimber

    cs.CVarXiv:2105.04328v12021
  37. R-CNNs for Pose Estimation and Action Detection

    Georgia Gkioxari, Bharath Hariharan, Ross Girshick +1

    cs.CVarXiv:1406.5212v12014
  38. Hallucination Augmented Contrastive Learning for Multimodal Large Language Model

    Chaoya Jiang, Haiyang Xu, Mengfan Dong +7

    cs.CVarXiv:2312.06968v42023
  39. Agentic Visual Generation: From Generative Models to Agentic Control

    Yinming Huang, Shuyuan Tu, Xi Yan +7

    cs.CVarXiv:2609.06758v12026
  40. AFTer-UNet: Axial Fusion Transformer UNet for Medical Image Segmentation

    Xiangyi Yan, Hao Tang, Shanlin Sun +3

    eess.IVcs.CVcs.LGarXiv:2110.10403v12021
  41. LocalBins: Improving Depth Estimation by Learning Local Distributions

    Shariq Farooq Bhat, Ibraheem Alhashim, Peter Wonka

    cs.CVarXiv:2203.15132v12022
  42. Cluster-to-Conquer: A Framework for End-to-End Multi-Instance Learning for Whole Slide Image Classification

    Yash Sharma, Aman Shrivastava, Lubaina Ehsan +3

    eess.IVcs.CVcs.LGarXiv:2103.10626v22021
  43. ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions

    Chunlong Xia, Xinliang Wang, Feng Lv +2

    cs.CVarXiv:2403.07392v32024
  44. OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

    Ling Fu, Zhebin Kuang, Jiajun Song +21

    cs.CVcs.AIarXiv:2501.00321v22024
  45. Exploiting multi-CNN features in CNN-RNN based Dimensional Emotion Recognition on the OMG in-the-wild Dataset

    Dimitrios Kollias, Stefanos Zafeiriou

    cs.LGcs.CVstat.MLarXiv:1910.01417v22019
  46. A parameterless scale-space approach to find meaningful modes in histograms - Application to image and spectrum segmentation

    Jérôme Gilles, Kathryn Heal

    cs.CVarXiv:1401.2686v12014
  47. Hierarchical LSTM with Adjusted Temporal Attention for Video Captioning

    Jingkuan Song, Zhao Guo, Lianli Gao +3

    cs.CVarXiv:1706.01231v12017
  48. Deep-6DPose: Recovering 6D Object Pose from a Single RGB Image

    Thanh-Toan Do, Ming Cai, Trung Pham +1

    cs.CVcs.ROarXiv:1802.10367v12018
  49. The varifold representation of non-oriented shapes for diffeomorphic registration

    Nicolas Charon, Alain Trouvé

    cs.CGcs.CVmath.DGarXiv:1304.6108v12013
  50. DriveZero: End-to-End Driving Beyond Human Demonstrations

    Hao He, Chengcheng Hu, Zirun Su +17

    cs.CVarXiv:2609.06055v12026
  51. Guidelines and Evaluation of Clinical Explainable AI in Medical Image Analysis

    Weina Jin, Xiaoxiao Li, Mostafa Fatehi +1

    cs.LGcs.AIcs.CVarXiv:2202.10553v32022
  52. Stochastic Latent Residual Video Prediction

    Jean-Yves Franceschi, Edouard Delasalles, Mickaël Chen +2

    cs.CVcs.LGstat.MLarXiv:2002.09219v42020
  53. Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning

    Ye-Chan Kim, Seunghee Choi, SeungJu Cha +4

    cs.CVcs.AIarXiv:2609.04183v12026
  54. Fashion Forward: Forecasting Visual Style in Fashion

    Ziad Al-Halah, Rainer Stiefelhagen, Kristen Grauman

    cs.CVarXiv:1705.06394v32017
  55. A Theoretical Explanation for Perplexing Behaviors of Backpropagation-based Visualizations

    Weili Nie, Yang Zhang, Ankit Patel

    cs.CVcs.AIarXiv:1805.07039v42018
  56. Channel Splitting Network for Single MR Image Super-Resolution

    Xiaole Zhao, Yulun Zhang, Tao Zhang +1

    cs.CVarXiv:1810.06453v32018
  57. COVID_MTNet: COVID-19 Detection with Multi-Task Deep Learning Approaches

    Md Zahangir Alom, M M Shaifur Rahman, Mst Shamima Nasrin +2

    eess.IVcs.CVcs.LGarXiv:2004.03747v32020
  58. Binding Touch to Everything: Learning Unified Multimodal Tactile Representations

    Fengyu Yang, Chao Feng, Ziyang Chen +8

    cs.CVcs.ROarXiv:2401.18084v12024
  59. Discrimination-aware Network Pruning for Deep Model Compression

    Jing Liu, Bohan Zhuang, Zhuangwei Zhuang +4

    cs.CVarXiv:2001.01050v22020
  60. Experiments of Federated Learning for COVID-19 Chest X-ray Images

    Boyi Liu, Bingjie Yan, Yize Zhou +2

    eess.IVcs.CVcs.LGarXiv:2007.05592v12020