Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

2,821 to 2,880 of 18,815

  1. EgoBody: Human Body Shape and Motion of Interacting People from Head-Mounted Devices

    Siwei Zhang, Qianli Ma, Yan Zhang +5

    cs.CVcs.AIarXiv:2112.07642v32021
  2. MinneApple: A Benchmark Dataset for Apple Detection and Segmentation

    Nicolai Häni, Pravakar Roy, Volkan Isler

    cs.CVarXiv:1909.06441v22019
  3. Learning Accurate Dense Correspondences and When to Trust Them

    Prune Truong, Martin Danelljan, Luc Van Gool +1

    cs.CVarXiv:2101.01710v22021
  4. CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements

    Hongxiang Zhao, Mutian Xu, Zeyu Jin +3

    cs.ROcs.CVarXiv:2609.07498v12026
  5. Deep Metric Learning for Practical Person Re-Identification

    Dong Yi, Zhen Lei, Stan Z. Li

    cs.CVcs.LGcs.NEarXiv:1407.4979v12014
  6. Talk-to-Edit: Fine-Grained Facial Editing via Dialog

    Yuming Jiang, Ziqi Huang, Xingang Pan +2

    cs.CVarXiv:2109.04425v12021
  7. HDR-NeRF: High Dynamic Range Neural Radiance Fields

    Xin Huang, Qi Zhang, Ying Feng +3

    cs.CVarXiv:2111.14451v42021
  8. Dense Cross-Query-and-Support Attention Weighted Mask Aggregation for Few-Shot Segmentation

    Xinyu Shi, Dong Wei, Yu Zhang +5

    cs.CVarXiv:2207.08549v12022
  9. Feature Erasing and Diffusion Network for Occluded Person Re-Identification

    Zhikang Wang, Feng Zhu, Shixiang Tang +3

    cs.CVarXiv:2112.08740v22021
  10. In-Context LoRA for Diffusion Transformers

    Lianghua Huang, Wei Wang, Zhi-Fan Wu +6

    cs.CVcs.GRarXiv:2410.23775v32024
  11. Fine-grained Visual Textual Alignment for Cross-Modal Retrieval using Transformer Encoders

    Nicola Messina, Giuseppe Amato, Andrea Esuli +3

    cs.CVarXiv:2008.05231v22020
  12. Read, Watch, and Move: Reinforcement Learning for Temporally Grounding Natural Language Descriptions in Videos

    Dongliang He, Xiang Zhao, Jizhou Huang +3

    cs.CVarXiv:1901.06829v12019
  13. CLIP-Dissect: Automatic Description of Neuron Representations in Deep Vision Networks

    Tuomas Oikarinen, Tsui-Wei Weng

    cs.CVcs.AIcs.LGarXiv:2204.10965v52022
  14. Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation

    Jiawei Mao, Haoqin Tu, Hardy Chen +8

    cs.CVarXiv:2609.06373v12026
  15. Hyper-SD: Trajectory Segmented Consistency Model for Efficient Image Synthesis

    Yuxi Ren, Xin Xia, Yanzuo Lu +5

    cs.CVarXiv:2404.13686v32024
  16. CODA: A Real-World Road Corner Case Dataset for Object Detection in Autonomous Driving

    Kaican Li, Kai Chen, Haoyu Wang +10

    cs.CVcs.LGcs.ROarXiv:2203.07724v32022
  17. Learnable PINs: Cross-Modal Embeddings for Person Identity

    Arsha Nagrani, Samuel Albanie, Andrew Zisserman

    cs.CVarXiv:1805.00833v22018
  18. SO-Pose: Exploiting Self-Occlusion for Direct 6D Pose Estimation

    Yan Di, Fabian Manhardt, Gu Wang +3

    cs.CVarXiv:2108.08367v12021
  19. Unsupervised Real-world Image Super Resolution via Domain-distance Aware Training

    Yunxuan Wei, Shuhang Gu, Yawei Li +1

    cs.CVarXiv:2004.01178v12020
  20. Learning Discriminative Representations for Skeleton Based Action Recognition

    Huanyu Zhou, Qingjie Liu, Yunhong Wang

    cs.CVarXiv:2303.03729v32023
  21. Region-based Quality Estimation Network for Large-scale Person Re-identification

    Guanglu Song, Biao Leng, Yu Liu +2

    cs.CVarXiv:1711.08766v22017
  22. NeuRAD: Neural Rendering for Autonomous Driving

    Adam Tonderski, Carl Lindström, Georg Hess +3

    cs.CVarXiv:2311.15260v32023
  23. Black-box Explanation of Object Detectors via Saliency Maps

    Vitali Petsiuk, Rajiv Jain, Varun Manjunatha +4

    cs.CVcs.AIcs.LGarXiv:2006.03204v22020
  24. 3D Human Motion Estimation via Motion Compression and Refinement

    Zhengyi Luo, S. Alireza Golestaneh, Kris M. Kitani

    cs.CVarXiv:2008.03789v22020
  25. Weakly- and Semi-Supervised Panoptic Segmentation

    Qizhu Li, Anurag Arnab, Philip H. S. Torr

    cs.CVarXiv:1808.03575v32018
  26. GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

    AgiBot Research Team, Renhang Liu, Wenzhi Zhao +42

    cs.ROcs.CVarXiv:2609.05588v12026
  27. Edit Probability for Scene Text Recognition

    Fan Bai, Zhanzhan Cheng, Yi Niu +2

    cs.CVarXiv:1805.03384v12018
  28. TARN: Temporal Attentive Relation Network for Few-Shot and Zero-Shot Action Recognition

    Mina Bishay, Georgios Zoumpourlis, Ioannis Patras

    cs.CVarXiv:1907.09021v12019
  29. Semantic-Guided Zero-Shot Learning for Low-Light Image/Video Enhancement

    Shen Zheng, Gaurav Gupta

    cs.CVarXiv:2110.00970v42021
  30. How to Exploit Hyperspherical Embeddings for Out-of-Distribution Detection?

    Yifei Ming, Yiyou Sun, Ousmane Dia +1

    cs.CVcs.LGarXiv:2203.04450v32022
  31. Epithelium segmentation using deep learning in H&E-stained prostate specimens with immunohistochemistry as reference standard

    Wouter Bulten, Péter Bándi, Jeffrey Hoven +7

    cs.CVarXiv:1808.05883v22018
  32. Gradient Step Denoiser for convergent Plug-and-Play

    Samuel Hurault, Arthur Leclaire, Nicolas Papadakis

    cs.CVeess.IVmath.OCarXiv:2110.03220v22021
  33. Anti-DreamBooth: Protecting users from personalized text-to-image synthesis

    Thanh Van Le, Hao Phung, Thuan Hoang Nguyen +3

    cs.CVcs.CRcs.LGarXiv:2303.15433v22023
  34. Unsupervised Depth Completion from Visual Inertial Odometry

    Alex Wong, Xiaohan Fei, Stephanie Tsuei +1

    cs.CVcs.AIcs.LGarXiv:1905.08616v42019
  35. PSGAN: Pose and Expression Robust Spatial-Aware GAN for Customizable Makeup Transfer

    Wentao Jiang, Si Liu, Chen Gao +4

    cs.CVarXiv:1909.06956v22019
  36. Improving Diffusion Inverse Problem Solving with Decoupled Noise Annealing

    Bingliang Zhang, Wenda Chu, Julius Berner +3

    cs.LGcs.AIcs.CVarXiv:2407.01521v32024
  37. RGB-D Object Detection and Semantic Segmentation for Autonomous Manipulation in Clutter

    Max Schwarz, Anton Milan, Arul Selvam Periyasamy +1

    cs.CVarXiv:1810.00818v12018
  38. SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution

    Xingjian Ran, Xiaoye Mo, Sihao Liu +3

    cs.CVarXiv:2609.05594v12026
  39. Everybody's Talkin': Let Me Talk as You Want

    Linsen Song, Wayne Wu, Chen Qian +2

    cs.CVcs.GRcs.MMarXiv:2001.05201v12020
  40. VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes

    Yan Ma, Jiadi Su, Zhulin Hu +4

    cs.CVarXiv:2609.06652v12026
  41. Joint Image Filtering with Deep Convolutional Networks

    Yijun Li, Jia-Bin Huang, Narendra Ahuja +1

    cs.CVarXiv:1710.04200v52017
  42. Self-Supervised Monocular Depth and Ego-Motion Estimation in Endoscopy: Appearance Flow to the Rescue

    Shuwei Shao, Zhongcai Pei, Weihai Chen +4

    cs.CVarXiv:2112.08122v12021
  43. An Autonomous Drone for Search and Rescue in Forests using Airborne Optical Sectioning

    D. C. Schedl, I. Kurmi, O. Bimber

    cs.CVarXiv:2105.04328v12021
  44. R-CNNs for Pose Estimation and Action Detection

    Georgia Gkioxari, Bharath Hariharan, Ross Girshick +1

    cs.CVarXiv:1406.5212v12014
  45. Hallucination Augmented Contrastive Learning for Multimodal Large Language Model

    Chaoya Jiang, Haiyang Xu, Mengfan Dong +7

    cs.CVarXiv:2312.06968v42023
  46. Agentic Visual Generation: From Generative Models to Agentic Control

    Yinming Huang, Shuyuan Tu, Xi Yan +7

    cs.CVarXiv:2609.06758v12026
  47. AFTer-UNet: Axial Fusion Transformer UNet for Medical Image Segmentation

    Xiangyi Yan, Hao Tang, Shanlin Sun +3

    eess.IVcs.CVcs.LGarXiv:2110.10403v12021
  48. LocalBins: Improving Depth Estimation by Learning Local Distributions

    Shariq Farooq Bhat, Ibraheem Alhashim, Peter Wonka

    cs.CVarXiv:2203.15132v12022
  49. Cluster-to-Conquer: A Framework for End-to-End Multi-Instance Learning for Whole Slide Image Classification

    Yash Sharma, Aman Shrivastava, Lubaina Ehsan +3

    eess.IVcs.CVcs.LGarXiv:2103.10626v22021
  50. ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions

    Chunlong Xia, Xinliang Wang, Feng Lv +2

    cs.CVarXiv:2403.07392v32024
  51. OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

    Ling Fu, Zhebin Kuang, Jiajun Song +21

    cs.CVcs.AIarXiv:2501.00321v22024
  52. Exploiting multi-CNN features in CNN-RNN based Dimensional Emotion Recognition on the OMG in-the-wild Dataset

    Dimitrios Kollias, Stefanos Zafeiriou

    cs.LGcs.CVstat.MLarXiv:1910.01417v22019
  53. A parameterless scale-space approach to find meaningful modes in histograms - Application to image and spectrum segmentation

    Jérôme Gilles, Kathryn Heal

    cs.CVarXiv:1401.2686v12014
  54. Hierarchical LSTM with Adjusted Temporal Attention for Video Captioning

    Jingkuan Song, Zhao Guo, Lianli Gao +3

    cs.CVarXiv:1706.01231v12017
  55. Deep-6DPose: Recovering 6D Object Pose from a Single RGB Image

    Thanh-Toan Do, Ming Cai, Trung Pham +1

    cs.CVcs.ROarXiv:1802.10367v12018
  56. The varifold representation of non-oriented shapes for diffeomorphic registration

    Nicolas Charon, Alain Trouvé

    cs.CGcs.CVmath.DGarXiv:1304.6108v12013
  57. DriveZero: End-to-End Driving Beyond Human Demonstrations

    Hao He, Chengcheng Hu, Zirun Su +17

    cs.CVarXiv:2609.06055v12026
  58. Guidelines and Evaluation of Clinical Explainable AI in Medical Image Analysis

    Weina Jin, Xiaoxiao Li, Mostafa Fatehi +1

    cs.LGcs.AIcs.CVarXiv:2202.10553v32022
  59. Stochastic Latent Residual Video Prediction

    Jean-Yves Franceschi, Edouard Delasalles, Mickaël Chen +2

    cs.CVcs.LGstat.MLarXiv:2002.09219v42020
  60. Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning

    Ye-Chan Kim, Seunghee Choi, SeungJu Cha +4

    cs.CVcs.AIarXiv:2609.04183v12026