Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
2,821 to 2,880 of 18,815
EgoBody: Human Body Shape and Motion of Interacting People from Head-Mounted Devices
Siwei Zhang, Qianli Ma, Yan Zhang +5
cs.CVcs.AIarXiv:2112.07642v32021MinneApple: A Benchmark Dataset for Apple Detection and Segmentation
Nicolai Häni, Pravakar Roy, Volkan Isler
cs.CVarXiv:1909.06441v22019Learning Accurate Dense Correspondences and When to Trust Them
Prune Truong, Martin Danelljan, Luc Van Gool +1
cs.CVarXiv:2101.01710v22021CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements
Hongxiang Zhao, Mutian Xu, Zeyu Jin +3
cs.ROcs.CVarXiv:2609.07498v12026Deep Metric Learning for Practical Person Re-Identification
Dong Yi, Zhen Lei, Stan Z. Li
cs.CVcs.LGcs.NEarXiv:1407.4979v12014Talk-to-Edit: Fine-Grained Facial Editing via Dialog
Yuming Jiang, Ziqi Huang, Xingang Pan +2
cs.CVarXiv:2109.04425v12021HDR-NeRF: High Dynamic Range Neural Radiance Fields
Xin Huang, Qi Zhang, Ying Feng +3
cs.CVarXiv:2111.14451v42021Dense Cross-Query-and-Support Attention Weighted Mask Aggregation for Few-Shot Segmentation
Xinyu Shi, Dong Wei, Yu Zhang +5
cs.CVarXiv:2207.08549v12022Feature Erasing and Diffusion Network for Occluded Person Re-Identification
Zhikang Wang, Feng Zhu, Shixiang Tang +3
cs.CVarXiv:2112.08740v22021In-Context LoRA for Diffusion Transformers
Lianghua Huang, Wei Wang, Zhi-Fan Wu +6
cs.CVcs.GRarXiv:2410.23775v32024Fine-grained Visual Textual Alignment for Cross-Modal Retrieval using Transformer Encoders
Nicola Messina, Giuseppe Amato, Andrea Esuli +3
cs.CVarXiv:2008.05231v22020Read, Watch, and Move: Reinforcement Learning for Temporally Grounding Natural Language Descriptions in Videos
Dongliang He, Xiang Zhao, Jizhou Huang +3
cs.CVarXiv:1901.06829v12019CLIP-Dissect: Automatic Description of Neuron Representations in Deep Vision Networks
Tuomas Oikarinen, Tsui-Wei Weng
cs.CVcs.AIcs.LGarXiv:2204.10965v52022Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation
Jiawei Mao, Haoqin Tu, Hardy Chen +8
cs.CVarXiv:2609.06373v12026Hyper-SD: Trajectory Segmented Consistency Model for Efficient Image Synthesis
Yuxi Ren, Xin Xia, Yanzuo Lu +5
cs.CVarXiv:2404.13686v32024CODA: A Real-World Road Corner Case Dataset for Object Detection in Autonomous Driving
Kaican Li, Kai Chen, Haoyu Wang +10
cs.CVcs.LGcs.ROarXiv:2203.07724v32022Learnable PINs: Cross-Modal Embeddings for Person Identity
Arsha Nagrani, Samuel Albanie, Andrew Zisserman
cs.CVarXiv:1805.00833v22018SO-Pose: Exploiting Self-Occlusion for Direct 6D Pose Estimation
Yan Di, Fabian Manhardt, Gu Wang +3
cs.CVarXiv:2108.08367v12021Unsupervised Real-world Image Super Resolution via Domain-distance Aware Training
Yunxuan Wei, Shuhang Gu, Yawei Li +1
cs.CVarXiv:2004.01178v12020Learning Discriminative Representations for Skeleton Based Action Recognition
Huanyu Zhou, Qingjie Liu, Yunhong Wang
cs.CVarXiv:2303.03729v32023Region-based Quality Estimation Network for Large-scale Person Re-identification
Guanglu Song, Biao Leng, Yu Liu +2
cs.CVarXiv:1711.08766v22017NeuRAD: Neural Rendering for Autonomous Driving
Adam Tonderski, Carl Lindström, Georg Hess +3
cs.CVarXiv:2311.15260v32023Black-box Explanation of Object Detectors via Saliency Maps
Vitali Petsiuk, Rajiv Jain, Varun Manjunatha +4
cs.CVcs.AIcs.LGarXiv:2006.03204v220203D Human Motion Estimation via Motion Compression and Refinement
Zhengyi Luo, S. Alireza Golestaneh, Kris M. Kitani
cs.CVarXiv:2008.03789v22020Weakly- and Semi-Supervised Panoptic Segmentation
Qizhu Li, Anurag Arnab, Philip H. S. Torr
cs.CVarXiv:1808.03575v32018GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation
AgiBot Research Team, Renhang Liu, Wenzhi Zhao +42
cs.ROcs.CVarXiv:2609.05588v12026Edit Probability for Scene Text Recognition
Fan Bai, Zhanzhan Cheng, Yi Niu +2
cs.CVarXiv:1805.03384v12018TARN: Temporal Attentive Relation Network for Few-Shot and Zero-Shot Action Recognition
Mina Bishay, Georgios Zoumpourlis, Ioannis Patras
cs.CVarXiv:1907.09021v12019Semantic-Guided Zero-Shot Learning for Low-Light Image/Video Enhancement
Shen Zheng, Gaurav Gupta
cs.CVarXiv:2110.00970v42021How to Exploit Hyperspherical Embeddings for Out-of-Distribution Detection?
Yifei Ming, Yiyou Sun, Ousmane Dia +1
cs.CVcs.LGarXiv:2203.04450v32022Epithelium segmentation using deep learning in H&E-stained prostate specimens with immunohistochemistry as reference standard
Wouter Bulten, Péter Bándi, Jeffrey Hoven +7
cs.CVarXiv:1808.05883v22018Gradient Step Denoiser for convergent Plug-and-Play
Samuel Hurault, Arthur Leclaire, Nicolas Papadakis
cs.CVeess.IVmath.OCarXiv:2110.03220v22021Anti-DreamBooth: Protecting users from personalized text-to-image synthesis
Thanh Van Le, Hao Phung, Thuan Hoang Nguyen +3
cs.CVcs.CRcs.LGarXiv:2303.15433v22023Unsupervised Depth Completion from Visual Inertial Odometry
Alex Wong, Xiaohan Fei, Stephanie Tsuei +1
cs.CVcs.AIcs.LGarXiv:1905.08616v42019PSGAN: Pose and Expression Robust Spatial-Aware GAN for Customizable Makeup Transfer
Wentao Jiang, Si Liu, Chen Gao +4
cs.CVarXiv:1909.06956v22019Improving Diffusion Inverse Problem Solving with Decoupled Noise Annealing
Bingliang Zhang, Wenda Chu, Julius Berner +3
cs.LGcs.AIcs.CVarXiv:2407.01521v32024RGB-D Object Detection and Semantic Segmentation for Autonomous Manipulation in Clutter
Max Schwarz, Anton Milan, Arul Selvam Periyasamy +1
cs.CVarXiv:1810.00818v12018SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution
Xingjian Ran, Xiaoye Mo, Sihao Liu +3
cs.CVarXiv:2609.05594v12026Everybody's Talkin': Let Me Talk as You Want
Linsen Song, Wayne Wu, Chen Qian +2
cs.CVcs.GRcs.MMarXiv:2001.05201v12020VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes
Yan Ma, Jiadi Su, Zhulin Hu +4
cs.CVarXiv:2609.06652v12026Joint Image Filtering with Deep Convolutional Networks
Yijun Li, Jia-Bin Huang, Narendra Ahuja +1
cs.CVarXiv:1710.04200v52017Self-Supervised Monocular Depth and Ego-Motion Estimation in Endoscopy: Appearance Flow to the Rescue
Shuwei Shao, Zhongcai Pei, Weihai Chen +4
cs.CVarXiv:2112.08122v12021An Autonomous Drone for Search and Rescue in Forests using Airborne Optical Sectioning
D. C. Schedl, I. Kurmi, O. Bimber
cs.CVarXiv:2105.04328v12021R-CNNs for Pose Estimation and Action Detection
Georgia Gkioxari, Bharath Hariharan, Ross Girshick +1
cs.CVarXiv:1406.5212v12014Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
Chaoya Jiang, Haiyang Xu, Mengfan Dong +7
cs.CVarXiv:2312.06968v42023Agentic Visual Generation: From Generative Models to Agentic Control
Yinming Huang, Shuyuan Tu, Xi Yan +7
cs.CVarXiv:2609.06758v12026AFTer-UNet: Axial Fusion Transformer UNet for Medical Image Segmentation
Xiangyi Yan, Hao Tang, Shanlin Sun +3
eess.IVcs.CVcs.LGarXiv:2110.10403v12021LocalBins: Improving Depth Estimation by Learning Local Distributions
Shariq Farooq Bhat, Ibraheem Alhashim, Peter Wonka
cs.CVarXiv:2203.15132v12022Cluster-to-Conquer: A Framework for End-to-End Multi-Instance Learning for Whole Slide Image Classification
Yash Sharma, Aman Shrivastava, Lubaina Ehsan +3
eess.IVcs.CVcs.LGarXiv:2103.10626v22021ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions
Chunlong Xia, Xinliang Wang, Feng Lv +2
cs.CVarXiv:2403.07392v32024OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Ling Fu, Zhebin Kuang, Jiajun Song +21
cs.CVcs.AIarXiv:2501.00321v22024Exploiting multi-CNN features in CNN-RNN based Dimensional Emotion Recognition on the OMG in-the-wild Dataset
Dimitrios Kollias, Stefanos Zafeiriou
cs.LGcs.CVstat.MLarXiv:1910.01417v22019A parameterless scale-space approach to find meaningful modes in histograms - Application to image and spectrum segmentation
Jérôme Gilles, Kathryn Heal
cs.CVarXiv:1401.2686v12014Hierarchical LSTM with Adjusted Temporal Attention for Video Captioning
Jingkuan Song, Zhao Guo, Lianli Gao +3
cs.CVarXiv:1706.01231v12017Deep-6DPose: Recovering 6D Object Pose from a Single RGB Image
Thanh-Toan Do, Ming Cai, Trung Pham +1
cs.CVcs.ROarXiv:1802.10367v12018The varifold representation of non-oriented shapes for diffeomorphic registration
Nicolas Charon, Alain Trouvé
cs.CGcs.CVmath.DGarXiv:1304.6108v12013DriveZero: End-to-End Driving Beyond Human Demonstrations
Hao He, Chengcheng Hu, Zirun Su +17
cs.CVarXiv:2609.06055v12026Guidelines and Evaluation of Clinical Explainable AI in Medical Image Analysis
Weina Jin, Xiaoxiao Li, Mostafa Fatehi +1
cs.LGcs.AIcs.CVarXiv:2202.10553v32022Stochastic Latent Residual Video Prediction
Jean-Yves Franceschi, Edouard Delasalles, Mickaël Chen +2
cs.CVcs.LGstat.MLarXiv:2002.09219v42020Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning
Ye-Chan Kim, Seunghee Choi, SeungJu Cha +4
cs.CVcs.AIarXiv:2609.04183v12026