Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
2,941 to 3,000 of 18,943
Neural Architecture Transfer
Zhichao Lu, Gautam Sreekumar, Erik Goodman +3
cs.CVcs.LGcs.NEarXiv:2005.05859v22020Multi-Modal Transformer for Accelerated MR Imaging
Chun-Mei Feng, Yunlu Yan, Geng Chen +3
eess.IVcs.CVarXiv:2106.14248v32021Reason Through the Latent! Making Latent Visual Reasoning Necessary
Suhyeong Park, Junha Jung, Jaewoo Kang
cs.AIcs.CLcs.CVarXiv:2609.06746v12026Hard to Track Objects with Irregular Motions and Similar Appearances? Make It Easier by Buffering the Matching Space
Fan Yang, Shigeyuki Odashima, Shoichi Masui +1
cs.CVcs.MMarXiv:2211.14317v32022One Loss for All: Deep Hashing with a Single Cosine Similarity based Learning Objective
Jiun Tian Hoe, Kam Woh Ng, Tianyu Zhang +3
cs.CVcs.LGarXiv:2109.14449v12021Places205-VGGNet Models for Scene Recognition
Limin Wang, Sheng Guo, Weilin Huang +1
cs.CVarXiv:1508.01667v12015Boosting Adversarial Training with Hypersphere Embedding
Tianyu Pang, Xiao Yang, Yinpeng Dong +3
cs.LGcs.CRcs.CVarXiv:2002.08619v32020Minimizing Energy Consumption Leads to the Emergence of Gaits in Legged Robots
Zipeng Fu, Ashish Kumar, Jitendra Malik +1
cs.ROcs.AIcs.CVarXiv:2111.01674v12021EgoBody: Human Body Shape and Motion of Interacting People from Head-Mounted Devices
Siwei Zhang, Qianli Ma, Yan Zhang +5
cs.CVcs.AIarXiv:2112.07642v32021MinneApple: A Benchmark Dataset for Apple Detection and Segmentation
Nicolai Häni, Pravakar Roy, Volkan Isler
cs.CVarXiv:1909.06441v22019Learning Accurate Dense Correspondences and When to Trust Them
Prune Truong, Martin Danelljan, Luc Van Gool +1
cs.CVarXiv:2101.01710v22021CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements
Hongxiang Zhao, Mutian Xu, Zeyu Jin +3
cs.ROcs.CVarXiv:2609.07498v12026Deep Metric Learning for Practical Person Re-Identification
Dong Yi, Zhen Lei, Stan Z. Li
cs.CVcs.LGcs.NEarXiv:1407.4979v12014Talk-to-Edit: Fine-Grained Facial Editing via Dialog
Yuming Jiang, Ziqi Huang, Xingang Pan +2
cs.CVarXiv:2109.04425v12021HDR-NeRF: High Dynamic Range Neural Radiance Fields
Xin Huang, Qi Zhang, Ying Feng +3
cs.CVarXiv:2111.14451v42021Dense Cross-Query-and-Support Attention Weighted Mask Aggregation for Few-Shot Segmentation
Xinyu Shi, Dong Wei, Yu Zhang +5
cs.CVarXiv:2207.08549v12022Feature Erasing and Diffusion Network for Occluded Person Re-Identification
Zhikang Wang, Feng Zhu, Shixiang Tang +3
cs.CVarXiv:2112.08740v22021In-Context LoRA for Diffusion Transformers
Lianghua Huang, Wei Wang, Zhi-Fan Wu +6
cs.CVcs.GRarXiv:2410.23775v32024Fine-grained Visual Textual Alignment for Cross-Modal Retrieval using Transformer Encoders
Nicola Messina, Giuseppe Amato, Andrea Esuli +3
cs.CVarXiv:2008.05231v22020Read, Watch, and Move: Reinforcement Learning for Temporally Grounding Natural Language Descriptions in Videos
Dongliang He, Xiang Zhao, Jizhou Huang +3
cs.CVarXiv:1901.06829v12019CLIP-Dissect: Automatic Description of Neuron Representations in Deep Vision Networks
Tuomas Oikarinen, Tsui-Wei Weng
cs.CVcs.AIcs.LGarXiv:2204.10965v52022Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation
Jiawei Mao, Haoqin Tu, Hardy Chen +8
cs.CVarXiv:2609.06373v12026Hyper-SD: Trajectory Segmented Consistency Model for Efficient Image Synthesis
Yuxi Ren, Xin Xia, Yanzuo Lu +5
cs.CVarXiv:2404.13686v32024CODA: A Real-World Road Corner Case Dataset for Object Detection in Autonomous Driving
Kaican Li, Kai Chen, Haoyu Wang +10
cs.CVcs.LGcs.ROarXiv:2203.07724v32022Learnable PINs: Cross-Modal Embeddings for Person Identity
Arsha Nagrani, Samuel Albanie, Andrew Zisserman
cs.CVarXiv:1805.00833v22018SO-Pose: Exploiting Self-Occlusion for Direct 6D Pose Estimation
Yan Di, Fabian Manhardt, Gu Wang +3
cs.CVarXiv:2108.08367v12021Unsupervised Real-world Image Super Resolution via Domain-distance Aware Training
Yunxuan Wei, Shuhang Gu, Yawei Li +1
cs.CVarXiv:2004.01178v12020Learning Discriminative Representations for Skeleton Based Action Recognition
Huanyu Zhou, Qingjie Liu, Yunhong Wang
cs.CVarXiv:2303.03729v32023Region-based Quality Estimation Network for Large-scale Person Re-identification
Guanglu Song, Biao Leng, Yu Liu +2
cs.CVarXiv:1711.08766v22017NeuRAD: Neural Rendering for Autonomous Driving
Adam Tonderski, Carl Lindström, Georg Hess +3
cs.CVarXiv:2311.15260v32023Black-box Explanation of Object Detectors via Saliency Maps
Vitali Petsiuk, Rajiv Jain, Varun Manjunatha +4
cs.CVcs.AIcs.LGarXiv:2006.03204v220203D Human Motion Estimation via Motion Compression and Refinement
Zhengyi Luo, S. Alireza Golestaneh, Kris M. Kitani
cs.CVarXiv:2008.03789v22020Weakly- and Semi-Supervised Panoptic Segmentation
Qizhu Li, Anurag Arnab, Philip H. S. Torr
cs.CVarXiv:1808.03575v32018GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation
AgiBot Research Team, Renhang Liu, Wenzhi Zhao +42
cs.ROcs.CVarXiv:2609.05588v12026Edit Probability for Scene Text Recognition
Fan Bai, Zhanzhan Cheng, Yi Niu +2
cs.CVarXiv:1805.03384v12018TARN: Temporal Attentive Relation Network for Few-Shot and Zero-Shot Action Recognition
Mina Bishay, Georgios Zoumpourlis, Ioannis Patras
cs.CVarXiv:1907.09021v12019Semantic-Guided Zero-Shot Learning for Low-Light Image/Video Enhancement
Shen Zheng, Gaurav Gupta
cs.CVarXiv:2110.00970v42021How to Exploit Hyperspherical Embeddings for Out-of-Distribution Detection?
Yifei Ming, Yiyou Sun, Ousmane Dia +1
cs.CVcs.LGarXiv:2203.04450v32022Epithelium segmentation using deep learning in H&E-stained prostate specimens with immunohistochemistry as reference standard
Wouter Bulten, Péter Bándi, Jeffrey Hoven +7
cs.CVarXiv:1808.05883v22018Gradient Step Denoiser for convergent Plug-and-Play
Samuel Hurault, Arthur Leclaire, Nicolas Papadakis
cs.CVeess.IVmath.OCarXiv:2110.03220v22021Anti-DreamBooth: Protecting users from personalized text-to-image synthesis
Thanh Van Le, Hao Phung, Thuan Hoang Nguyen +3
cs.CVcs.CRcs.LGarXiv:2303.15433v22023Unsupervised Depth Completion from Visual Inertial Odometry
Alex Wong, Xiaohan Fei, Stephanie Tsuei +1
cs.CVcs.AIcs.LGarXiv:1905.08616v42019PSGAN: Pose and Expression Robust Spatial-Aware GAN for Customizable Makeup Transfer
Wentao Jiang, Si Liu, Chen Gao +4
cs.CVarXiv:1909.06956v22019Improving Diffusion Inverse Problem Solving with Decoupled Noise Annealing
Bingliang Zhang, Wenda Chu, Julius Berner +3
cs.LGcs.AIcs.CVarXiv:2407.01521v32024RGB-D Object Detection and Semantic Segmentation for Autonomous Manipulation in Clutter
Max Schwarz, Anton Milan, Arul Selvam Periyasamy +1
cs.CVarXiv:1810.00818v12018SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution
Xingjian Ran, Xiaoye Mo, Sihao Liu +3
cs.CVarXiv:2609.05594v12026Everybody's Talkin': Let Me Talk as You Want
Linsen Song, Wayne Wu, Chen Qian +2
cs.CVcs.GRcs.MMarXiv:2001.05201v12020VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes
Yan Ma, Jiadi Su, Zhulin Hu +4
cs.CVarXiv:2609.06652v12026Joint Image Filtering with Deep Convolutional Networks
Yijun Li, Jia-Bin Huang, Narendra Ahuja +1
cs.CVarXiv:1710.04200v52017Self-Supervised Monocular Depth and Ego-Motion Estimation in Endoscopy: Appearance Flow to the Rescue
Shuwei Shao, Zhongcai Pei, Weihai Chen +4
cs.CVarXiv:2112.08122v12021An Autonomous Drone for Search and Rescue in Forests using Airborne Optical Sectioning
D. C. Schedl, I. Kurmi, O. Bimber
cs.CVarXiv:2105.04328v12021R-CNNs for Pose Estimation and Action Detection
Georgia Gkioxari, Bharath Hariharan, Ross Girshick +1
cs.CVarXiv:1406.5212v12014Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
Chaoya Jiang, Haiyang Xu, Mengfan Dong +7
cs.CVarXiv:2312.06968v42023Agentic Visual Generation: From Generative Models to Agentic Control
Yinming Huang, Shuyuan Tu, Xi Yan +7
cs.CVarXiv:2609.06758v12026AFTer-UNet: Axial Fusion Transformer UNet for Medical Image Segmentation
Xiangyi Yan, Hao Tang, Shanlin Sun +3
eess.IVcs.CVcs.LGarXiv:2110.10403v12021LocalBins: Improving Depth Estimation by Learning Local Distributions
Shariq Farooq Bhat, Ibraheem Alhashim, Peter Wonka
cs.CVarXiv:2203.15132v12022Cluster-to-Conquer: A Framework for End-to-End Multi-Instance Learning for Whole Slide Image Classification
Yash Sharma, Aman Shrivastava, Lubaina Ehsan +3
eess.IVcs.CVcs.LGarXiv:2103.10626v22021ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions
Chunlong Xia, Xinliang Wang, Feng Lv +2
cs.CVarXiv:2403.07392v32024OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Ling Fu, Zhebin Kuang, Jiajun Song +21
cs.CVcs.AIarXiv:2501.00321v22024Exploiting multi-CNN features in CNN-RNN based Dimensional Emotion Recognition on the OMG in-the-wild Dataset
Dimitrios Kollias, Stefanos Zafeiriou
cs.LGcs.CVstat.MLarXiv:1910.01417v22019