Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

2,941 to 3,000 of 18,943

  1. Neural Architecture Transfer

    Zhichao Lu, Gautam Sreekumar, Erik Goodman +3

    cs.CVcs.LGcs.NEarXiv:2005.05859v22020
  2. Multi-Modal Transformer for Accelerated MR Imaging

    Chun-Mei Feng, Yunlu Yan, Geng Chen +3

    eess.IVcs.CVarXiv:2106.14248v32021
  3. Reason Through the Latent! Making Latent Visual Reasoning Necessary

    Suhyeong Park, Junha Jung, Jaewoo Kang

    cs.AIcs.CLcs.CVarXiv:2609.06746v12026
  4. Hard to Track Objects with Irregular Motions and Similar Appearances? Make It Easier by Buffering the Matching Space

    Fan Yang, Shigeyuki Odashima, Shoichi Masui +1

    cs.CVcs.MMarXiv:2211.14317v32022
  5. One Loss for All: Deep Hashing with a Single Cosine Similarity based Learning Objective

    Jiun Tian Hoe, Kam Woh Ng, Tianyu Zhang +3

    cs.CVcs.LGarXiv:2109.14449v12021
  6. Places205-VGGNet Models for Scene Recognition

    Limin Wang, Sheng Guo, Weilin Huang +1

    cs.CVarXiv:1508.01667v12015
  7. Boosting Adversarial Training with Hypersphere Embedding

    Tianyu Pang, Xiao Yang, Yinpeng Dong +3

    cs.LGcs.CRcs.CVarXiv:2002.08619v32020
  8. Minimizing Energy Consumption Leads to the Emergence of Gaits in Legged Robots

    Zipeng Fu, Ashish Kumar, Jitendra Malik +1

    cs.ROcs.AIcs.CVarXiv:2111.01674v12021
  9. EgoBody: Human Body Shape and Motion of Interacting People from Head-Mounted Devices

    Siwei Zhang, Qianli Ma, Yan Zhang +5

    cs.CVcs.AIarXiv:2112.07642v32021
  10. MinneApple: A Benchmark Dataset for Apple Detection and Segmentation

    Nicolai Häni, Pravakar Roy, Volkan Isler

    cs.CVarXiv:1909.06441v22019
  11. Learning Accurate Dense Correspondences and When to Trust Them

    Prune Truong, Martin Danelljan, Luc Van Gool +1

    cs.CVarXiv:2101.01710v22021
  12. CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements

    Hongxiang Zhao, Mutian Xu, Zeyu Jin +3

    cs.ROcs.CVarXiv:2609.07498v12026
  13. Deep Metric Learning for Practical Person Re-Identification

    Dong Yi, Zhen Lei, Stan Z. Li

    cs.CVcs.LGcs.NEarXiv:1407.4979v12014
  14. Talk-to-Edit: Fine-Grained Facial Editing via Dialog

    Yuming Jiang, Ziqi Huang, Xingang Pan +2

    cs.CVarXiv:2109.04425v12021
  15. HDR-NeRF: High Dynamic Range Neural Radiance Fields

    Xin Huang, Qi Zhang, Ying Feng +3

    cs.CVarXiv:2111.14451v42021
  16. Dense Cross-Query-and-Support Attention Weighted Mask Aggregation for Few-Shot Segmentation

    Xinyu Shi, Dong Wei, Yu Zhang +5

    cs.CVarXiv:2207.08549v12022
  17. Feature Erasing and Diffusion Network for Occluded Person Re-Identification

    Zhikang Wang, Feng Zhu, Shixiang Tang +3

    cs.CVarXiv:2112.08740v22021
  18. In-Context LoRA for Diffusion Transformers

    Lianghua Huang, Wei Wang, Zhi-Fan Wu +6

    cs.CVcs.GRarXiv:2410.23775v32024
  19. Fine-grained Visual Textual Alignment for Cross-Modal Retrieval using Transformer Encoders

    Nicola Messina, Giuseppe Amato, Andrea Esuli +3

    cs.CVarXiv:2008.05231v22020
  20. Read, Watch, and Move: Reinforcement Learning for Temporally Grounding Natural Language Descriptions in Videos

    Dongliang He, Xiang Zhao, Jizhou Huang +3

    cs.CVarXiv:1901.06829v12019
  21. CLIP-Dissect: Automatic Description of Neuron Representations in Deep Vision Networks

    Tuomas Oikarinen, Tsui-Wei Weng

    cs.CVcs.AIcs.LGarXiv:2204.10965v52022
  22. Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation

    Jiawei Mao, Haoqin Tu, Hardy Chen +8

    cs.CVarXiv:2609.06373v12026
  23. Hyper-SD: Trajectory Segmented Consistency Model for Efficient Image Synthesis

    Yuxi Ren, Xin Xia, Yanzuo Lu +5

    cs.CVarXiv:2404.13686v32024
  24. CODA: A Real-World Road Corner Case Dataset for Object Detection in Autonomous Driving

    Kaican Li, Kai Chen, Haoyu Wang +10

    cs.CVcs.LGcs.ROarXiv:2203.07724v32022
  25. Learnable PINs: Cross-Modal Embeddings for Person Identity

    Arsha Nagrani, Samuel Albanie, Andrew Zisserman

    cs.CVarXiv:1805.00833v22018
  26. SO-Pose: Exploiting Self-Occlusion for Direct 6D Pose Estimation

    Yan Di, Fabian Manhardt, Gu Wang +3

    cs.CVarXiv:2108.08367v12021
  27. Unsupervised Real-world Image Super Resolution via Domain-distance Aware Training

    Yunxuan Wei, Shuhang Gu, Yawei Li +1

    cs.CVarXiv:2004.01178v12020
  28. Learning Discriminative Representations for Skeleton Based Action Recognition

    Huanyu Zhou, Qingjie Liu, Yunhong Wang

    cs.CVarXiv:2303.03729v32023
  29. Region-based Quality Estimation Network for Large-scale Person Re-identification

    Guanglu Song, Biao Leng, Yu Liu +2

    cs.CVarXiv:1711.08766v22017
  30. NeuRAD: Neural Rendering for Autonomous Driving

    Adam Tonderski, Carl Lindström, Georg Hess +3

    cs.CVarXiv:2311.15260v32023
  31. Black-box Explanation of Object Detectors via Saliency Maps

    Vitali Petsiuk, Rajiv Jain, Varun Manjunatha +4

    cs.CVcs.AIcs.LGarXiv:2006.03204v22020
  32. 3D Human Motion Estimation via Motion Compression and Refinement

    Zhengyi Luo, S. Alireza Golestaneh, Kris M. Kitani

    cs.CVarXiv:2008.03789v22020
  33. Weakly- and Semi-Supervised Panoptic Segmentation

    Qizhu Li, Anurag Arnab, Philip H. S. Torr

    cs.CVarXiv:1808.03575v32018
  34. GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

    AgiBot Research Team, Renhang Liu, Wenzhi Zhao +42

    cs.ROcs.CVarXiv:2609.05588v12026
  35. Edit Probability for Scene Text Recognition

    Fan Bai, Zhanzhan Cheng, Yi Niu +2

    cs.CVarXiv:1805.03384v12018
  36. TARN: Temporal Attentive Relation Network for Few-Shot and Zero-Shot Action Recognition

    Mina Bishay, Georgios Zoumpourlis, Ioannis Patras

    cs.CVarXiv:1907.09021v12019
  37. Semantic-Guided Zero-Shot Learning for Low-Light Image/Video Enhancement

    Shen Zheng, Gaurav Gupta

    cs.CVarXiv:2110.00970v42021
  38. How to Exploit Hyperspherical Embeddings for Out-of-Distribution Detection?

    Yifei Ming, Yiyou Sun, Ousmane Dia +1

    cs.CVcs.LGarXiv:2203.04450v32022
  39. Epithelium segmentation using deep learning in H&E-stained prostate specimens with immunohistochemistry as reference standard

    Wouter Bulten, Péter Bándi, Jeffrey Hoven +7

    cs.CVarXiv:1808.05883v22018
  40. Gradient Step Denoiser for convergent Plug-and-Play

    Samuel Hurault, Arthur Leclaire, Nicolas Papadakis

    cs.CVeess.IVmath.OCarXiv:2110.03220v22021
  41. Anti-DreamBooth: Protecting users from personalized text-to-image synthesis

    Thanh Van Le, Hao Phung, Thuan Hoang Nguyen +3

    cs.CVcs.CRcs.LGarXiv:2303.15433v22023
  42. Unsupervised Depth Completion from Visual Inertial Odometry

    Alex Wong, Xiaohan Fei, Stephanie Tsuei +1

    cs.CVcs.AIcs.LGarXiv:1905.08616v42019
  43. PSGAN: Pose and Expression Robust Spatial-Aware GAN for Customizable Makeup Transfer

    Wentao Jiang, Si Liu, Chen Gao +4

    cs.CVarXiv:1909.06956v22019
  44. Improving Diffusion Inverse Problem Solving with Decoupled Noise Annealing

    Bingliang Zhang, Wenda Chu, Julius Berner +3

    cs.LGcs.AIcs.CVarXiv:2407.01521v32024
  45. RGB-D Object Detection and Semantic Segmentation for Autonomous Manipulation in Clutter

    Max Schwarz, Anton Milan, Arul Selvam Periyasamy +1

    cs.CVarXiv:1810.00818v12018
  46. SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution

    Xingjian Ran, Xiaoye Mo, Sihao Liu +3

    cs.CVarXiv:2609.05594v12026
  47. Everybody's Talkin': Let Me Talk as You Want

    Linsen Song, Wayne Wu, Chen Qian +2

    cs.CVcs.GRcs.MMarXiv:2001.05201v12020
  48. VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes

    Yan Ma, Jiadi Su, Zhulin Hu +4

    cs.CVarXiv:2609.06652v12026
  49. Joint Image Filtering with Deep Convolutional Networks

    Yijun Li, Jia-Bin Huang, Narendra Ahuja +1

    cs.CVarXiv:1710.04200v52017
  50. Self-Supervised Monocular Depth and Ego-Motion Estimation in Endoscopy: Appearance Flow to the Rescue

    Shuwei Shao, Zhongcai Pei, Weihai Chen +4

    cs.CVarXiv:2112.08122v12021
  51. An Autonomous Drone for Search and Rescue in Forests using Airborne Optical Sectioning

    D. C. Schedl, I. Kurmi, O. Bimber

    cs.CVarXiv:2105.04328v12021
  52. R-CNNs for Pose Estimation and Action Detection

    Georgia Gkioxari, Bharath Hariharan, Ross Girshick +1

    cs.CVarXiv:1406.5212v12014
  53. Hallucination Augmented Contrastive Learning for Multimodal Large Language Model

    Chaoya Jiang, Haiyang Xu, Mengfan Dong +7

    cs.CVarXiv:2312.06968v42023
  54. Agentic Visual Generation: From Generative Models to Agentic Control

    Yinming Huang, Shuyuan Tu, Xi Yan +7

    cs.CVarXiv:2609.06758v12026
  55. AFTer-UNet: Axial Fusion Transformer UNet for Medical Image Segmentation

    Xiangyi Yan, Hao Tang, Shanlin Sun +3

    eess.IVcs.CVcs.LGarXiv:2110.10403v12021
  56. LocalBins: Improving Depth Estimation by Learning Local Distributions

    Shariq Farooq Bhat, Ibraheem Alhashim, Peter Wonka

    cs.CVarXiv:2203.15132v12022
  57. Cluster-to-Conquer: A Framework for End-to-End Multi-Instance Learning for Whole Slide Image Classification

    Yash Sharma, Aman Shrivastava, Lubaina Ehsan +3

    eess.IVcs.CVcs.LGarXiv:2103.10626v22021
  58. ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions

    Chunlong Xia, Xinliang Wang, Feng Lv +2

    cs.CVarXiv:2403.07392v32024
  59. OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

    Ling Fu, Zhebin Kuang, Jiajun Song +21

    cs.CVcs.AIarXiv:2501.00321v22024
  60. Exploiting multi-CNN features in CNN-RNN based Dimensional Emotion Recognition on the OMG in-the-wild Dataset

    Dimitrios Kollias, Stefanos Zafeiriou

    cs.LGcs.CVstat.MLarXiv:1910.01417v22019