Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

16,801 to 16,860 of 18,830

  1. Don't Decay the Learning Rate, Increase the Batch Size

    Samuel L. Smith, Pieter-Jan Kindermans, Chris Ying +1

    cs.LGcs.CVcs.DCarXiv:1711.00489v22017
  2. Rethinking ImageNet Pre-training

    Kaiming He, Ross Girshick, Piotr Dollár

    cs.CVarXiv:1811.08883v12018
  3. DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras

    Zachary Teed, Jia Deng

    cs.CVarXiv:2108.10869v22021
  4. Dynamic Network Surgery for Efficient DNNs

    Yiwen Guo, Anbang Yao, Yurong Chen

    cs.NEcs.CVcs.LGarXiv:1608.04493v22016
  5. MaPLe: Multi-modal Prompt Learning

    Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz +2

    cs.CVarXiv:2210.03117v32022
  6. iBOT: Image BERT Pre-Training with Online Tokenizer

    Jinghao Zhou, Chen Wei, Huiyu Wang +4

    cs.CVarXiv:2111.07832v32021
  7. Tversky loss function for image segmentation using 3D fully convolutional deep networks

    Seyed Sadegh Mohseni Salehi, Deniz Erdogmus, Ali Gholipour

    cs.CVarXiv:1706.05721v12017
  8. Scaling Vision with Sparse Mixture of Experts

    Carlos Riquelme, Joan Puigcerver, Basil Mustafa +5

    cs.CVcs.LGstat.MLarXiv:2106.05974v12021
  9. DETR3D: 3D Object Detection from Multi-view Images via 3D-to-2D Queries

    Yue Wang, Vitor Guizilini, Tianyuan Zhang +3

    cs.CVcs.AIcs.LGarXiv:2110.06922v12021
  10. CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

    Wenyi Hong, Ming Ding, Wendi Zheng +2

    cs.CVcs.CLcs.LGarXiv:2205.15868v12022
  11. Editing Models with Task Arithmetic

    Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman +4

    cs.LGcs.CLcs.CVarXiv:2212.04089v32022
  12. Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling

    Xumin Yu, Lulu Tang, Yongming Rao +3

    cs.CVcs.AIcs.LGarXiv:2111.14819v22021
  13. Depth-supervised NeRF: Fewer Views and Faster Training for Free

    Kangle Deng, Andrew Liu, Jun-Yan Zhu +1

    cs.CVcs.GRcs.LGarXiv:2107.02791v32021
  14. Soft Anisotropic Diagrams for Differentiable Image Representation

    Laki Iinbor, Zhiyang Dou, Wojciech Matusik

    cs.CVarXiv:2604.21984v22026
  15. SketchVLM: Vision language models can annotate images to explain thoughts and guide users

    Brandon Collins, Logan Bolton, Hung Huy Nguyen +3

    cs.CVcs.AIarXiv:2604.22875v22026
    Summaries:한국어
  16. Person Re-identification: Past, Present and Future

    Liang Zheng, Yi Yang, Alexander G. Hauptmann

    cs.CVarXiv:1610.02984v12016
  17. TALL: Temporal Activity Localization via Language Query

    Jiyang Gao, Chen Sun, Zhenheng Yang +1

    cs.CVarXiv:1705.02101v22017
  18. Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models

    Mohammed Safi Ur Rahman Khan, Sanjay Suryanarayanan, Tushar Anand +1

    cs.CVcs.CLarXiv:2604.21523v12026
  19. RaV-IDP: A Reconstruction-as-Validation Framework for Faithful Intelligent Document Processing

    Pritesh Jha

    cs.CVcs.AIarXiv:2604.23644v12026
  20. Deep Double Descent: Where Bigger Models and More Data Hurt

    Preetum Nakkiran, Gal Kaplun, Yamini Bansal +3

    cs.LGcs.CVcs.NEarXiv:1912.02292v12019
  21. ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks

    Mohit Shridhar, Jesse Thomason, Daniel Gordon +5

    cs.CVcs.AIcs.CLarXiv:1912.01734v22019
  22. Attention Augmented Convolutional Networks

    Irwan Bello, Barret Zoph, Ashish Vaswani +2

    cs.CVarXiv:1904.09925v52019
  23. Neural 3D Mesh Renderer

    Hiroharu Kato, Yoshitaka Ushiku, Tatsuya Harada

    cs.CVcs.LGarXiv:1711.07566v12017
  24. InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions

    Wenhai Wang, Jifeng Dai, Zhe Chen +9

    cs.CVarXiv:2211.05778v42022
  25. Probing Visual Planning in Image Editing Models

    Zhimu Zhou, Yanpeng Zhao, Qiuyu Liao +2

    cs.CVcs.AIarXiv:2604.22868v12026
  26. AnalogRetriever: Learning Cross-Modal Representations for Analog Circuit Retrieval

    Yihan Wang, Lei Li, Yao Lai +2

    cs.CVcs.AIarXiv:2604.23195v12026
  27. DeeperCut: A Deeper, Stronger, and Faster Multi-Person Pose Estimation Model

    Eldar Insafutdinov, Leonid Pishchulin, Bjoern Andres +2

    cs.CVarXiv:1605.03170v32016
  28. OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models

    Yida Xue, Ningyu Zhang, Tingwei Wu +5

    cs.MMcs.AIcs.CLarXiv:2605.00877v22026
  29. Data Augmentation Generative Adversarial Networks

    Antreas Antoniou, Amos Storkey, Harrison Edwards

    stat.MLcs.CVcs.LGarXiv:1711.04340v32017
  30. Learning to Identify Out-of-Distribution Objects for 3D LiDAR Anomaly Segmentation

    Simone Mosco, Daniel Fusaro, Alberto Pretto

    cs.CVcs.ROarXiv:2604.23604v12026
  31. VIBE: Video Inference for Human Body Pose and Shape Estimation

    Muhammed Kocabas, Nikos Athanasiou, Michael J. Black

    cs.CVarXiv:1912.05656v32019
  32. DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion

    Chen Wang, Danfei Xu, Yuke Zhu +4

    cs.CVcs.ROarXiv:1901.04780v12019
  33. Learning Spatio-Temporal Transformer for Visual Tracking

    Bin Yan, Houwen Peng, Jianlong Fu +2

    cs.CVarXiv:2103.17154v12021
  34. AutoGUI-v2: A Comprehensive Multi-Modal GUI Functionality Understanding Benchmark

    Hongxin Li, Xiping Wang, Jingran Su +4

    cs.CVarXiv:2604.24441v12026
  35. CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

    Huaishao Luo, Lei Ji, Ming Zhong +4

    cs.CVarXiv:2104.08860v22021
  36. GoClick: Lightweight Element Grounding Model for Autonomous GUI Interaction

    Hongxin Li, Yuntao Chen, Zhaoxiang Zhang

    cs.CVarXiv:2604.23941v12026
  37. StackGAN++: Realistic Image Synthesis with Stacked Generative Adversarial Networks

    Han Zhang, Tao Xu, Hongsheng Li +4

    cs.CVcs.AIstat.MLarXiv:1710.10916v32017
  38. ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

    Lin Chen, Jinsong Li, Xiaoyi Dong +5

    cs.CVarXiv:2311.12793v22023
  39. Symmetric Cross Entropy for Robust Learning with Noisy Labels

    Yisen Wang, Xingjun Ma, Zaiyi Chen +3

    cs.LGcs.CVstat.MLarXiv:1908.06112v12019
  40. Deep Learning for Medical Image Processing: Overview, Challenges and Future

    Muhammad Imran Razzak, Saeeda Naz, Ahmad Zaib

    cs.CVarXiv:1704.06825v12017
  41. MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding

    Aishwarya Kamath, Mannat Singh, Yann LeCun +3

    cs.CVcs.CLcs.LGarXiv:2104.12763v22021
  42. FDA: Fourier Domain Adaptation for Semantic Segmentation

    Yanchao Yang, Stefano Soatto

    cs.CVarXiv:2004.05498v12020
  43. Voxel R-CNN: Towards High Performance Voxel-based 3D Object Detection

    Jiajun Deng, Shaoshuai Shi, Peiwei Li +3

    cs.CVarXiv:2012.15712v22020
  44. X2SAM: Any Segmentation in Images and Videos

    Hao Wang, Limeng Qiao, Chi Zhang +4

    cs.CVcs.AIarXiv:2605.00891v12026
  45. Accelerating 3D Deep Learning with PyTorch3D

    Nikhila Ravi, Jeremy Reizenstein, David Novotny +4

    cs.CVcs.GRcs.LGarXiv:2007.08501v12020
  46. FcaNet: Frequency Channel Attention Networks

    Zequn Qin, Pengyi Zhang, Fei Wu +1

    cs.CVarXiv:2012.11879v42020
  47. Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models

    Jiayi Guo, Linqing Wang, Jiangshan Wang +6

    cs.CVarXiv:2604.25636v12026
  48. Importance Estimation for Neural Network Pruning

    Pavlo Molchanov, Arun Mallya, Stephen Tyree +2

    cs.LGcs.CVstat.MLarXiv:1906.10771v12019
  49. RADIO-ViPE: Online Tightly Coupled Multi-Modal Fusion for Open-Vocabulary Semantic SLAM in Dynamic Environments

    Zaid Nasser, Mikhail Iumanov, Tianhao Li +3

    cs.CVarXiv:2604.26067v12026
  50. Learning to Reconstruct 3D Human Pose and Shape via Model-fitting in the Loop

    Nikos Kolotouros, Georgios Pavlakos, Michael J. Black +1

    cs.CVarXiv:1909.12828v12019
  51. Pseudo-LiDAR from Visual Depth Estimation: Bridging the Gap in 3D Object Detection for Autonomous Driving

    Yan Wang, Wei-Lun Chao, Divyansh Garg +3

    cs.CVarXiv:1812.07179v62018
  52. FASH-iCNN: Making Editorial Fashion Identity Inspectable Through Multimodal CNN Probing

    Morayo Danielle Adeyemi, Ryan A. Rossi, Franck Dernoncourt

    cs.CVcs.HCcs.IRarXiv:2604.26186v12026
  53. Structural-RNN: Deep Learning on Spatio-Temporal Graphs

    Ashesh Jain, Amir R. Zamir, Silvio Savarese +1

    cs.CVcs.LGcs.NEarXiv:1511.05298v32015
  54. From Coarse to Fine: Robust Hierarchical Localization at Large Scale

    Paul-Edouard Sarlin, Cesar Cadena, Roland Siegwart +1

    cs.CVarXiv:1812.03506v22018
  55. Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation

    Jay Zhangjie Wu, Yixiao Ge, Xintao Wang +7

    cs.CVarXiv:2212.11565v22022
  56. Action Recognition with Trajectory-Pooled Deep-Convolutional Descriptors

    Limin Wang, Yu Qiao, Xiaoou Tang

    cs.CVarXiv:1505.04868v12015
  57. DRAEM -- A discriminatively trained reconstruction embedding for surface anomaly detection

    Vitjan Zavrtanik, Matej Kristan, Danijel Skočaj

    cs.CVarXiv:2108.07610v22021
  58. Image De-raining Using a Conditional Generative Adversarial Network

    He Zhang, Vishwanath Sindagi, Vishal M. Patel

    cs.CVarXiv:1701.05957v42017
  59. Oriented R-CNN for Object Detection

    Xingxing Xie, Gong Cheng, Jiabao Wang +2

    cs.CVarXiv:2108.05699v12021
  60. Neural Module Networks

    Jacob Andreas, Marcus Rohrbach, Trevor Darrell +1

    cs.CVcs.CLcs.LGarXiv:1511.02799v42015