Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

16,921 to 16,980 of 18,855

  1. Transformers in Medical Imaging: A Survey

    Fahad Shamshad, Salman Khan, Syed Waqas Zamir +4

    eess.IVcs.CVarXiv:2201.09873v12022
  2. Compressing Deep Convolutional Networks using Vector Quantization

    Yunchao Gong, Liu Liu, Ming Yang +1

    cs.CVcs.LGcs.NEarXiv:1412.6115v12014
  3. AdaBins: Depth Estimation using Adaptive Bins

    Shariq Farooq Bhat, Ibraheem Alhashim, Peter Wonka

    cs.CVarXiv:2011.14141v12020
  4. Meta-Transfer Learning for Few-Shot Learning

    Qianru Sun, Yaoyao Liu, Tat-Seng Chua +1

    cs.CVarXiv:1812.02391v32018
  5. Quantized Convolutional Neural Networks for Mobile Devices

    Jiaxiang Wu, Cong Leng, Yuhang Wang +2

    cs.CVarXiv:1512.06473v32015
  6. MiniCPM-V: A GPT-4V Level MLLM on Your Phone

    Yuan Yao, Tianyu Yu, Ao Zhang +20

    cs.CVarXiv:2408.01800v12024
  7. NeRF++: Analyzing and Improving Neural Radiance Fields

    Kai Zhang, Gernot Riegler, Noah Snavely +1

    cs.CVarXiv:2010.07492v22020
  8. Tensor field networks: Rotation- and translation-equivariant neural networks for 3D point clouds

    Nathaniel Thomas, Tess Smidt, Steven Kearnes +4

    cs.LGcs.AIcs.CVarXiv:1802.08219v32018
  9. Spatial As Deep: Spatial CNN for Traffic Scene Understanding

    Xingang Pan, Xiaohang Zhan, Jianping Shi +3

    cs.CVarXiv:1712.06080v22017
  10. DenseCap: Fully Convolutional Localization Networks for Dense Captioning

    Justin Johnson, Andrej Karpathy, Li Fei-Fei

    cs.CVcs.LGarXiv:1511.07571v12015
  11. DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification

    Yongming Rao, Wenliang Zhao, Benlin Liu +3

    cs.CVcs.AIcs.LGarXiv:2106.02034v22021
  12. UCTransNet: Rethinking the Skip Connections in U-Net from a Channel-wise Perspective with Transformer

    Haonan Wang, Peng Cao, Jiaqi Wang +1

    cs.CVcs.LGeess.IVarXiv:2109.04335v32021
  13. WorldSimBench: Towards Video Generation Models as World Simulators

    Yiran Qin, Zhelun Shi, Jiwen Yu +10

    cs.CVarXiv:2410.18072v12024
  14. Spatio-Temporal LSTM with Trust Gates for 3D Human Action Recognition

    Jun Liu, Amir Shahroudy, Dong Xu +1

    cs.CVcs.AIcs.LGarXiv:1607.07043v12016
  15. Virtual Worlds as Proxy for Multi-Object Tracking Analysis

    Adrien Gaidon, Qiao Wang, Yohann Cabon +1

    cs.CVcs.LGcs.NEarXiv:1605.06457v12016
  16. DN-DETR: Accelerate DETR Training by Introducing Query DeNoising

    Feng Li, Hao Zhang, Shilong Liu +3

    cs.CVcs.AIarXiv:2203.01305v32022
  17. PIXOR: Real-time 3D Object Detection from Point Clouds

    Bin Yang, Wenjie Luo, Raquel Urtasun

    cs.CVarXiv:1902.06326v32019
  18. Plug-and-Play Image Restoration with Deep Denoiser Prior

    Kai Zhang, Yawei Li, Wangmeng Zuo +3

    eess.IVcs.CVarXiv:2008.13751v22020
  19. Going Deeper in Spiking Neural Networks: VGG and Residual Architectures

    Abhronil Sengupta, Yuting Ye, Robert Wang +2

    cs.CVarXiv:1802.02627v42018
  20. Multi-Label Image Recognition with Graph Convolutional Networks

    Zhao-Min Chen, Xiu-Shen Wei, Peng Wang +1

    cs.CVcs.LGarXiv:1904.03582v12019
  21. Semantic Image Inpainting with Deep Generative Models

    Raymond A. Yeh, Chen Chen, Teck Yian Lim +3

    cs.CVarXiv:1607.07539v32016
  22. A Lip Sync Expert Is All You Need for Speech to Lip Generation In The Wild

    K R Prajwal, Rudrabha Mukhopadhyay, Vinay Namboodiri +1

    cs.CVcs.LGcs.SDarXiv:2008.10010v12020
  23. DynaSLAM: Tracking, Mapping and Inpainting in Dynamic Scenes

    Berta Bescos, José M. Fácil, Javier Civera +1

    cs.CVarXiv:1806.05620v22018
  24. Unmasking Clever Hans Predictors and Assessing What Machines Really Learn

    Sebastian Lapuschkin, Stephan Wäldchen, Alexander Binder +3

    cs.AIcs.CVcs.LGarXiv:1902.10178v12019
  25. Face X-ray for More General Face Forgery Detection

    Lingzhi Li, Jianmin Bao, Ting Zhang +4

    cs.CVarXiv:1912.13458v22019
  26. First Order Motion Model for Image Animation

    Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov +2

    cs.CVcs.AIarXiv:2003.00196v32020
  27. Localizing Moments in Video with Natural Language

    Lisa Anne Hendricks, Oliver Wang, Eli Shechtman +3

    cs.CVarXiv:1708.01641v12017
  28. CNN-RNN: A Unified Framework for Multi-label Image Classification

    Jiang Wang, Yi Yang, Junhua Mao +3

    cs.CVcs.LGcs.NEarXiv:1604.04573v12016
  29. Representation Engineering: A Top-Down Approach to AI Transparency

    Andy Zou, Long Phan, Sarah Chen +18

    cs.LGcs.AIcs.CLarXiv:2310.01405v42023
  30. ParseNet: Looking Wider to See Better

    Wei Liu, Andrew Rabinovich, Alexander C. Berg

    cs.CVarXiv:1506.04579v22015
  31. Interpreting the Latent Space of GANs for Semantic Face Editing

    Yujun Shen, Jinjin Gu, Xiaoou Tang +1

    cs.CVarXiv:1907.10786v32019
  32. A Deep Cascade of Convolutional Neural Networks for Dynamic MR Image Reconstruction

    Jo Schlemper, Jose Caballero, Joseph V. Hajnal +2

    cs.CVarXiv:1704.02422v22017
  33. A Learned Representation For Artistic Style

    Vincent Dumoulin, Jonathon Shlens, Manjunath Kudlur

    cs.CVcs.LGarXiv:1610.07629v52016
  34. Graph Convolutional Networks for Hyperspectral Image Classification

    Danfeng Hong, Lianru Gao, Jing Yao +3

    cs.CVarXiv:2008.02457v22020
  35. An Analysis of Deep Neural Network Models for Practical Applications

    Alfredo Canziani, Adam Paszke, Eugenio Culurciello

    cs.CVarXiv:1605.07678v42016
  36. Visual Relationship Detection with Language Priors

    Cewu Lu, Ranjay Krishna, Michael Bernstein +1

    cs.CVarXiv:1608.00187v12016
  37. Learning to Track at 100 FPS with Deep Regression Networks

    David Held, Sebastian Thrun, Silvio Savarese

    cs.CVcs.AIcs.LGarXiv:1604.01802v22016
  38. CutPaste: Self-Supervised Learning for Anomaly Detection and Localization

    Chun-Liang Li, Kihyuk Sohn, Jinsung Yoon +1

    cs.CVarXiv:2104.04015v12021
  39. A Survey of Model Compression and Acceleration for Deep Neural Networks

    Yu Cheng, Duo Wang, Pan Zhou +1

    cs.LGcs.CVarXiv:1710.09282v92017
  40. Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

    Tianhe Ren, Shilong Liu, Ailing Zeng +14

    cs.CVarXiv:2401.14159v12024
  41. Filter Pruning via Geometric Median for Deep Convolutional Neural Networks Acceleration

    Yang He, Ping Liu, Ziwei Wang +2

    cs.CVarXiv:1811.00250v32018
  42. Regularization With Stochastic Transformations and Perturbations for Deep Semi-Supervised Learning

    Mehdi Sajjadi, Mehran Javanmardi, Tolga Tasdizen

    cs.CVarXiv:1606.04586v12016
  43. Inside-Outside Net: Detecting Objects in Context with Skip Pooling and Recurrent Neural Networks

    Sean Bell, C. Lawrence Zitnick, Kavita Bala +1

    cs.CVarXiv:1512.04143v12015
  44. Instance-aware Semantic Segmentation via Multi-task Network Cascades

    Jifeng Dai, Kaiming He, Jian Sun

    cs.CVarXiv:1512.04412v12015
  45. Gradient Descent Finds Global Minima of Deep Neural Networks

    Simon S. Du, Jason D. Lee, Haochuan Li +2

    cs.LGcs.AIcs.CVarXiv:1811.03804v42018
  46. Defense-GAN: Protecting Classifiers Against Adversarial Attacks Using Generative Models

    Pouya Samangouei, Maya Kabkab, Rama Chellappa

    cs.CVcs.LGstat.MLarXiv:1805.06605v22018
  47. TransReID: Transformer-based Object Re-Identification

    Shuting He, Hao Luo, Pichao Wang +3

    cs.CVarXiv:2102.04378v22021
  48. Video Generation with Predictive Latents

    Yian Zhao, Feng Wang, Qiushan Guo +4

    cs.CVarXiv:2605.02134v12026
  49. Audio-Visual Intelligence in Large Foundation Models

    You Qin, Kai Liu, Shengqiong Wu +12

    cs.CVarXiv:2605.04045v12026
  50. Encoding in Style: a StyleGAN Encoder for Image-to-Image Translation

    Elad Richardson, Yuval Alaluf, Or Patashnik +4

    cs.CVarXiv:2008.00951v22020
  51. 3D human pose estimation in video with temporal convolutions and semi-supervised training

    Dario Pavllo, Christoph Feichtenhofer, David Grangier +1

    cs.CVarXiv:1811.11742v22018
  52. StableI2I: Spotting Unintended Changes in Image-to-Image Transition

    Jiayang Li, Shuo Cao, Xiaohui Li +6

    cs.CVcs.AIarXiv:2605.04453v12026
  53. Recurrent Residual Convolutional Neural Network based on U-Net (R2U-Net) for Medical Image Segmentation

    Md Zahangir Alom, Mahmudul Hasan, Chris Yakopcic +2

    cs.CVarXiv:1802.06955v52018
  54. A Hybrid Approach for Closing the Sim2real Appearance Gap in Game Engine Synthetic Datasets

    Stefanos Pasios

    cs.CVarXiv:2605.02291v12026
  55. Perceptual Flow Network for Visually Grounded Reasoning

    Yangfu Li, Yuning Gong, Hongjian Zhan +8

    cs.CVcs.AIarXiv:2605.02730v12026
  56. Similarity-Preserving Knowledge Distillation

    Frederick Tung, Greg Mori

    cs.CVarXiv:1907.09682v22019
  57. GeoNet: Unsupervised Learning of Dense Depth, Optical Flow and Camera Pose

    Zhichao Yin, Jianping Shi

    cs.CVarXiv:1803.02276v22018
  58. Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)

    Junhua Mao, Wei Xu, Yi Yang +3

    cs.CVcs.CLcs.LGarXiv:1412.6632v52014
  59. Kosmos-2: Grounding Multimodal Large Language Models to the World

    Zhiliang Peng, Wenhui Wang, Li Dong +4

    cs.CLcs.CVarXiv:2306.14824v32023
  60. Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

    Muhammad Maaz, Hanoona Rasheed, Salman Khan +1

    cs.CVarXiv:2306.05424v22023