Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

16,861 to 16,920 of 18,837

  1. From Coarse to Fine: Robust Hierarchical Localization at Large Scale

    Paul-Edouard Sarlin, Cesar Cadena, Roland Siegwart +1

    cs.CVarXiv:1812.03506v22018
  2. Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation

    Jay Zhangjie Wu, Yixiao Ge, Xintao Wang +7

    cs.CVarXiv:2212.11565v22022
  3. Action Recognition with Trajectory-Pooled Deep-Convolutional Descriptors

    Limin Wang, Yu Qiao, Xiaoou Tang

    cs.CVarXiv:1505.04868v12015
  4. DRAEM -- A discriminatively trained reconstruction embedding for surface anomaly detection

    Vitjan Zavrtanik, Matej Kristan, Danijel Skočaj

    cs.CVarXiv:2108.07610v22021
  5. Image De-raining Using a Conditional Generative Adversarial Network

    He Zhang, Vishwanath Sindagi, Vishal M. Patel

    cs.CVarXiv:1701.05957v42017
  6. Oriented R-CNN for Object Detection

    Xingxing Xie, Gong Cheng, Jiabao Wang +2

    cs.CVarXiv:2108.05699v12021
  7. Neural Module Networks

    Jacob Andreas, Marcus Rohrbach, Trevor Darrell +1

    cs.CVcs.CLcs.LGarXiv:1511.02799v42015
  8. Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?

    Rameen Abdal, Yipeng Qin, Peter Wonka

    cs.CVarXiv:1904.03189v22019
  9. Object-Centric Learning with Slot Attention

    Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner +5

    cs.LGcs.CVstat.MLarXiv:2006.15055v22020
  10. Prototypical Contrastive Learning of Unsupervised Representations

    Junnan Li, Pan Zhou, Caiming Xiong +1

    cs.CVcs.LGarXiv:2005.04966v52020
  11. 3DSSD: Point-based 3D Single Stage Object Detector

    Zetong Yang, Yanan Sun, Shu Liu +1

    cs.CVarXiv:2002.10187v12020
  12. Recovering Realistic Texture in Image Super-resolution by Deep Spatial Feature Transform

    Xintao Wang, Ke Yu, Chao Dong +1

    cs.CVarXiv:1804.02815v12018
  13. PointNeXt: Revisiting PointNet++ with Improved Training and Scaling Strategies

    Guocheng Qian, Yuchen Li, Houwen Peng +4

    cs.CVcs.AIarXiv:2206.04670v22022
  14. FINN: A Framework for Fast, Scalable Binarized Neural Network Inference

    Yaman Umuroglu, Nicholas J. Fraser, Giulio Gambardella +4

    cs.CVcs.ARcs.LGarXiv:1612.07119v12016
  15. TransBTS: Multimodal Brain Tumor Segmentation Using Transformer

    Wenxuan Wang, Chen Chen, Meng Ding +3

    cs.CVcs.AIarXiv:2103.04430v22021
  16. A disciplined approach to neural network hyper-parameters: Part 1 -- learning rate, batch size, momentum, and weight decay

    Leslie N. Smith

    cs.LGcs.CVcs.NEarXiv:1803.09820v22018
  17. Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs

    Xinming Wang, Weinong Wang, Hongming Yang +13

    cs.CVarXiv:2608.12781v22026
  18. Bottleneck Transformers for Visual Recognition

    Aravind Srinivas, Tsung-Yi Lin, Niki Parmar +3

    cs.CVcs.AIcs.LGarXiv:2101.11605v22021
  19. Thinking in Frequency: Face Forgery Detection by Mining Frequency-aware Clues

    Yuyang Qian, Guojun Yin, Lu Sheng +2

    cs.CVarXiv:2007.09355v22020
  20. InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter

    Yunze Tong, Mushui Liu, Canyu Zhao +9

    cs.CVarXiv:2608.20910v12026
  21. Generating Images with Perceptual Similarity Metrics based on Deep Networks

    Alexey Dosovitskiy, Thomas Brox

    cs.LGcs.CVcs.NEarXiv:1602.02644v22016
  22. Exploring Plain Vision Transformer Backbones for Object Detection

    Yanghao Li, Hanzi Mao, Ross Girshick +1

    cs.CVarXiv:2203.16527v22022
  23. Discriminative Scale Space Tracking

    Martin Danelljan, Gustav Häger, Fahad Shahbaz Khan +1

    cs.CVarXiv:1609.06141v12016
  24. Online Self-Calibration Against Hallucination in Vision-Language Models

    Minghui Chen, Chenxu Yang, Hengjie Zhu +3

    cs.CVcs.LGarXiv:2605.00323v12026
  25. Supersizing Self-supervision: Learning to Grasp from 50K Tries and 700 Robot Hours

    Lerrel Pinto, Abhinav Gupta

    cs.LGcs.CVcs.ROarXiv:1509.06825v12015
  26. Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object Detection

    Jinyuan Liu, Xin Fan, Zhanbo Huang +4

    cs.CVarXiv:2203.16220v12022
  27. Detection and Tracking Meet Drones Challenge

    Pengfei Zhu, Longyin Wen, Dawei Du +4

    cs.CVarXiv:2001.06303v32020
  28. Motion-Aware Caching for Efficient Autoregressive Video Generation

    Jing Xu, Yuexiao Ma, Xuzhe Zheng +7

    cs.CVcs.AIarXiv:2605.01725v22026
  29. SplAttN: Bridging 2D and 3D with Gaussian Soft Splatting and Attention for Point Cloud Completion

    Zhaoyang Li, Zhichao You, Tianrui Li

    cs.CVcs.LGarXiv:2605.01466v22026
  30. Let ViT Speak: Generative Language-Image Pre-training

    Yan Fang, Mengcheng Lan, Zilong Huang +7

    cs.CVarXiv:2605.00809v22026
  31. BlenderRAG: High-Fidelity 3D Object Generation via Retrieval-Augmented Code Synthesis

    Massimo Rondelli, Francesco Pivi, Maurizio Gabbrielli

    cs.CVcs.AIcs.GRarXiv:2605.00632v12026
  32. A Comparative Study of Efficient Initialization Methods for the K-Means Clustering Algorithm

    M. Emre Celebi, Hassan A. Kingravi, Patricio A. Vela

    cs.LGcs.CVarXiv:1209.1960v12012
  33. Beyond triplet loss: a deep quadruplet network for person re-identification

    Weihua Chen, Xiaotang Chen, Jianguo Zhang +1

    cs.CVarXiv:1704.01719v12017
  34. HiDDeN: Hiding Data With Deep Networks

    Jiren Zhu, Russell Kaplan, Justin Johnson +1

    cs.CVcs.LGarXiv:1807.09937v12018
  35. AdaptFormer: Adapting Vision Transformers for Scalable Visual Recognition

    Shoufa Chen, Chongjian Ge, Zhan Tong +4

    cs.CVarXiv:2205.13535v32022
  36. Hand Keypoint Detection in Single Images using Multiview Bootstrapping

    Tomas Simon, Hanbyul Joo, Iain Matthews +1

    cs.CVarXiv:1704.07809v12017
  37. Exposing DeepFake Videos By Detecting Face Warping Artifacts

    Yuezun Li, Siwei Lyu

    cs.CVarXiv:1811.00656v32018
  38. TT4D: A Pipeline and Dataset for Table Tennis 4D Reconstruction From Monocular Videos

    Nima Rahmanian, Daniel Kienzle, Thomas Gossard +3

    cs.CVarXiv:2605.01234v12026
  39. AVA: A Video Dataset of Spatio-temporally Localized Atomic Visual Actions

    Chunhui Gu, Chen Sun, David A. Ross +9

    cs.CVarXiv:1705.08421v42017
  40. Dynamic Few-Shot Visual Learning without Forgetting

    Spyros Gidaris, Nikos Komodakis

    cs.CVcs.LGarXiv:1804.09458v12018
  41. Reading Text in the Wild with Convolutional Neural Networks

    Max Jaderberg, Karen Simonyan, Andrea Vedaldi +1

    cs.CVarXiv:1412.1842v12014
  42. EDU-CIRCUIT-HW: Evaluating Multimodal Large Language Models on Real-World University-Level STEM Student Handwritten Solutions

    Weiyu Sun, Liangliang Chen, Yongnuo Cai +3

    cs.CVcs.AIcs.CYarXiv:2602.00095v32026
  43. Transformers in Medical Imaging: A Survey

    Fahad Shamshad, Salman Khan, Syed Waqas Zamir +4

    eess.IVcs.CVarXiv:2201.09873v12022
  44. Compressing Deep Convolutional Networks using Vector Quantization

    Yunchao Gong, Liu Liu, Ming Yang +1

    cs.CVcs.LGcs.NEarXiv:1412.6115v12014
  45. AdaBins: Depth Estimation using Adaptive Bins

    Shariq Farooq Bhat, Ibraheem Alhashim, Peter Wonka

    cs.CVarXiv:2011.14141v12020
  46. Meta-Transfer Learning for Few-Shot Learning

    Qianru Sun, Yaoyao Liu, Tat-Seng Chua +1

    cs.CVarXiv:1812.02391v32018
  47. Quantized Convolutional Neural Networks for Mobile Devices

    Jiaxiang Wu, Cong Leng, Yuhang Wang +2

    cs.CVarXiv:1512.06473v32015
  48. MiniCPM-V: A GPT-4V Level MLLM on Your Phone

    Yuan Yao, Tianyu Yu, Ao Zhang +20

    cs.CVarXiv:2408.01800v12024
  49. NeRF++: Analyzing and Improving Neural Radiance Fields

    Kai Zhang, Gernot Riegler, Noah Snavely +1

    cs.CVarXiv:2010.07492v22020
  50. Tensor field networks: Rotation- and translation-equivariant neural networks for 3D point clouds

    Nathaniel Thomas, Tess Smidt, Steven Kearnes +4

    cs.LGcs.AIcs.CVarXiv:1802.08219v32018
  51. Spatial As Deep: Spatial CNN for Traffic Scene Understanding

    Xingang Pan, Xiaohang Zhan, Jianping Shi +3

    cs.CVarXiv:1712.06080v22017
  52. DenseCap: Fully Convolutional Localization Networks for Dense Captioning

    Justin Johnson, Andrej Karpathy, Li Fei-Fei

    cs.CVcs.LGarXiv:1511.07571v12015
  53. DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification

    Yongming Rao, Wenliang Zhao, Benlin Liu +3

    cs.CVcs.AIcs.LGarXiv:2106.02034v22021
  54. UCTransNet: Rethinking the Skip Connections in U-Net from a Channel-wise Perspective with Transformer

    Haonan Wang, Peng Cao, Jiaqi Wang +1

    cs.CVcs.LGeess.IVarXiv:2109.04335v32021
  55. WorldSimBench: Towards Video Generation Models as World Simulators

    Yiran Qin, Zhelun Shi, Jiwen Yu +10

    cs.CVarXiv:2410.18072v12024
  56. Spatio-Temporal LSTM with Trust Gates for 3D Human Action Recognition

    Jun Liu, Amir Shahroudy, Dong Xu +1

    cs.CVcs.AIcs.LGarXiv:1607.07043v12016
  57. Virtual Worlds as Proxy for Multi-Object Tracking Analysis

    Adrien Gaidon, Qiao Wang, Yohann Cabon +1

    cs.CVcs.LGcs.NEarXiv:1605.06457v12016
  58. DN-DETR: Accelerate DETR Training by Introducing Query DeNoising

    Feng Li, Hao Zhang, Shilong Liu +3

    cs.CVcs.AIarXiv:2203.01305v32022
  59. PIXOR: Real-time 3D Object Detection from Point Clouds

    Bin Yang, Wenjie Luo, Raquel Urtasun

    cs.CVarXiv:1902.06326v32019
  60. Plug-and-Play Image Restoration with Deep Denoiser Prior

    Kai Zhang, Yawei Li, Wangmeng Zuo +3

    eess.IVcs.CVarXiv:2008.13751v22020