Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
16,861 to 16,920 of 18,837
From Coarse to Fine: Robust Hierarchical Localization at Large Scale
Paul-Edouard Sarlin, Cesar Cadena, Roland Siegwart +1
cs.CVarXiv:1812.03506v22018Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation
Jay Zhangjie Wu, Yixiao Ge, Xintao Wang +7
cs.CVarXiv:2212.11565v22022Action Recognition with Trajectory-Pooled Deep-Convolutional Descriptors
Limin Wang, Yu Qiao, Xiaoou Tang
cs.CVarXiv:1505.04868v12015DRAEM -- A discriminatively trained reconstruction embedding for surface anomaly detection
Vitjan Zavrtanik, Matej Kristan, Danijel Skočaj
cs.CVarXiv:2108.07610v22021Image De-raining Using a Conditional Generative Adversarial Network
He Zhang, Vishwanath Sindagi, Vishal M. Patel
cs.CVarXiv:1701.05957v42017Oriented R-CNN for Object Detection
Xingxing Xie, Gong Cheng, Jiabao Wang +2
cs.CVarXiv:2108.05699v12021Neural Module Networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell +1
cs.CVcs.CLcs.LGarXiv:1511.02799v42015Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?
Rameen Abdal, Yipeng Qin, Peter Wonka
cs.CVarXiv:1904.03189v22019Object-Centric Learning with Slot Attention
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner +5
cs.LGcs.CVstat.MLarXiv:2006.15055v22020Prototypical Contrastive Learning of Unsupervised Representations
Junnan Li, Pan Zhou, Caiming Xiong +1
cs.CVcs.LGarXiv:2005.04966v520203DSSD: Point-based 3D Single Stage Object Detector
Zetong Yang, Yanan Sun, Shu Liu +1
cs.CVarXiv:2002.10187v12020Recovering Realistic Texture in Image Super-resolution by Deep Spatial Feature Transform
Xintao Wang, Ke Yu, Chao Dong +1
cs.CVarXiv:1804.02815v12018PointNeXt: Revisiting PointNet++ with Improved Training and Scaling Strategies
Guocheng Qian, Yuchen Li, Houwen Peng +4
cs.CVcs.AIarXiv:2206.04670v22022FINN: A Framework for Fast, Scalable Binarized Neural Network Inference
Yaman Umuroglu, Nicholas J. Fraser, Giulio Gambardella +4
cs.CVcs.ARcs.LGarXiv:1612.07119v12016TransBTS: Multimodal Brain Tumor Segmentation Using Transformer
Wenxuan Wang, Chen Chen, Meng Ding +3
cs.CVcs.AIarXiv:2103.04430v22021A disciplined approach to neural network hyper-parameters: Part 1 -- learning rate, batch size, momentum, and weight decay
Leslie N. Smith
cs.LGcs.CVcs.NEarXiv:1803.09820v22018Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs
Xinming Wang, Weinong Wang, Hongming Yang +13
cs.CVarXiv:2608.12781v22026Bottleneck Transformers for Visual Recognition
Aravind Srinivas, Tsung-Yi Lin, Niki Parmar +3
cs.CVcs.AIcs.LGarXiv:2101.11605v22021Thinking in Frequency: Face Forgery Detection by Mining Frequency-aware Clues
Yuyang Qian, Guojun Yin, Lu Sheng +2
cs.CVarXiv:2007.09355v22020InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter
Yunze Tong, Mushui Liu, Canyu Zhao +9
cs.CVarXiv:2608.20910v12026Generating Images with Perceptual Similarity Metrics based on Deep Networks
Alexey Dosovitskiy, Thomas Brox
cs.LGcs.CVcs.NEarXiv:1602.02644v22016Exploring Plain Vision Transformer Backbones for Object Detection
Yanghao Li, Hanzi Mao, Ross Girshick +1
cs.CVarXiv:2203.16527v22022Discriminative Scale Space Tracking
Martin Danelljan, Gustav Häger, Fahad Shahbaz Khan +1
cs.CVarXiv:1609.06141v12016Online Self-Calibration Against Hallucination in Vision-Language Models
Minghui Chen, Chenxu Yang, Hengjie Zhu +3
cs.CVcs.LGarXiv:2605.00323v12026Supersizing Self-supervision: Learning to Grasp from 50K Tries and 700 Robot Hours
Lerrel Pinto, Abhinav Gupta
cs.LGcs.CVcs.ROarXiv:1509.06825v12015Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object Detection
Jinyuan Liu, Xin Fan, Zhanbo Huang +4
cs.CVarXiv:2203.16220v12022Detection and Tracking Meet Drones Challenge
Pengfei Zhu, Longyin Wen, Dawei Du +4
cs.CVarXiv:2001.06303v32020Motion-Aware Caching for Efficient Autoregressive Video Generation
Jing Xu, Yuexiao Ma, Xuzhe Zheng +7
cs.CVcs.AIarXiv:2605.01725v22026SplAttN: Bridging 2D and 3D with Gaussian Soft Splatting and Attention for Point Cloud Completion
Zhaoyang Li, Zhichao You, Tianrui Li
cs.CVcs.LGarXiv:2605.01466v22026Let ViT Speak: Generative Language-Image Pre-training
Yan Fang, Mengcheng Lan, Zilong Huang +7
cs.CVarXiv:2605.00809v22026BlenderRAG: High-Fidelity 3D Object Generation via Retrieval-Augmented Code Synthesis
Massimo Rondelli, Francesco Pivi, Maurizio Gabbrielli
cs.CVcs.AIcs.GRarXiv:2605.00632v12026A Comparative Study of Efficient Initialization Methods for the K-Means Clustering Algorithm
M. Emre Celebi, Hassan A. Kingravi, Patricio A. Vela
cs.LGcs.CVarXiv:1209.1960v12012Beyond triplet loss: a deep quadruplet network for person re-identification
Weihua Chen, Xiaotang Chen, Jianguo Zhang +1
cs.CVarXiv:1704.01719v12017HiDDeN: Hiding Data With Deep Networks
Jiren Zhu, Russell Kaplan, Justin Johnson +1
cs.CVcs.LGarXiv:1807.09937v12018AdaptFormer: Adapting Vision Transformers for Scalable Visual Recognition
Shoufa Chen, Chongjian Ge, Zhan Tong +4
cs.CVarXiv:2205.13535v32022Hand Keypoint Detection in Single Images using Multiview Bootstrapping
Tomas Simon, Hanbyul Joo, Iain Matthews +1
cs.CVarXiv:1704.07809v12017Exposing DeepFake Videos By Detecting Face Warping Artifacts
Yuezun Li, Siwei Lyu
cs.CVarXiv:1811.00656v32018TT4D: A Pipeline and Dataset for Table Tennis 4D Reconstruction From Monocular Videos
Nima Rahmanian, Daniel Kienzle, Thomas Gossard +3
cs.CVarXiv:2605.01234v12026AVA: A Video Dataset of Spatio-temporally Localized Atomic Visual Actions
Chunhui Gu, Chen Sun, David A. Ross +9
cs.CVarXiv:1705.08421v42017Dynamic Few-Shot Visual Learning without Forgetting
Spyros Gidaris, Nikos Komodakis
cs.CVcs.LGarXiv:1804.09458v12018Reading Text in the Wild with Convolutional Neural Networks
Max Jaderberg, Karen Simonyan, Andrea Vedaldi +1
cs.CVarXiv:1412.1842v12014EDU-CIRCUIT-HW: Evaluating Multimodal Large Language Models on Real-World University-Level STEM Student Handwritten Solutions
Weiyu Sun, Liangliang Chen, Yongnuo Cai +3
cs.CVcs.AIcs.CYarXiv:2602.00095v32026Transformers in Medical Imaging: A Survey
Fahad Shamshad, Salman Khan, Syed Waqas Zamir +4
eess.IVcs.CVarXiv:2201.09873v12022Compressing Deep Convolutional Networks using Vector Quantization
Yunchao Gong, Liu Liu, Ming Yang +1
cs.CVcs.LGcs.NEarXiv:1412.6115v12014AdaBins: Depth Estimation using Adaptive Bins
Shariq Farooq Bhat, Ibraheem Alhashim, Peter Wonka
cs.CVarXiv:2011.14141v12020Meta-Transfer Learning for Few-Shot Learning
Qianru Sun, Yaoyao Liu, Tat-Seng Chua +1
cs.CVarXiv:1812.02391v32018Quantized Convolutional Neural Networks for Mobile Devices
Jiaxiang Wu, Cong Leng, Yuhang Wang +2
cs.CVarXiv:1512.06473v32015MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Yuan Yao, Tianyu Yu, Ao Zhang +20
cs.CVarXiv:2408.01800v12024NeRF++: Analyzing and Improving Neural Radiance Fields
Kai Zhang, Gernot Riegler, Noah Snavely +1
cs.CVarXiv:2010.07492v22020Tensor field networks: Rotation- and translation-equivariant neural networks for 3D point clouds
Nathaniel Thomas, Tess Smidt, Steven Kearnes +4
cs.LGcs.AIcs.CVarXiv:1802.08219v32018Spatial As Deep: Spatial CNN for Traffic Scene Understanding
Xingang Pan, Xiaohang Zhan, Jianping Shi +3
cs.CVarXiv:1712.06080v22017DenseCap: Fully Convolutional Localization Networks for Dense Captioning
Justin Johnson, Andrej Karpathy, Li Fei-Fei
cs.CVcs.LGarXiv:1511.07571v12015DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification
Yongming Rao, Wenliang Zhao, Benlin Liu +3
cs.CVcs.AIcs.LGarXiv:2106.02034v22021UCTransNet: Rethinking the Skip Connections in U-Net from a Channel-wise Perspective with Transformer
Haonan Wang, Peng Cao, Jiaqi Wang +1
cs.CVcs.LGeess.IVarXiv:2109.04335v32021WorldSimBench: Towards Video Generation Models as World Simulators
Yiran Qin, Zhelun Shi, Jiwen Yu +10
cs.CVarXiv:2410.18072v12024Spatio-Temporal LSTM with Trust Gates for 3D Human Action Recognition
Jun Liu, Amir Shahroudy, Dong Xu +1
cs.CVcs.AIcs.LGarXiv:1607.07043v12016Virtual Worlds as Proxy for Multi-Object Tracking Analysis
Adrien Gaidon, Qiao Wang, Yohann Cabon +1
cs.CVcs.LGcs.NEarXiv:1605.06457v12016DN-DETR: Accelerate DETR Training by Introducing Query DeNoising
Feng Li, Hao Zhang, Shilong Liu +3
cs.CVcs.AIarXiv:2203.01305v32022PIXOR: Real-time 3D Object Detection from Point Clouds
Bin Yang, Wenjie Luo, Raquel Urtasun
cs.CVarXiv:1902.06326v32019Plug-and-Play Image Restoration with Deep Denoiser Prior
Kai Zhang, Yawei Li, Wangmeng Zuo +3
eess.IVcs.CVarXiv:2008.13751v22020