Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
16,921 to 16,980 of 18,855
Transformers in Medical Imaging: A Survey
Fahad Shamshad, Salman Khan, Syed Waqas Zamir +4
eess.IVcs.CVarXiv:2201.09873v12022Compressing Deep Convolutional Networks using Vector Quantization
Yunchao Gong, Liu Liu, Ming Yang +1
cs.CVcs.LGcs.NEarXiv:1412.6115v12014AdaBins: Depth Estimation using Adaptive Bins
Shariq Farooq Bhat, Ibraheem Alhashim, Peter Wonka
cs.CVarXiv:2011.14141v12020Meta-Transfer Learning for Few-Shot Learning
Qianru Sun, Yaoyao Liu, Tat-Seng Chua +1
cs.CVarXiv:1812.02391v32018Quantized Convolutional Neural Networks for Mobile Devices
Jiaxiang Wu, Cong Leng, Yuhang Wang +2
cs.CVarXiv:1512.06473v32015MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Yuan Yao, Tianyu Yu, Ao Zhang +20
cs.CVarXiv:2408.01800v12024NeRF++: Analyzing and Improving Neural Radiance Fields
Kai Zhang, Gernot Riegler, Noah Snavely +1
cs.CVarXiv:2010.07492v22020Tensor field networks: Rotation- and translation-equivariant neural networks for 3D point clouds
Nathaniel Thomas, Tess Smidt, Steven Kearnes +4
cs.LGcs.AIcs.CVarXiv:1802.08219v32018Spatial As Deep: Spatial CNN for Traffic Scene Understanding
Xingang Pan, Xiaohang Zhan, Jianping Shi +3
cs.CVarXiv:1712.06080v22017DenseCap: Fully Convolutional Localization Networks for Dense Captioning
Justin Johnson, Andrej Karpathy, Li Fei-Fei
cs.CVcs.LGarXiv:1511.07571v12015DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification
Yongming Rao, Wenliang Zhao, Benlin Liu +3
cs.CVcs.AIcs.LGarXiv:2106.02034v22021UCTransNet: Rethinking the Skip Connections in U-Net from a Channel-wise Perspective with Transformer
Haonan Wang, Peng Cao, Jiaqi Wang +1
cs.CVcs.LGeess.IVarXiv:2109.04335v32021WorldSimBench: Towards Video Generation Models as World Simulators
Yiran Qin, Zhelun Shi, Jiwen Yu +10
cs.CVarXiv:2410.18072v12024Spatio-Temporal LSTM with Trust Gates for 3D Human Action Recognition
Jun Liu, Amir Shahroudy, Dong Xu +1
cs.CVcs.AIcs.LGarXiv:1607.07043v12016Virtual Worlds as Proxy for Multi-Object Tracking Analysis
Adrien Gaidon, Qiao Wang, Yohann Cabon +1
cs.CVcs.LGcs.NEarXiv:1605.06457v12016DN-DETR: Accelerate DETR Training by Introducing Query DeNoising
Feng Li, Hao Zhang, Shilong Liu +3
cs.CVcs.AIarXiv:2203.01305v32022PIXOR: Real-time 3D Object Detection from Point Clouds
Bin Yang, Wenjie Luo, Raquel Urtasun
cs.CVarXiv:1902.06326v32019Plug-and-Play Image Restoration with Deep Denoiser Prior
Kai Zhang, Yawei Li, Wangmeng Zuo +3
eess.IVcs.CVarXiv:2008.13751v22020Going Deeper in Spiking Neural Networks: VGG and Residual Architectures
Abhronil Sengupta, Yuting Ye, Robert Wang +2
cs.CVarXiv:1802.02627v42018Multi-Label Image Recognition with Graph Convolutional Networks
Zhao-Min Chen, Xiu-Shen Wei, Peng Wang +1
cs.CVcs.LGarXiv:1904.03582v12019Semantic Image Inpainting with Deep Generative Models
Raymond A. Yeh, Chen Chen, Teck Yian Lim +3
cs.CVarXiv:1607.07539v32016A Lip Sync Expert Is All You Need for Speech to Lip Generation In The Wild
K R Prajwal, Rudrabha Mukhopadhyay, Vinay Namboodiri +1
cs.CVcs.LGcs.SDarXiv:2008.10010v12020DynaSLAM: Tracking, Mapping and Inpainting in Dynamic Scenes
Berta Bescos, José M. Fácil, Javier Civera +1
cs.CVarXiv:1806.05620v22018Unmasking Clever Hans Predictors and Assessing What Machines Really Learn
Sebastian Lapuschkin, Stephan Wäldchen, Alexander Binder +3
cs.AIcs.CVcs.LGarXiv:1902.10178v12019Face X-ray for More General Face Forgery Detection
Lingzhi Li, Jianmin Bao, Ting Zhang +4
cs.CVarXiv:1912.13458v22019First Order Motion Model for Image Animation
Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov +2
cs.CVcs.AIarXiv:2003.00196v32020Localizing Moments in Video with Natural Language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman +3
cs.CVarXiv:1708.01641v12017CNN-RNN: A Unified Framework for Multi-label Image Classification
Jiang Wang, Yi Yang, Junhua Mao +3
cs.CVcs.LGcs.NEarXiv:1604.04573v12016Representation Engineering: A Top-Down Approach to AI Transparency
Andy Zou, Long Phan, Sarah Chen +18
cs.LGcs.AIcs.CLarXiv:2310.01405v42023ParseNet: Looking Wider to See Better
Wei Liu, Andrew Rabinovich, Alexander C. Berg
cs.CVarXiv:1506.04579v22015Interpreting the Latent Space of GANs for Semantic Face Editing
Yujun Shen, Jinjin Gu, Xiaoou Tang +1
cs.CVarXiv:1907.10786v32019A Deep Cascade of Convolutional Neural Networks for Dynamic MR Image Reconstruction
Jo Schlemper, Jose Caballero, Joseph V. Hajnal +2
cs.CVarXiv:1704.02422v22017A Learned Representation For Artistic Style
Vincent Dumoulin, Jonathon Shlens, Manjunath Kudlur
cs.CVcs.LGarXiv:1610.07629v52016Graph Convolutional Networks for Hyperspectral Image Classification
Danfeng Hong, Lianru Gao, Jing Yao +3
cs.CVarXiv:2008.02457v22020An Analysis of Deep Neural Network Models for Practical Applications
Alfredo Canziani, Adam Paszke, Eugenio Culurciello
cs.CVarXiv:1605.07678v42016Visual Relationship Detection with Language Priors
Cewu Lu, Ranjay Krishna, Michael Bernstein +1
cs.CVarXiv:1608.00187v12016Learning to Track at 100 FPS with Deep Regression Networks
David Held, Sebastian Thrun, Silvio Savarese
cs.CVcs.AIcs.LGarXiv:1604.01802v22016CutPaste: Self-Supervised Learning for Anomaly Detection and Localization
Chun-Liang Li, Kihyuk Sohn, Jinsung Yoon +1
cs.CVarXiv:2104.04015v12021A Survey of Model Compression and Acceleration for Deep Neural Networks
Yu Cheng, Duo Wang, Pan Zhou +1
cs.LGcs.CVarXiv:1710.09282v92017Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
Tianhe Ren, Shilong Liu, Ailing Zeng +14
cs.CVarXiv:2401.14159v12024Filter Pruning via Geometric Median for Deep Convolutional Neural Networks Acceleration
Yang He, Ping Liu, Ziwei Wang +2
cs.CVarXiv:1811.00250v32018Regularization With Stochastic Transformations and Perturbations for Deep Semi-Supervised Learning
Mehdi Sajjadi, Mehran Javanmardi, Tolga Tasdizen
cs.CVarXiv:1606.04586v12016Inside-Outside Net: Detecting Objects in Context with Skip Pooling and Recurrent Neural Networks
Sean Bell, C. Lawrence Zitnick, Kavita Bala +1
cs.CVarXiv:1512.04143v12015Instance-aware Semantic Segmentation via Multi-task Network Cascades
Jifeng Dai, Kaiming He, Jian Sun
cs.CVarXiv:1512.04412v12015Gradient Descent Finds Global Minima of Deep Neural Networks
Simon S. Du, Jason D. Lee, Haochuan Li +2
cs.LGcs.AIcs.CVarXiv:1811.03804v42018Defense-GAN: Protecting Classifiers Against Adversarial Attacks Using Generative Models
Pouya Samangouei, Maya Kabkab, Rama Chellappa
cs.CVcs.LGstat.MLarXiv:1805.06605v22018TransReID: Transformer-based Object Re-Identification
Shuting He, Hao Luo, Pichao Wang +3
cs.CVarXiv:2102.04378v22021Video Generation with Predictive Latents
Yian Zhao, Feng Wang, Qiushan Guo +4
cs.CVarXiv:2605.02134v12026Audio-Visual Intelligence in Large Foundation Models
You Qin, Kai Liu, Shengqiong Wu +12
cs.CVarXiv:2605.04045v12026Encoding in Style: a StyleGAN Encoder for Image-to-Image Translation
Elad Richardson, Yuval Alaluf, Or Patashnik +4
cs.CVarXiv:2008.00951v220203D human pose estimation in video with temporal convolutions and semi-supervised training
Dario Pavllo, Christoph Feichtenhofer, David Grangier +1
cs.CVarXiv:1811.11742v22018StableI2I: Spotting Unintended Changes in Image-to-Image Transition
Jiayang Li, Shuo Cao, Xiaohui Li +6
cs.CVcs.AIarXiv:2605.04453v12026Recurrent Residual Convolutional Neural Network based on U-Net (R2U-Net) for Medical Image Segmentation
Md Zahangir Alom, Mahmudul Hasan, Chris Yakopcic +2
cs.CVarXiv:1802.06955v52018A Hybrid Approach for Closing the Sim2real Appearance Gap in Game Engine Synthetic Datasets
Stefanos Pasios
cs.CVarXiv:2605.02291v12026Perceptual Flow Network for Visually Grounded Reasoning
Yangfu Li, Yuning Gong, Hongjian Zhan +8
cs.CVcs.AIarXiv:2605.02730v12026Similarity-Preserving Knowledge Distillation
Frederick Tung, Greg Mori
cs.CVarXiv:1907.09682v22019GeoNet: Unsupervised Learning of Dense Depth, Optical Flow and Camera Pose
Zhichao Yin, Jianping Shi
cs.CVarXiv:1803.02276v22018Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
Junhua Mao, Wei Xu, Yi Yang +3
cs.CVcs.CLcs.LGarXiv:1412.6632v52014Kosmos-2: Grounding Multimodal Large Language Models to the World
Zhiliang Peng, Wenhui Wang, Li Dong +4
cs.CLcs.CVarXiv:2306.14824v32023Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Muhammad Maaz, Hanoona Rasheed, Salman Khan +1
cs.CVarXiv:2306.05424v22023