Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
17,581 to 17,640 of 18,831
Do ImageNet Classifiers Generalize to ImageNet?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt +1
cs.CVcs.LGstat.MLarXiv:1902.10811v22019PV-RCNN: Point-Voxel Feature Set Abstraction for 3D Object Detection
Shaoshuai Shi, Chaoxu Guo, Li Jiang +4
cs.CVcs.LGeess.IVarXiv:1912.13192v22019Deep Learning for 3D Point Clouds: A Survey
Yulan Guo, Hanyun Wang, Qingyong Hu +3
cs.CVcs.LGcs.ROarXiv:1912.12033v22019SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations
Chenlin Meng, Yutong He, Yang Song +4
cs.CVcs.AIarXiv:2108.01073v22021SPICE: Semantic Propositional Image Caption Evaluation
Peter Anderson, Basura Fernando, Mark Johnson +1
cs.CVcs.CLarXiv:1607.08822v12016PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes
Yu Xiang, Tanner Schmidt, Venkatraman Narayanan +1
cs.CVcs.ROarXiv:1711.00199v32017Moment Matching for Multi-Source Domain Adaptation
Xingchao Peng, Qinxun Bai, Xide Xia +3
cs.CVarXiv:1812.01754v42018Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Jinze Bai, Shuai Bai, Shusheng Yang +6
cs.CVcs.CLarXiv:2308.12966v32023ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation
Adam Paszke, Abhishek Chaurasia, Sangpil Kim +1
cs.CVarXiv:1606.02147v12016NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction
Peng Wang, Lingjie Liu, Yuan Liu +3
cs.CVcs.GRarXiv:2106.10689v32021Beyond Part Models: Person Retrieval with Refined Part Pooling (and a Strong Convolutional Baseline)
Yifan Sun, Liang Zheng, Yi Yang +2
cs.CVarXiv:1711.09349v32017An Empirical Study of Training Self-Supervised Vision Transformers
Xinlei Chen, Saining Xie, Kaiming He
cs.CVcs.LGarXiv:2104.02057v420214D Spatio-Temporal ConvNets: Minkowski Convolutional Neural Networks
Christopher Choy, JunYoung Gwak, Silvio Savarese
cs.CVcs.AIarXiv:1904.08755v42019Conditional Prompt Learning for Vision-Language Models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy +1
cs.CVcs.AIcs.CLarXiv:2203.05557v22022MMBench: Is Your Multi-modal Model an All-around Player?
Yuan Liu, Haodong Duan, Yuanhan Zhang +9
cs.CVcs.CLarXiv:2307.06281v52023Visualizing the Loss Landscape of Neural Nets
Hao Li, Zheng Xu, Gavin Taylor +2
cs.LGcs.CVstat.MLarXiv:1712.09913v32017Representation Learning on Graphs with Jumping Knowledge Networks
Keyulu Xu, Chengtao Li, Yonglong Tian +3
cs.LGcs.AIcs.CVarXiv:1806.03536v22018Generating Diverse High-Fidelity Images with VQ-VAE-2
Ali Razavi, Aaron van den Oord, Oriol Vinyals
cs.LGcs.CVstat.MLarXiv:1906.00446v12019Beyond Short Snippets: Deep Networks for Video Classification
Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan +3
cs.CVarXiv:1503.08909v22015Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks
Agrim Gupta, Justin Johnson, Li Fei-Fei +2
cs.CVarXiv:1803.10892v12018MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Xiang Yue, Yuansheng Ni, Kai Zhang +19
cs.CLcs.AIcs.CVarXiv:2311.16502v42023MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer
Sachin Mehta, Mohammad Rastegari
cs.CVcs.AIcs.LGarXiv:2110.02178v22021The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization
Dan Hendrycks, Steven Basart, Norman Mu +10
cs.CVcs.LGstat.MLarXiv:2006.16241v32020Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors
Hanxun Yu, Xuan Qu, Lei Ke +4
cs.CVarXiv:2606.06891v12026Struct-Searcher: Agentic Structural Thinking Advances Multimodal Deep Information Seeking
Fan Zhang, Vireo Zhang, Shengju Qian +7
cs.CVarXiv:2606.07689v12026AirSim: High-Fidelity Visual and Physical Simulation for Autonomous Vehicles
Shital Shah, Debadeepta Dey, Chris Lovett +1
cs.ROcs.AIcs.CVarXiv:1705.05065v22017Deeply-Supervised Nets
Chen-Yu Lee, Saining Xie, Patrick Gallagher +2
stat.MLcs.CVcs.LGarXiv:1409.5185v22014Expressive Body Capture: 3D Hands, Face, and Body from a Single Image
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani +4
cs.CVarXiv:1904.05866v12019Deep Multi-scale Convolutional Neural Network for Dynamic Scene Deblurring
Seungjun Nah, Tae Hyun Kim, Kyoung Mu Lee
cs.CVarXiv:1612.02177v22016Unsupervised Anomaly Detection with Generative Adversarial Networks to Guide Marker Discovery
Thomas Schlegl, Philipp Seeböck, Sebastian M. Waldstein +2
cs.CVcs.LGarXiv:1703.05921v12017Continual Learning with Deep Generative Replay
Hanul Shin, Jung Kwon Lee, Jaehong Kim +1
cs.AIcs.CVcs.LGarXiv:1705.08690v32017RepVGG: Making VGG-style ConvNets Great Again
Xiaohan Ding, Xiangyu Zhang, Ningning Ma +3
cs.CVcs.AIcs.LGarXiv:2101.03697v32021Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly +1
cs.LGcs.CLcs.CVarXiv:1506.03099v32015EnlightenGAN: Deep Light Enhancement without Paired Supervision
Yifan Jiang, Xinyu Gong, Ding Liu +6
cs.CVeess.IVarXiv:1906.06972v22019ECO: Efficient Convolution Operators for Tracking
Martin Danelljan, Goutam Bhat, Fahad Shahbaz Khan +1
cs.CVarXiv:1611.09224v22016AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization
Yu Li, Menghan Xia, Gongye Liu +8
cs.CVarXiv:2606.07326v12026Robotic Policy Adaptation via Weight-Space Meta-Learning
Christian Bianchi, Siamak Yousefi, Alessio Sampieri +4
cs.ROcs.CVcs.LGarXiv:2606.07217v12026Swin UNETR: Swin Transformers for Semantic Segmentation of Brain Tumors in MRI Images
Ali Hatamizadeh, Vishwesh Nath, Yucheng Tang +3
eess.IVcs.CVcs.LGarXiv:2201.01266v12022Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks
Francesco Croce, Matthias Hein
cs.LGcs.CVstat.MLarXiv:2003.01690v22020CvT: Introducing Convolutions to Vision Transformers
Haiping Wu, Bin Xiao, Noel Codella +4
cs.CVarXiv:2103.15808v12021Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet
Li Yuan, Yunpeng Chen, Tao Wang +6
cs.CVarXiv:2101.11986v32021Skill-3D: Evolving Scene-Aware Skills for Agentic 3D Spatial Reasoning
Haoyuan Li, Zhengdong Hu, Jun Wang +2
cs.CVarXiv:2606.07436v22026Deep Hashing Network for Unsupervised Domain Adaptation
Hemanth Venkateswara, Jose Eusebio, Shayok Chakraborty +1
cs.CVarXiv:1706.07522v12017Set-Based Transformer for Atmospheric Compensation in Standoff LWIR Hyperspectral Imaging
Fabian Perez, Nicolas Quintero, Jeferson Acevedo +1
cs.CVcs.AIarXiv:2606.08324v12026Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Pan Lu, Swaroop Mishra, Tony Xia +6
cs.CLcs.AIcs.CVarXiv:2209.09513v22022Generative Image Inpainting with Contextual Attention
Jiahui Yu, Zhe Lin, Jimei Yang +3
cs.CVcs.GRarXiv:1801.07892v22018Big Self-Supervised Models are Strong Semi-Supervised Learners
Ting Chen, Simon Kornblith, Kevin Swersky +2
cs.LGcs.CVstat.MLarXiv:2006.10029v22020SwiftVR: Real-Time One-Step Generative Video Restoration
Jiaqi Yan, Xiangyu Chen, Xinlin Zhong +5
cs.CVarXiv:2606.09516v12026EIE: Efficient Inference Engine on Compressed Deep Neural Network
Song Han, Xingyu Liu, Huizi Mao +4
cs.CVcs.ARarXiv:1602.01528v22016Frustum PointNets for 3D Object Detection from RGB-D Data
Charles R. Qi, Wei Liu, Chenxia Wu +2
cs.CVarXiv:1711.08488v22017Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere
Tongzhou Wang, Phillip Isola
cs.LGcs.CVstat.MLarXiv:2005.10242v102020Revisiting Articulated Parts Perception in Robot Manipulation
Xiaoqian Wu, Yejie Guo, Xiaoyang Chen +3
cs.ROcs.CVarXiv:2606.08103v12026Phase Marginalization for Patch-Grid Instability in Vision Transformers
Oğuzhan Ercan
cs.CVcs.LGarXiv:2606.08132v12026PVT v2: Improved Baselines with Pyramid Vision Transformer
Wenhai Wang, Enze Xie, Xiang Li +6
cs.CVarXiv:2106.13797v72021Sanity Checks for Saliency Maps
Julius Adebayo, Justin Gilmer, Michael Muelly +3
cs.CVcs.LGstat.MLarXiv:1810.03292v32018OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning
Jiahao Wang, An Ping, Yanghai Wang +13
cs.CVarXiv:2606.08572v12026MaskAlign: Token-Subset Representation Alignment for Efficient Diffusion Training
Lianyu Pang, Tianlin Pan, Cheng Da +5
cs.CVarXiv:2606.08788v12026OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics
Mingxian Lin, Shengju Qian, Yuqi Liu +9
cs.CVcs.AIarXiv:2606.09826v12026One pixel attack for fooling deep neural networks
Jiawei Su, Danilo Vasconcellos Vargas, Sakurai Kouichi
cs.LGcs.CVstat.MLarXiv:1710.08864v72017ByteTrack: Multi-Object Tracking by Associating Every Detection Box
Yifu Zhang, Peize Sun, Yi Jiang +6
cs.CVarXiv:2110.06864v32021