Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
16,261 to 16,320 of 18,975
SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation Guidelines
Yinda Xu, Zeyu Wang, Zuoxin Li +2
cs.CVarXiv:1911.06188v42019The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
Douwe Kiela, Hamed Firooz, Aravind Mohan +4
cs.AIcs.CLcs.CVarXiv:2005.04790v32020RMA: Rapid Motor Adaptation for Legged Robots
Ashish Kumar, Zipeng Fu, Deepak Pathak +1
cs.LGcs.AIcs.CVarXiv:2107.04034v12021Large Batch Training of Convolutional Networks
Yang You, Igor Gitman, Boris Ginsburg
cs.CVarXiv:1708.03888v32017Weakly- and Semi-Supervised Learning of a DCNN for Semantic Image Segmentation
George Papandreou, Liang-Chieh Chen, Kevin Murphy +1
cs.CVarXiv:1502.02734v32015A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge
Dustin Schwenk, Apoorv Khandelwal, Christopher Clark +2
cs.CVcs.CLarXiv:2206.01718v12022The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems
Xiaoze Liu, Ruowang Zhang, Weichen Yu +7
cs.CLcs.CVcs.LGarXiv:2602.15382v22026MWM: Mobile World Models for Action-Conditioned Consistent Prediction
Han Yan, Zishang Xiang, Zeyu Zhang +1
cs.CVcs.ROarXiv:2603.07799v12026HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising
Kai Zou, Dian Zheng, Hongbo Liu +3
cs.CVarXiv:2603.08703v12026LIVE: Long-horizon Interactive Video World Modeling
Junchao Huang, Ziyang Ye, Xinting Hu +5
cs.CVarXiv:2602.03747v12026Temporal Segment Networks for Action Recognition in Videos
Limin Wang, Yuanjun Xiong, Zhe Wang +4
cs.CVarXiv:1705.02953v12017NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions
Junbin Xiao, Xindi Shang, Angela Yao +1
cs.CVcs.AIarXiv:2105.08276v22021pi-GAN: Periodic Implicit Generative Adversarial Networks for 3D-Aware Image Synthesis
Eric R. Chan, Marco Monteiro, Petr Kellnhofer +2
cs.CVcs.GRarXiv:2012.00926v22020Medical Image Segmentation Review: The success of U-Net
Reza Azad, Ehsan Khodapanah Aghdam, Amelie Rauland +7
eess.IVcs.CVarXiv:2211.14830v12022LOME: Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model
Quankai Gao, Jiawei Yang, Qiangeng Xu +2
cs.CVarXiv:2603.27449v12026Mobile-GS: Real-time Gaussian Splatting for Mobile Devices
Xiaobiao Du, Yida Wang, Kun Zhan +1
cs.CVarXiv:2603.11531v12026ILVR: Conditioning Method for Denoising Diffusion Probabilistic Models
Jooyoung Choi, Sungwon Kim, Yonghyun Jeong +2
cs.CVarXiv:2108.02938v22021Understanding data augmentation for classification: when to warp?
Sebastien C. Wong, Adam Gatt, Victor Stamatescu +1
cs.CVarXiv:1609.08764v22016DeepID3: Face Recognition with Very Deep Neural Networks
Yi Sun, Ding Liang, Xiaogang Wang +1
cs.CVarXiv:1502.00873v12015Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training
Gen Li, Nan Duan, Yuejian Fang +3
cs.CVarXiv:1908.06066v32019Making Convolutional Networks Shift-Invariant Again
Richard Zhang
cs.CVcs.LGarXiv:1904.11486v22019PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference
Xiaofeng Mao, Shaohao Rui, Kaining Ying +4
cs.CVcs.AIarXiv:2603.25730v12026Sparsity Invariant CNNs
Jonas Uhrig, Nick Schneider, Lukas Schneider +3
cs.CVarXiv:1708.06500v22017Semi-Supervised Semantic Segmentation with Cross-Consistency Training
Yassine Ouali, Céline Hudelot, Myriam Tami
cs.CVarXiv:2003.09005v32020MAttNet: Modular Attention Network for Referring Expression Comprehension
Licheng Yu, Zhe Lin, Xiaohui Shen +4
cs.CVcs.AIcs.CLarXiv:1801.08186v32018Zero-Shot Learning -- The Good, the Bad and the Ugly
Yongqin Xian, Bernt Schiele, Zeynep Akata
cs.CVarXiv:1703.04394v22017Cascade Cost Volume for High-Resolution Multi-View Stereo and Stereo Matching
Xiaodong Gu, Zhiwen Fan, Zuozhuo Dai +3
cs.CVarXiv:1912.06378v32019Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
Andrej Karpathy, Armand Joulin, Li Fei-Fei
cs.CVcs.CLcs.LGarXiv:1406.5679v12014TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers
Bin Yu, Shijie Lian, Xiaopeng Lin +8
cs.ROcs.CVarXiv:2601.14133v22026MViTv2: Improved Multiscale Vision Transformers for Classification and Detection
Yanghao Li, Chao-Yuan Wu, Haoqi Fan +4
cs.CVarXiv:2112.01526v22021KITTI-360: A Novel Dataset and Benchmarks for Urban Scene Understanding in 2D and 3D
Yiyi Liao, Jun Xie, Andreas Geiger
cs.CVarXiv:2109.13410v22021BoT-SORT: Robust Associations Multi-Pedestrian Tracking
Nir Aharon, Roy Orfaig, Ben-Zion Bobrovsky
cs.CVarXiv:2206.14651v22022SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Boyuan Chen, Zhuo Xu, Sean Kirmani +6
cs.CVcs.CLcs.LGarXiv:2401.12168v12024Repurposing Geometric Foundation Models for Multi-view Diffusion
Wooseok Jang, Seonghu Jeon, Jisang Han +5
cs.CVarXiv:2603.22275v12026CLIPort: What and Where Pathways for Robotic Manipulation
Mohit Shridhar, Lucas Manuelli, Dieter Fox
cs.ROcs.CLcs.CVarXiv:2109.12098v12021MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
Yejin Kim, Wilbert Pumacay, Omar Rayyan +23
cs.ROcs.AIcs.CVarXiv:2602.11337v22026One-Shot Video Object Segmentation
Sergi Caelles, Kevis-Kokitsi Maninis, Jordi Pont-Tuset +3
cs.CVarXiv:1611.05198v42016Translating Videos to Natural Language Using Deep Recurrent Neural Networks
Subhashini Venugopalan, Huijuan Xu, Jeff Donahue +3
cs.CVcs.CLarXiv:1412.4729v32014Designing a Practical Degradation Model for Deep Blind Image Super-Resolution
Kai Zhang, Jingyun Liang, Luc Van Gool +1
eess.IVcs.CVarXiv:2103.14006v22021SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Bohao Li, Rui Wang, Guangzhi Wang +3
cs.CLcs.CVarXiv:2307.16125v22023Temporal Action Detection with Structured Segment Networks
Yue Zhao, Yuanjun Xiong, Limin Wang +3
cs.CVarXiv:1704.06228v22017"Zero-Shot" Super-Resolution using Deep Internal Learning
Assaf Shocher, Nadav Cohen, Michal Irani
cs.CVcs.LGcs.NEarXiv:1712.06087v12017Coarse-to-Fine Volumetric Prediction for Single-Image 3D Human Pose
Georgios Pavlakos, Xiaowei Zhou, Konstantinos G. Derpanis +1
cs.CVarXiv:1611.07828v22016Leveraging Frequency Analysis for Deep Fake Image Recognition
Joel Frank, Thorsten Eisenhofer, Lea Schönherr +3
cs.CVeess.IVarXiv:2003.08685v32020Land-Cover Classification with High-Resolution Remote Sensing Images Using Transferable Deep Models
Xin-Yi Tong, Gui-Song Xia, Qikai Lu +4
cs.CVarXiv:1807.05713v32018Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models
Meiqi Wu, Zhixin Cai, Fufangchen Zhao +13
cs.CVarXiv:2603.22212v12026GLIGEN: Open-Set Grounded Text-to-Image Generation
Yuheng Li, Haotian Liu, Qingyang Wu +5
cs.CVcs.AIcs.CLarXiv:2301.07093v22023Invariant Information Clustering for Unsupervised Image Classification and Segmentation
Xu Ji, João F. Henriques, Andrea Vedaldi
cs.CVcs.LGarXiv:1807.06653v42018Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding
Xianjin Wu, Dingkang Liang, Tianrui Feng +5
cs.CVcs.ROarXiv:2603.19235v32026PhyCritic: Multimodal Critic Models for Physical AI
Tianyi Xiong, Shihao Wang, Guilin Liu +5
cs.CVarXiv:2602.11124v12026Anomaly Detection via Reverse Distillation from One-Class Embedding
Hanqiu Deng, Xingyu Li
cs.CVarXiv:2201.10703v22022Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Keqin Chen, Zhao Zhang, Weili Zeng +3
cs.CVarXiv:2306.15195v22023Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition
Max Jaderberg, Karen Simonyan, Andrea Vedaldi +1
cs.CVarXiv:1406.2227v42014Learning Continuous Image Representation with Local Implicit Image Function
Yinbo Chen, Sifei Liu, Xiaolong Wang
cs.CVcs.LGarXiv:2012.09161v22020Dynamic Head: Unifying Object Detection Heads with Attentions
Xiyang Dai, Yinpeng Chen, Bin Xiao +4
cs.CVarXiv:2106.08322v12021StrongSORT: Make DeepSORT Great Again
Yunhao Du, Zhicheng Zhao, Yang Song +4
cs.CVarXiv:2202.13514v22022Pros and Cons of GAN Evaluation Measures
Ali Borji
cs.CVarXiv:1802.03446v52018Making Reconstruction FID Predictive of Diffusion Generation FID
Tongda Xu, Mingwei He, Shady Abu-Hussein +6
cs.CVcs.LGarXiv:2603.05630v22026Uninformed Students: Student-Teacher Anomaly Detection with Discriminative Latent Embeddings
Paul Bergmann, Michael Fauser, David Sattlegger +1
cs.CVarXiv:1911.02357v22019K-Planes: Explicit Radiance Fields in Space, Time, and Appearance
Sara Fridovich-Keil, Giacomo Meanti, Frederik Warburg +2
cs.CVarXiv:2301.10241v22023