Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
17,161 to 17,220 of 18,827
VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation
Yiming Zhao, Yu Zeng, Wenxuan Huang +11
cs.CVcs.AIcs.HCarXiv:2605.16079v12026Neural Sparse Voxel Fields
Lingjie Liu, Jiatao Gu, Kyaw Zaw Lin +2
cs.CVcs.GRcs.LGarXiv:2007.11571v22020Meta-Learning with Latent Embedding Optimization
Andrei A. Rusu, Dushyant Rao, Jakub Sygnowski +4
cs.LGcs.CVstat.MLarXiv:1807.05960v32018Geom-GCN: Geometric Graph Convolutional Networks
Hongbin Pei, Bingzhe Wei, Kevin Chen-Chuan Chang +2
cs.LGcs.CVstat.MLarXiv:2002.05287v22020A simple yet effective baseline for 3d human pose estimation
Julieta Martinez, Rayat Hossain, Javier Romero +1
cs.CVarXiv:1705.03098v22017Panoptic Feature Pyramid Networks
Alexander Kirillov, Ross Girshick, Kaiming He +1
cs.CVarXiv:1901.02446v22019Efficient Multi-Scale Attention Module with Cross-Spatial Learning
Daliang Ouyang, Su He, Guozhong Zhang +4
cs.CVcs.AIarXiv:2305.13563v22023Planning-oriented Autonomous Driving
Yihan Hu, Jiazhi Yang, Li Chen +13
cs.CVcs.ROarXiv:2212.10156v22022Receptive Field Block Net for Accurate and Fast Object Detection
Songtao Liu, Di Huang, Yunhong Wang
cs.CVarXiv:1711.07767v32017OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
Ruixiang Zhao, Jie Yang, Zijie Xin +4
cs.CVarXiv:2605.18577v12026Semantic Generative Tuning for Unified Multimodal Models
Songsong Yu, Yuxin Chen, Ying Shan +1
cs.CVcs.AIarXiv:2605.18714v22026End-to-End Learning of Geometry and Context for Deep Stereo Regression
Alex Kendall, Hayk Martirosyan, Saumitro Dasgupta +4
cs.CVcs.NEarXiv:1703.04309v12017See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
Boyuan Sun, Bowen Yin, Yuanming Li +2
cs.CVcs.AIcs.HCarXiv:2605.18018v12026Efficient Geometry-aware 3D Generative Adversarial Networks
Eric R. Chan, Connor Z. Lin, Matthew A. Chan +9
cs.CVcs.AIcs.GRarXiv:2112.07945v22021Score-CAM: Score-Weighted Visual Explanations for Convolutional Neural Networks
Haofan Wang, Zifan Wang, Mengnan Du +5
cs.CVarXiv:1910.01279v22019Matérn Noise for Triangulation-Agnostic Flow Matching on Meshes
Tianshu Kuai, Arman Maesumi, Daniel Ritchie +1
cs.GRcs.CVcs.LGarXiv:2605.19305v12026Fast 4D Mesh Generation by Spatio-Temporal Attention Chains
Dvir Samuel, Yuval Atzmon, Gal Chechik +1
cs.CVarXiv:2605.19786v12026ADVENT: Adversarial Entropy Minimization for Domain Adaptation in Semantic Segmentation
Tuan-Hung Vu, Himalaya Jain, Maxime Bucher +2
cs.CVarXiv:1811.12833v22018Hybrid Task Cascade for Instance Segmentation
Kai Chen, Jiangmiao Pang, Jiaqi Wang +9
cs.CVarXiv:1901.07518v22019Deep Layer Aggregation
Fisher Yu, Dequan Wang, Evan Shelhamer +1
cs.CVcs.LGarXiv:1707.06484v32017Lost in the Folds: When Cross-Validation Is Not a Deep Ensemble for Uncertainty Estimation
Tristan Kirscher, Markus Bujotzek, Yannick Kirchhoff +5
cs.CVcs.LGarXiv:2605.18329v22026Large Scale Incremental Learning
Yue Wu, Yinpeng Chen, Lijuan Wang +4
cs.CVarXiv:1905.13260v12019A Unified Multi-scale Deep Convolutional Neural Network for Fast Object Detection
Zhaowei Cai, Quanfu Fan, Rogerio S. Feris +1
cs.CVarXiv:1607.07155v12016Reproducible scaling laws for contrastive language-image learning
Mehdi Cherti, Romain Beaumont, Ross Wightman +6
cs.LGcs.AIcs.CVarXiv:2212.07143v22022Learning Rich Features from RGB-D Images for Object Detection and Segmentation
Saurabh Gupta, Ross Girshick, Pablo Arbeláez +1
cs.CVcs.ROarXiv:1407.5736v12014Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models
Guangzhi Xiong, Qiao Jin, Sanchit Sinha +2
cs.CVcs.AIcs.CLarXiv:2605.20158v12026CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing
Haobo Hu, Xiangwu Guo, Zhiheng Chen +4
cs.CVcs.AIcs.GRarXiv:2605.19484v12026Q-ARVD: Quantizing Autoregressive Video Diffusion Models
Siao Tang, Xinyin Ma, Gongfan Fang +2
cs.CVarXiv:2605.21072v12026From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
Juncheng Wu, Hardy Chen, Haoqin Tu +6
cs.CLcs.CVarXiv:2605.20177v12026Disentangling Sampling from Training Budget in Class-Imbalanced CT Body Composition Segmentation
Iason Skylitsis, Dimitrios Karkalousos, Ivana Išgum
eess.IVcs.AIcs.CVarXiv:2605.20405v12026Do Better ImageNet Models Transfer Better?
Simon Kornblith, Jonathon Shlens, Quoc V. Le
cs.CVcs.LGstat.MLarXiv:1805.08974v32018Synthetic Data for Text Localisation in Natural Images
Ankush Gupta, Andrea Vedaldi, Andrew Zisserman
cs.CVarXiv:1604.06646v12016AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment
Kuei-Chun Kao, Daixuan Huo, Yuanhao Ban +1
cs.AIcs.CVcs.LGarXiv:2605.17602v22026A Survey on Multimodal Large Language Models
Shukang Yin, Chaoyou Fu, Sirui Zhao +4
cs.CVcs.AIcs.CLarXiv:2306.13549v42023CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition
Hongji Yang, Songlian Li, Yucheng Zhou +4
cs.CVarXiv:2605.19995v12026Speeding up Convolutional Neural Networks with Low Rank Expansions
Max Jaderberg, Andrea Vedaldi, Andrew Zisserman
cs.CVarXiv:1405.3866v12014Decision-Based Adversarial Attacks: Reliable Attacks Against Black-Box Machine Learning Models
Wieland Brendel, Jonas Rauber, Matthias Bethge
stat.MLcs.CRcs.CVarXiv:1712.04248v22017Learning to See in the Dark
Chen Chen, Qifeng Chen, Jia Xu +1
cs.CVcs.GRcs.LGarXiv:1805.01934v12018Minimalist Visual Inertial Odometry
Francesco Pasti, Jeremy Klotz, Nicola Bellotto +1
cs.ROcs.CVcs.LGarXiv:2605.19990v12026Data-Efficient Image Recognition with Contrastive Predictive Coding
Olivier J. Hénaff, Aravind Srinivas, Jeffrey De Fauw +4
cs.CVcs.LGarXiv:1905.09272v32019Platonic Representations in the Human Brain: Unsupervised Recovery of Universal Geometry
Pablo Marcos-Manchón, Rishi Jha, Lluís Fuentemilla
q-bio.NCcs.CVarXiv:2605.20496v12026PCANet: A Simple Deep Learning Baseline for Image Classification?
Tsung-Han Chan, Kui Jia, Shenghua Gao +3
cs.CVcs.LGcs.NEarXiv:1404.3606v22014Perceiver: General Perception with Iterative Attention
Andrew Jaegle, Felix Gimeno, Andrew Brock +3
cs.CVcs.AIcs.LGarXiv:2103.03206v22021Cross-stitch Networks for Multi-task Learning
Ishan Misra, Abhinav Shrivastava, Abhinav Gupta +1
cs.CVcs.LGarXiv:1604.03539v12016IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools
Rongbin Tan, Fangfang Lin, Zhenlong Yuan +10
cs.CVarXiv:2605.20682v12026BranchyNet: Fast Inference via Early Exiting from Deep Neural Networks
Surat Teerapittayanon, Bradley McDanel, H. T. Kung
cs.NEcs.CVcs.LGarXiv:1709.01686v12017Generating Videos with Scene Dynamics
Carl Vondrick, Hamed Pirsiavash, Antonio Torralba
cs.CVcs.GRcs.LGarXiv:1609.02612v32016What Makes for Good Views for Contrastive Learning?
Yonglong Tian, Chen Sun, Ben Poole +3
cs.CVcs.LGarXiv:2005.10243v32020Fine-tuning CNN Image Retrieval with No Human Annotation
Filip Radenović, Giorgos Tolias, Ondřej Chum
cs.CVarXiv:1711.02512v22017Domain Adaptive Faster R-CNN for Object Detection in the Wild
Yuhua Chen, Wen Li, Christos Sakaridis +2
cs.CVarXiv:1803.03243v12018RISE: Randomized Input Sampling for Explanation of Black-box Models
Vitali Petsiuk, Abir Das, Kate Saenko
cs.CVarXiv:1806.07421v32018This Looks Like That: Deep Learning for Interpretable Image Recognition
Chaofan Chen, Oscar Li, Chaofan Tao +3
cs.LGcs.AIcs.CVarXiv:1806.10574v52018Joint 3D Proposal Generation and Object Detection from View Aggregation
Jason Ku, Melissa Mozifian, Jungwook Lee +2
cs.CVarXiv:1712.02294v42017MUSIQ: Multi-scale Image Quality Transformer
Junjie Ke, Qifei Wang, Yilin Wang +2
cs.CVarXiv:2108.05997v12021MotiMotion: Motion-Controlled Video Generation with Visual Reasoning
Lee Hsin-Ying, Hanwen Jiang, Yiqun Mei +3
cs.CVarXiv:2605.22818v12026DocVQA: A Dataset for VQA on Document Images
Minesh Mathew, Dimosthenis Karatzas, C. V. Jawahar
cs.CVcs.IRarXiv:2007.00398v32020EMMA: Extracting Multiple physical parameters from Multimodal Data
Farhat Shaikh, Ayan Banerjee, Sandeep Gupta
cs.CVarXiv:2605.24047v12026SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation
Xiaolong Zhou, Yifei Liu, Ziyang Gong +8
cs.CVcs.CLarXiv:2605.22536v22026DecQ: Detail-Condensing Queries for Enhanced Reconstruction and Generation in Representation Autoencoders
Tianhang Wang, Yitong Chen, Wei Song +3
cs.CVarXiv:2605.22777v12026HunyuanVideo: A Systematic Framework For Large Video Generative Models
Weijie Kong, Qi Tian, Zijian Zhang +49
cs.CVarXiv:2412.03603v62024