Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

17,161 to 17,220 of 18,827

  1. VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation

    Yiming Zhao, Yu Zeng, Wenxuan Huang +11

    cs.CVcs.AIcs.HCarXiv:2605.16079v12026
  2. Neural Sparse Voxel Fields

    Lingjie Liu, Jiatao Gu, Kyaw Zaw Lin +2

    cs.CVcs.GRcs.LGarXiv:2007.11571v22020
  3. Meta-Learning with Latent Embedding Optimization

    Andrei A. Rusu, Dushyant Rao, Jakub Sygnowski +4

    cs.LGcs.CVstat.MLarXiv:1807.05960v32018
  4. Geom-GCN: Geometric Graph Convolutional Networks

    Hongbin Pei, Bingzhe Wei, Kevin Chen-Chuan Chang +2

    cs.LGcs.CVstat.MLarXiv:2002.05287v22020
  5. A simple yet effective baseline for 3d human pose estimation

    Julieta Martinez, Rayat Hossain, Javier Romero +1

    cs.CVarXiv:1705.03098v22017
  6. Panoptic Feature Pyramid Networks

    Alexander Kirillov, Ross Girshick, Kaiming He +1

    cs.CVarXiv:1901.02446v22019
  7. Efficient Multi-Scale Attention Module with Cross-Spatial Learning

    Daliang Ouyang, Su He, Guozhong Zhang +4

    cs.CVcs.AIarXiv:2305.13563v22023
  8. Planning-oriented Autonomous Driving

    Yihan Hu, Jiazhi Yang, Li Chen +13

    cs.CVcs.ROarXiv:2212.10156v22022
  9. Receptive Field Block Net for Accurate and Fast Object Detection

    Songtao Liu, Di Huang, Yunhong Wang

    cs.CVarXiv:1711.07767v32017
  10. OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding

    Ruixiang Zhao, Jie Yang, Zijie Xin +4

    cs.CVarXiv:2605.18577v12026
  11. Semantic Generative Tuning for Unified Multimodal Models

    Songsong Yu, Yuxin Chen, Ying Shan +1

    cs.CVcs.AIarXiv:2605.18714v22026
  12. End-to-End Learning of Geometry and Context for Deep Stereo Regression

    Alex Kendall, Hayk Martirosyan, Saumitro Dasgupta +4

    cs.CVcs.NEarXiv:1703.04309v12017
  13. See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding

    Boyuan Sun, Bowen Yin, Yuanming Li +2

    cs.CVcs.AIcs.HCarXiv:2605.18018v12026
  14. Efficient Geometry-aware 3D Generative Adversarial Networks

    Eric R. Chan, Connor Z. Lin, Matthew A. Chan +9

    cs.CVcs.AIcs.GRarXiv:2112.07945v22021
  15. Score-CAM: Score-Weighted Visual Explanations for Convolutional Neural Networks

    Haofan Wang, Zifan Wang, Mengnan Du +5

    cs.CVarXiv:1910.01279v22019
  16. Matérn Noise for Triangulation-Agnostic Flow Matching on Meshes

    Tianshu Kuai, Arman Maesumi, Daniel Ritchie +1

    cs.GRcs.CVcs.LGarXiv:2605.19305v12026
  17. Fast 4D Mesh Generation by Spatio-Temporal Attention Chains

    Dvir Samuel, Yuval Atzmon, Gal Chechik +1

    cs.CVarXiv:2605.19786v12026
  18. ADVENT: Adversarial Entropy Minimization for Domain Adaptation in Semantic Segmentation

    Tuan-Hung Vu, Himalaya Jain, Maxime Bucher +2

    cs.CVarXiv:1811.12833v22018
  19. Hybrid Task Cascade for Instance Segmentation

    Kai Chen, Jiangmiao Pang, Jiaqi Wang +9

    cs.CVarXiv:1901.07518v22019
  20. Deep Layer Aggregation

    Fisher Yu, Dequan Wang, Evan Shelhamer +1

    cs.CVcs.LGarXiv:1707.06484v32017
  21. Lost in the Folds: When Cross-Validation Is Not a Deep Ensemble for Uncertainty Estimation

    Tristan Kirscher, Markus Bujotzek, Yannick Kirchhoff +5

    cs.CVcs.LGarXiv:2605.18329v22026
  22. Large Scale Incremental Learning

    Yue Wu, Yinpeng Chen, Lijuan Wang +4

    cs.CVarXiv:1905.13260v12019
  23. A Unified Multi-scale Deep Convolutional Neural Network for Fast Object Detection

    Zhaowei Cai, Quanfu Fan, Rogerio S. Feris +1

    cs.CVarXiv:1607.07155v12016
  24. Reproducible scaling laws for contrastive language-image learning

    Mehdi Cherti, Romain Beaumont, Ross Wightman +6

    cs.LGcs.AIcs.CVarXiv:2212.07143v22022
  25. Learning Rich Features from RGB-D Images for Object Detection and Segmentation

    Saurabh Gupta, Ross Girshick, Pablo Arbeláez +1

    cs.CVcs.ROarXiv:1407.5736v12014
  26. Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models

    Guangzhi Xiong, Qiao Jin, Sanchit Sinha +2

    cs.CVcs.AIcs.CLarXiv:2605.20158v12026
  27. CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing

    Haobo Hu, Xiangwu Guo, Zhiheng Chen +4

    cs.CVcs.AIcs.GRarXiv:2605.19484v12026
  28. Q-ARVD: Quantizing Autoregressive Video Diffusion Models

    Siao Tang, Xinyin Ma, Gongfan Fang +2

    cs.CVarXiv:2605.21072v12026
  29. From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models

    Juncheng Wu, Hardy Chen, Haoqin Tu +6

    cs.CLcs.CVarXiv:2605.20177v12026
  30. Disentangling Sampling from Training Budget in Class-Imbalanced CT Body Composition Segmentation

    Iason Skylitsis, Dimitrios Karkalousos, Ivana Išgum

    eess.IVcs.AIcs.CVarXiv:2605.20405v12026
  31. Do Better ImageNet Models Transfer Better?

    Simon Kornblith, Jonathon Shlens, Quoc V. Le

    cs.CVcs.LGstat.MLarXiv:1805.08974v32018
  32. Synthetic Data for Text Localisation in Natural Images

    Ankush Gupta, Andrea Vedaldi, Andrew Zisserman

    cs.CVarXiv:1604.06646v12016
  33. AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment

    Kuei-Chun Kao, Daixuan Huo, Yuanhao Ban +1

    cs.AIcs.CVcs.LGarXiv:2605.17602v22026
  34. A Survey on Multimodal Large Language Models

    Shukang Yin, Chaoyou Fu, Sirui Zhao +4

    cs.CVcs.AIcs.CLarXiv:2306.13549v42023
  35. CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition

    Hongji Yang, Songlian Li, Yucheng Zhou +4

    cs.CVarXiv:2605.19995v12026
  36. Speeding up Convolutional Neural Networks with Low Rank Expansions

    Max Jaderberg, Andrea Vedaldi, Andrew Zisserman

    cs.CVarXiv:1405.3866v12014
  37. Decision-Based Adversarial Attacks: Reliable Attacks Against Black-Box Machine Learning Models

    Wieland Brendel, Jonas Rauber, Matthias Bethge

    stat.MLcs.CRcs.CVarXiv:1712.04248v22017
  38. Learning to See in the Dark

    Chen Chen, Qifeng Chen, Jia Xu +1

    cs.CVcs.GRcs.LGarXiv:1805.01934v12018
  39. Minimalist Visual Inertial Odometry

    Francesco Pasti, Jeremy Klotz, Nicola Bellotto +1

    cs.ROcs.CVcs.LGarXiv:2605.19990v12026
  40. Data-Efficient Image Recognition with Contrastive Predictive Coding

    Olivier J. Hénaff, Aravind Srinivas, Jeffrey De Fauw +4

    cs.CVcs.LGarXiv:1905.09272v32019
  41. Platonic Representations in the Human Brain: Unsupervised Recovery of Universal Geometry

    Pablo Marcos-Manchón, Rishi Jha, Lluís Fuentemilla

    q-bio.NCcs.CVarXiv:2605.20496v12026
  42. PCANet: A Simple Deep Learning Baseline for Image Classification?

    Tsung-Han Chan, Kui Jia, Shenghua Gao +3

    cs.CVcs.LGcs.NEarXiv:1404.3606v22014
  43. Perceiver: General Perception with Iterative Attention

    Andrew Jaegle, Felix Gimeno, Andrew Brock +3

    cs.CVcs.AIcs.LGarXiv:2103.03206v22021
  44. Cross-stitch Networks for Multi-task Learning

    Ishan Misra, Abhinav Shrivastava, Abhinav Gupta +1

    cs.CVcs.LGarXiv:1604.03539v12016
  45. IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools

    Rongbin Tan, Fangfang Lin, Zhenlong Yuan +10

    cs.CVarXiv:2605.20682v12026
  46. BranchyNet: Fast Inference via Early Exiting from Deep Neural Networks

    Surat Teerapittayanon, Bradley McDanel, H. T. Kung

    cs.NEcs.CVcs.LGarXiv:1709.01686v12017
  47. Generating Videos with Scene Dynamics

    Carl Vondrick, Hamed Pirsiavash, Antonio Torralba

    cs.CVcs.GRcs.LGarXiv:1609.02612v32016
  48. What Makes for Good Views for Contrastive Learning?

    Yonglong Tian, Chen Sun, Ben Poole +3

    cs.CVcs.LGarXiv:2005.10243v32020
  49. Fine-tuning CNN Image Retrieval with No Human Annotation

    Filip Radenović, Giorgos Tolias, Ondřej Chum

    cs.CVarXiv:1711.02512v22017
  50. Domain Adaptive Faster R-CNN for Object Detection in the Wild

    Yuhua Chen, Wen Li, Christos Sakaridis +2

    cs.CVarXiv:1803.03243v12018
  51. RISE: Randomized Input Sampling for Explanation of Black-box Models

    Vitali Petsiuk, Abir Das, Kate Saenko

    cs.CVarXiv:1806.07421v32018
  52. This Looks Like That: Deep Learning for Interpretable Image Recognition

    Chaofan Chen, Oscar Li, Chaofan Tao +3

    cs.LGcs.AIcs.CVarXiv:1806.10574v52018
  53. Joint 3D Proposal Generation and Object Detection from View Aggregation

    Jason Ku, Melissa Mozifian, Jungwook Lee +2

    cs.CVarXiv:1712.02294v42017
  54. MUSIQ: Multi-scale Image Quality Transformer

    Junjie Ke, Qifei Wang, Yilin Wang +2

    cs.CVarXiv:2108.05997v12021
  55. MotiMotion: Motion-Controlled Video Generation with Visual Reasoning

    Lee Hsin-Ying, Hanwen Jiang, Yiqun Mei +3

    cs.CVarXiv:2605.22818v12026
  56. DocVQA: A Dataset for VQA on Document Images

    Minesh Mathew, Dimosthenis Karatzas, C. V. Jawahar

    cs.CVcs.IRarXiv:2007.00398v32020
  57. EMMA: Extracting Multiple physical parameters from Multimodal Data

    Farhat Shaikh, Ayan Banerjee, Sandeep Gupta

    cs.CVarXiv:2605.24047v12026
  58. SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation

    Xiaolong Zhou, Yifei Liu, Ziyang Gong +8

    cs.CVcs.CLarXiv:2605.22536v22026
  59. DecQ: Detail-Condensing Queries for Enhanced Reconstruction and Generation in Representation Autoencoders

    Tianhang Wang, Yitong Chen, Wei Song +3

    cs.CVarXiv:2605.22777v12026
  60. HunyuanVideo: A Systematic Framework For Large Video Generative Models

    Weijie Kong, Qi Tian, Zijian Zhang +49

    cs.CVarXiv:2412.03603v62024