Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

14,761 to 14,820 of 18,867

  1. Variational Adversarial Active Learning

    Samarth Sinha, Sayna Ebrahimi, Trevor Darrell

    cs.LGcs.CVstat.MLarXiv:1904.00370v32019
  2. How Do Vision Transformers Work?

    Namuk Park, Songkuk Kim

    cs.CVcs.LGarXiv:2202.06709v42022
  3. Autoregressive Image Generation without Vector Quantization

    Tianhong Li, Yonglong Tian, He Li +2

    cs.CVarXiv:2406.11838v32024
  4. Probabilistic Regression for Visual Tracking

    Martin Danelljan, Luc Van Gool, Radu Timofte

    cs.CVcs.LGarXiv:2003.12565v12020
  5. OmniLottie: Generating Vector Animations via Parameterized Lottie Tokens

    Yiying Yang, Wei Cheng, Sijin Chen +5

    cs.CVarXiv:2603.02138v12026
  6. GroupEnsemble: Efficient Uncertainty Estimation for DETR-based Object Detection

    Yutong Yang, Katarina Popović, Julian Wiederer +3

    cs.CVarXiv:2603.01847v12026
  7. SCAN: Learning to Classify Images without Labels

    Wouter Van Gansbeke, Simon Vandenhende, Stamatios Georgoulis +2

    cs.CVcs.LGarXiv:2005.12320v22020
  8. Learning Deep Models for Face Anti-Spoofing: Binary or Auxiliary Supervision

    Yaojie Liu, Amin Jourabloo, Xiaoming Liu

    cs.CVarXiv:1803.11097v12018
  9. Multi-Context Attention for Human Pose Estimation

    Xiao Chu, Wei Yang, Wanli Ouyang +3

    cs.CVarXiv:1702.07432v12017
  10. Comparative Assessment of Deep Learning Architectures for Underwater Subsurface Kelp Forest Segmentation with The Kelp-o-Tron

    Sundarabalan Balasubramanian, César Borja, Ana C. Murillo +5

    cs.CVarXiv:2608.24594v12026
  11. Fast and Furious: Real Time End-to-End 3D Detection, Tracking and Motion Forecasting with a Single Convolutional Net

    Wenjie Luo, Bin Yang, Raquel Urtasun

    cs.CVarXiv:2012.12395v12020
  12. MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models

    Zhongxi Wang, Yueqian Lin, Jingyang Zhang +2

    cs.LGcs.CLcs.CVarXiv:2603.02482v12026
  13. PointCLIP: Point Cloud Understanding by CLIP

    Renrui Zhang, Ziyu Guo, Wei Zhang +6

    cs.CVcs.AIcs.ROarXiv:2112.02413v12021
  14. Multimodal Deep Learning for Robust RGB-D Object Recognition

    Andreas Eitel, Jost Tobias Springenberg, Luciano Spinello +2

    cs.CVcs.LGcs.NEarXiv:1507.06821v22015
  15. Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation

    Chao Li, Tianhong Li, Sai Vidyaranya Nuthalapati +9

    cs.CVcs.LGarXiv:2603.02667v22026
  16. Rethinking the Faster R-CNN Architecture for Temporal Action Localization

    Yu-Wei Chao, Sudheendra Vijayanarasimhan, Bryan Seybold +3

    cs.CVarXiv:1804.07667v12018
  17. Kling-MotionControl Technical Report

    Kling Team, Jialu Chen, Yikang Ding +21

    cs.CVarXiv:2603.03160v12026
  18. Understanding Deep Convolutional Networks

    Stéphane Mallat

    stat.MLcs.CVcs.LGarXiv:1601.04920v12016
  19. S$^3$FD: Single Shot Scale-invariant Face Detector

    Shifeng Zhang, Xiangyu Zhu, Zhen Lei +3

    cs.CVarXiv:1708.05237v32017
  20. ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors

    Zihao Huang, Tianqi Liu, Zhaoxi Chen +7

    cs.CVarXiv:2603.04338v12026
  21. Bidirectional Learning for Domain Adaptation of Semantic Segmentation

    Yunsheng Li, Lu Yuan, Nuno Vasconcelos

    cs.CVarXiv:1904.10620v12019
  22. MRI-based Deep Radiomic Phenotyping of Neuromuscular Disorders: A Topology-driven Characterization

    Martyna Żur, Łukasz Piórecki, Marek Socha +5

    cs.CVarXiv:2608.24415v12026
  23. MASQuant: Modality-Aware Smoothing Quantization for Multimodal Large Language Models

    Lulu Hu, Wenhu Xiao, Xin Chen +4

    cs.CVarXiv:2603.04800v12026
  24. Augmentation for small object detection

    Mate Kisantal, Zbigniew Wojna, Jakub Murawski +2

    cs.CVarXiv:1902.07296v12019
  25. PolypSteer: Counterfactual Endoscopic Synthesis via Training-Free Activation Steering

    Trong-Thang Pham, Loc Nguyen, Anh Nguyen +2

    cs.CVcs.AIarXiv:2603.07066v22026
  26. PureCC: Pure Learning for Text-to-Image Concept Customization

    Zhichao Liao, Xiaole Xian, Qingyu Li +7

    cs.CVarXiv:2603.07561v22026
  27. Absorbing Gradient Conflicts: Modeling Semantic Variance via Kent Distributions for Cross-Modal Hashing

    Hengjie Zhu, Dayan Wu, Zihao Zhang +5

    cs.CVarXiv:2608.24010v12026
  28. Seeing biodiversity: perspectives in machine learning for wildlife conservation

    Devis Tuia, Benjamin Kellenberger, Sara Beery +15

    cs.LGcs.CVarXiv:2110.12951v12021
  29. TeamHOI: Learning a Unified Policy for Cooperative Human-Object Interactions with Any Team Size

    Stefan Lionar, Gim Hee Lee

    cs.CVcs.GRcs.MAarXiv:2603.07988v12026
  30. WaDi: Weight Direction-aware Distillation for One-step Image Synthesis

    Lei Wang, Yang Cheng, Senmao Li +3

    cs.CVarXiv:2603.08258v12026
  31. Multi-Task Multi-Sensor Fusion for 3D Object Detection

    Ming Liang, Bin Yang, Yun Chen +2

    cs.CVarXiv:2012.12397v12020
  32. HyPER-GAN: Hybrid Patch-Based Image-to-Image Translation for Real-Time Photorealism Enhancement in Game Engines

    Stefanos Pasios, Nikos Nikolaidis

    cs.CVarXiv:2603.10604v32026
  33. You Only Look One-level Feature

    Qiang Chen, Yingming Wang, Tong Yang +3

    cs.CVarXiv:2103.09460v12021
  34. HINet: Half Instance Normalization Network for Image Restoration

    Liangyu Chen, Xin Lu, Jie Zhang +2

    eess.IVcs.CVarXiv:2105.06086v22021
  35. FineRMoE: Dimension Expansion for Finer-Grained Expert with Its Upcycling Approach

    Ning Liao, Xiaoxing Wang, Xiaohan Qin +1

    cs.CVcs.AIarXiv:2603.13364v12026
  36. What Makes Training Multi-Modal Classification Networks Hard?

    Weiyao Wang, Du Tran, Matt Feiszli

    cs.CVcs.LGarXiv:1905.12681v52019
  37. Sub-Image Anomaly Detection with Deep Pyramid Correspondences

    Niv Cohen, Yedid Hoshen

    cs.CVcs.LGarXiv:2005.02357v32020
  38. Semi-Supervised Deep Learning for Monocular Depth Map Prediction

    Yevhen Kuznietsov, Jörg Stückler, Bastian Leibe

    cs.CVarXiv:1702.02706v32017
  39. Heterogeneous Federated Learning: State-of-the-art and Research Challenges

    Mang Ye, Xiuwen Fang, Bo Du +2

    cs.LGcs.AIcs.CVarXiv:2307.10616v22023
  40. Shape Completion using 3D-Encoder-Predictor CNNs and Shape Synthesis

    Angela Dai, Charles Ruizhongtai Qi, Matthias Nießner

    cs.CVarXiv:1612.00101v22016
  41. Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering

    Yura Choi, Roy Miles, Rolandos Alexandros Potamias +3

    cs.CVarXiv:2603.12533v22026
  42. MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning

    Haozhan Shen, Shilin Yan, Hongwei Xue +5

    cs.CVarXiv:2603.12266v12026
  43. Sparse Multi-Stage Expert-Agent Routing for Complex Clinical Reasoning

    Sike Xiang, Shuang Chen, Qian sun +3

    cs.CVarXiv:2608.21948v12026
  44. Taking Shortcuts for Categorical VQA Using Super Neurons

    Pierre Musacchio, Jaeyi Jeong, Dahun Kim +1

    cs.CVcs.AIcs.LGarXiv:2603.10781v12026
  45. Semantic Slots for Video Object-Centric Learning

    Khalil Sabri, Guillaume-Alexandre Bilodeau, Nicolas Saunier +1

    cs.CVarXiv:2608.21636v12026
  46. COT-FM: Cluster-wise Optimal Transport Flow Matching

    Chiensheng Chiang, Kuan-Hsun Tu, Jia-Wei Liao +2

    cs.CVcs.LGcs.ROarXiv:2603.13395v12026
  47. Strip Pooling: Rethinking Spatial Pooling for Scene Parsing

    Qibin Hou, Li Zhang, Ming-Ming Cheng +1

    cs.CVarXiv:2003.13328v12020
  48. COMIC: Agentic Sketch Comedy Generation

    Susung Hong, Brian Curless, Ira Kemelmacher-Shlizerman +1

    cs.CVcs.AIcs.CLarXiv:2603.11048v12026
  49. Generalisation in humans and deep neural networks

    Robert Geirhos, Carlos R. Medina Temme, Jonas Rauber +3

    cs.CVcs.AIcs.LGarXiv:1808.08750v32018
  50. Med3D: Transfer Learning for 3D Medical Image Analysis

    Sihong Chen, Kai Ma, Yefeng Zheng

    cs.CVarXiv:1904.00625v42019
  51. OPV2V: An Open Benchmark Dataset and Fusion Pipeline for Perception with Vehicle-to-Vehicle Communication

    Runsheng Xu, Hao Xiang, Xin Xia +3

    cs.CVcs.ROarXiv:2109.07644v52021
  52. Detecting Oriented Text in Natural Images by Linking Segments

    Baoguang Shi, Xiang Bai, Serge Belongie

    cs.CVarXiv:1703.06520v32017
  53. On the Automatic Generation of Medical Imaging Reports

    Baoyu Jing, Pengtao Xie, Eric Xing

    cs.CLcs.CVarXiv:1711.08195v32017
  54. Pyramid Feature Attention Network for Saliency detection

    Ting Zhao, Xiangqian Wu

    cs.CVarXiv:1903.00179v22019
  55. Coherent Human-Scene Reconstruction from Multi-Person Multi-View Video in a Single Pass

    Sangmin Kim, Minhyuk Hwang, Geonho Cha +2

    cs.CVarXiv:2603.12789v22026
  56. Unsupervised Learning of Monocular Depth Estimation and Visual Odometry with Deep Feature Reconstruction

    Huangying Zhan, Ravi Garg, Chamara Saroj Weerasekera +3

    cs.CVarXiv:1803.03893v32018
  57. Make it SING: Analyzing Semantic Invariants in Classifiers

    Harel Yadid, Meir Yossef Levi, Roy Betser +1

    cs.CVeess.IVarXiv:2603.14610v22026
  58. MSM-Mem: A Universal Medical Structured Multimodal Memory Framework for Medical AI Agents

    Md Asaduzzaman Jabin, Khoa Le, Lin Zhao +1

    cs.LGcs.CVarXiv:2608.21810v12026
  59. Differentiable Augmentation for Data-Efficient GAN Training

    Shengyu Zhao, Zhijian Liu, Ji Lin +2

    cs.CVcs.GRcs.LGarXiv:2006.10738v42020
  60. MedReaMM: Evaluating Large Multimodal Models on Expert-Level Clinical Diagnostic Synthesis

    Lai Wei, Yuchao Chen, Zhenbiao Cao +5

    cs.CVarXiv:2608.22323v12026