Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

15,001 to 15,060 of 18,933

  1. ADMIL: Attention-Distilled Multiple Instance Learning for Selective Foundation Model Inference in Pathology

    Duncan Stothers, Ren-Chin Wu, William Lotter

    cs.CVcs.AIcs.LGarXiv:2608.22066v12026
  2. Cut, Paste and Learn: Surprisingly Easy Synthesis for Instance Detection

    Debidatta Dwibedi, Ishan Misra, Martial Hebert

    cs.CVarXiv:1708.01642v12017
  3. Gate Voltage Effect on Pulse Detection Efficiency of Perimeter-Gated SPADs

    Hunter Guthrie, Md Sakibur Sajal, Zexi Liu +1

    physics.ins-detcs.CVarXiv:2608.21371v12026
  4. VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

    Wenhai Wang, Zhe Chen, Xiaokang Chen +8

    cs.CVarXiv:2305.11175v22023
  5. Adapting Dense Vision-Language Relationships for Multi-label Classification with Partial Label

    Cheng Chen, Yifan Zhao, Jia Li

    cs.CVarXiv:2608.22313v12026
  6. Using Transfer Learning for Image-Based Cassava Disease Detection

    Amanda Ramcharan, Kelsee Baranowski, Peter McCloskey +3

    cs.CVcs.CYarXiv:1707.03717v22017
  7. RoboShape: Information-Theoretic Point Cloud Representations for Privacy-Aware Robot Perception

    Oguzhan Baser, Mirac Sozen, Kaan Kale +2

    cs.ROcs.AIcs.CVarXiv:2608.21380v12026
  8. Improved denoising diffusion probabilistic models with efficient non-diagonal covariance modeling

    Rui Xia, Ayan Das, Artem Artemev +3

    cs.CVstat.MLarXiv:2608.21972v12026
  9. Recurrent MVSNet for High-resolution Multi-view Stereo Depth Inference

    Yao Yao, Zixin Luo, Shiwei Li +3

    cs.CVarXiv:1902.10556v12019
  10. Differentiable Voxelization of Surface Representations

    Tobias Djuren, Ugo Finnendahl, Markus Worchel +2

    cs.GRcs.CVarXiv:2608.15934v12026
  11. PointNetVLAD: Deep Point Cloud Based Retrieval for Large-Scale Place Recognition

    Mikaela Angelina Uy, Gim Hee Lee

    cs.CVarXiv:1804.03492v32018
  12. 3D Point Cloud from Close-Range Photogrammetry for Defect Characterisation of Rubberised Concrete

    Jiacheng Liu, Mohammed Alnahhal, Ailar Hajimohammadi +3

    cs.CVarXiv:2608.21468v12026
  13. Executing your Commands via Motion Diffusion in Latent Space

    Xin Chen, Biao Jiang, Wen Liu +5

    cs.CVcs.GRarXiv:2212.04048v32022
  14. Retrieval-Augmented Visual Prompting: Guiding Foundation Models in Two-Photon Imaging

    Salvatore Calcagno, Marco Finocchiaro, Giovanni Bellitto +3

    cs.CVarXiv:2608.21970v12026
  15. On the Choice of Tensor Estimation for Corner Detection, Optical Flow and Denoising

    Freddie Åström, Michael Felsberg

    cs.CVarXiv:2608.22314v12026
  16. Searching for Activation Functions

    Prajit Ramachandran, Barret Zoph, Quoc V. Le

    cs.NEcs.CVcs.LGarXiv:1710.05941v22017
  17. MedGAN: Medical Image Translation using GANs

    Karim Armanious, Chenming Jiang, Marc Fischer +4

    cs.CVarXiv:1806.06397v22018
  18. Dual Attention Networks for Multimodal Reasoning and Matching

    Hyeonseob Nam, Jung-Woo Ha, Jeonghee Kim

    cs.CVarXiv:1611.00471v22016
  19. V2X-ViT: Vehicle-to-Everything Cooperative Perception with Vision Transformer

    Runsheng Xu, Hao Xiang, Zhengzhong Tu +3

    cs.CVarXiv:2203.10638v32022
  20. Unsupervised Deep Feature Extraction for Remote Sensing Image Classification

    Adriana Romero, Carlo Gatta, Gustau Camps-Valls

    cs.CVarXiv:1511.08131v12015
  21. MobileFaceNets: Efficient CNNs for Accurate Real-Time Face Verification on Mobile Devices

    Sheng Chen, Yang Liu, Xiang Gao +1

    cs.CVcs.LGarXiv:1804.07573v42018
  22. Targeted Iterative Filtering

    Freddie Åström, Michael Felsberg, George Baravdish +1

    cs.CVarXiv:2608.22299v12026
  23. Learning Sample-wise Rank-aware Interpolation Weights for Composed Visual Data Retrieval

    Boseung Jeong, Taegyu Park, Donghyeon Kwon +2

    cs.CVarXiv:2608.22500v12026
  24. On Tensor-Based PDEs and their Corresponding Variational Formulations with Application to Color Image Denoising

    Freddie Åström, George Baravdish, Michael Felsberg

    cs.CVarXiv:2608.22302v12026
  25. LiST: Local-Simplex Test-Time LoRA Fusion

    Yihua Shao, Jia Li, Siyu Chen +12

    cs.CVarXiv:2608.22370v12026
  26. Learning Spatio-Temporal Features with 3D Residual Networks for Action Recognition

    Kensho Hara, Hirokatsu Kataoka, Yutaka Satoh

    cs.CVarXiv:1708.07632v12017
  27. Focal Guidance: Unlocking Controllability from Semantic-Weak Layers in Video Diffusion Models

    Yuanyang Yin, Yufan Deng, Shenghai Yuan +3

    cs.CVarXiv:2601.07287v12026
  28. What makes ImageNet good for transfer learning?

    Minyoung Huh, Pulkit Agrawal, Alexei A. Efros

    cs.CVcs.AIcs.LGarXiv:1608.08614v22016
  29. No Fuss Distance Metric Learning using Proxies

    Yair Movshovitz-Attias, Alexander Toshev, Thomas K. Leung +2

    cs.CVarXiv:1703.07464v32017
  30. SkinFlow: Efficient Information Transmission for Open Dermatological Diagnosis via Dynamic Visual Encoding and Staged RL

    Lijun Liu, Linwei Chen, Zhishou Zhang +7

    cs.CVcs.AIarXiv:2601.09136v12026
  31. mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

    Qinghao Ye, Haiyang Xu, Jiabo Ye +7

    cs.CLcs.CVarXiv:2311.04257v22023
  32. CMX: Cross-Modal Fusion for RGB-X Semantic Segmentation with Transformers

    Jiaming Zhang, Huayao Liu, Kailun Yang +3

    cs.CVcs.ROeess.IVarXiv:2203.04838v52022
  33. Exploring the Limitations of Behavior Cloning for Autonomous Driving

    Felipe Codevilla, Eder Santana, Antonio M. López +1

    cs.CVcs.AIarXiv:1904.08980v12019
  34. Bee Detection and Tracking at Hive Entrance using YOLO11 and ByteTrack

    Thi Thu Thao Nguyen, Johannes Reschke

    cs.CVarXiv:2608.23213v12026
  35. Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing

    Dohun Lee, Chun-Hao Paul Huang, Xuelin Chen +3

    cs.CVcs.AIcs.LGarXiv:2601.16296v22026
  36. DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset

    Hengyu Shen, Tiancheng Gu, Bin Qin +10

    cs.CVcs.AIarXiv:2601.10305v32026
  37. 3D Shape Generation and Completion through Point-Voxel Diffusion

    Linqi Zhou, Yilun Du, Jiajun Wu

    cs.CVarXiv:2104.03670v32021
  38. FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects

    Bowen Wen, Wei Yang, Jan Kautz +1

    cs.CVcs.AIcs.ROarXiv:2312.08344v22023
  39. MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing

    Kai Zhang, Lingbo Mo, Wenhu Chen +2

    cs.CVcs.AIcs.CLarXiv:2306.10012v32023
  40. Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing

    Tingyu Song, Yanzhao Zhang, Mingxin Li +6

    cs.CVcs.CLcs.IRarXiv:2601.16125v12026
  41. O-HAZE: a dehazing benchmark with real hazy and haze-free outdoor images

    Codruta O. Ancuti, Cosmin Ancuti, Radu Timofte +1

    cs.CVarXiv:1804.05101v12018
  42. Structured Knowledge Distillation for Dense Prediction

    Yifan Liu, Changyong Shun, Jingdong Wang +1

    cs.CVarXiv:1903.04197v72019
  43. Grounding Isn't Knowing: Do VLMs Need Object Localization for Spatial Reasoning?

    Xiwei Liu, Yulong Li, Xinlin Zhuang +5

    cs.CVarXiv:2608.23074v12026
  44. Domain Generalization for Object Recognition with Multi-task Autoencoders

    Muhammad Ghifary, W. Bastiaan Kleijn, Mengjie Zhang +1

    cs.CVcs.AIcs.LGarXiv:1508.07680v12015
  45. Following Motion for Sequential Modeling in Video Frame Interpolation

    Jaehyun Park, Nam Ik Cho

    cs.CVarXiv:2608.22861v12026
  46. Artificial Intelligence in the Creative Industries: A Review

    Nantheera Anantrasirichai, David Bull

    cs.CVcs.AIcs.LGarXiv:2007.12391v62020
  47. TokenTrim: Inference-Time Token Pruning for Autoregressive Long Video Generation

    Ariel Shaulov, Eitan Shaar, Amit Edenzon +1

    cs.CVcs.AIarXiv:2602.00268v12026
  48. Ten Years of Pedestrian Detection, What Have We Learned?

    Rodrigo Benenson, Mohamed Omran, Jan Hosang +1

    cs.CVarXiv:1411.4304v12014
  49. Rotation-invariant convolutional neural networks for galaxy morphology prediction

    Sander Dieleman, Kyle W. Willett, Joni Dambre

    astro-ph.IMastro-ph.GAcs.CVarXiv:1503.07077v12015
  50. Seeing the Unseen: Semantic-in-Gaussian for Sparse-View 3D Generalization

    Zeyang Bai, Yunpeng Wang, Yunbiao Wang +1

    cs.CVarXiv:2608.22740v12026
  51. Beyond Self-attention: External Attention using Two Linear Layers for Visual Tasks

    Meng-Hao Guo, Zheng-Ning Liu, Tai-Jiang Mu +1

    cs.CVarXiv:2105.02358v22021
  52. Deep Learning in Multimodal Remote Sensing Data Fusion: A Comprehensive Review

    Jiaxin Li, Danfeng Hong, Lianru Gao +4

    cs.CVcs.LGeess.SParXiv:2205.01380v12022
  53. Routing the Lottery: Adaptive Subnetworks for Heterogeneous Data

    Grzegorz Stefanski, Alberto Presta, Michal Byra

    cs.AIcs.CVcs.LGarXiv:2601.22141v12026
  54. AI-Generated Image Detectors Overrely on Global Artifacts: Evidence from Inpainting Exchange

    Elif Nebioglu, Emirhan Bilgiç, Adrian Popescu

    cs.CVcs.AIarXiv:2602.00192v12026
  55. DSAC - Differentiable RANSAC for Camera Localization

    Eric Brachmann, Alexander Krull, Sebastian Nowozin +4

    cs.CVarXiv:1611.05705v42016
  56. Searching for A Robust Neural Architecture in Four GPU Hours

    Xuanyi Dong, Yi Yang

    cs.CVarXiv:1910.04465v22019
  57. Neural-Symbolic VQA: Disentangling Reasoning from Vision and Language Understanding

    Kexin Yi, Jiajun Wu, Chuang Gan +3

    cs.AIcs.CLcs.CVarXiv:1810.02338v22018
  58. Reliable and Responsible Foundation Models: A Comprehensive Survey

    Xinyu Yang, Junlin Han, Rishi Bommasani +49

    cs.LGcs.AIcs.CLarXiv:2602.08145v12026
  59. PlantDoc: A Dataset for Visual Plant Disease Detection

    Davinder Singh, Naman Jain, Pranjali Jain +3

    cs.CVeess.IVarXiv:1911.10317v12019
  60. Contextualized Visual Personalization in Vision-Language Models

    Yeongtak Oh, Sangwon Yu, Junsung Park +3

    cs.CVarXiv:2602.03454v32026