Source-linked AI summary

Segment Anything Is Not Always Perfect: An Investigation of SAM on Different Real-world Applications

Wei Ji, Jingjing Li, Qi Bi, Tingwei Liu, Wenbo Li, Li Cheng

arXiv:2304.05750v3cs.CV

TL;DR

SAM shows strong generalization in common natural-image scenes, but dedicated pre-training data may not cover the unusual conditions, modalities, and applications found in real-world segmentation. This study evaluates SAM1 across diverse applications and benchmarks, finding both effective common-scene performance and substantial gaps in challenging scenarios, while identifying directions for more robust adaptation.

  • Problem

    A dedicated pre-training dataset may not encompass the diverse conditions, input modalities, and real-world applications encountered in segmentation.

  • Method

    The study investigates SAM1 across natural images, agriculture, manufacturing, remote sensing, and healthcare using eight representative benchmarks, while analyzing practical benefits and limitations.

  • Results

    SAM generalizes well to typical natural-image scenes with distinct target regions, but shows a considerable performance gap against top-performing models across eight benchmarks, especially in challenging scenarios.

  • Takeaways & Limitations

    The observations can guide development of more robust SAM algorithms and benchmarks, including learning strategies that adapt SAM to broader scenarios.

  • Takeaways & Limitations

    SAM requires more manual prompts and prior knowledge in complex scenes, and its foreground bias can hinder performance on tasks such as shadow detection.

Abstract

from arXiv · show

Recently, Meta AI Research approaches a general, promptable Segment Anything Model (SAM) pre-trained on an unprecedentedly large segmentation dataset (SA-1B). Without a doubt, the emergence of SAM will yield significant benefits for a wide array of practical image segmentation applications. In this study, we conduct a series of intriguing investigations into the performance of SAM across various applications, particularly in the fields of natural images, agriculture, manufacturing, remote sensing, and healthcare. We analyze and discuss the benefits and limitations of SAM, while also presenting an outlook on its future development in segmentation tasks. By doing so, we aim to give a comprehensive understanding of SAM's practical applications. This work is expected to provide insights that facilitate future research activities toward generic segmentation. Source code is publicly available.

1. Introduction

The paper investigates SAM across diverse real-world segmentation applications, finding strong generalization in common scenes but important limitations in low-contrast, professional, and small or irregular-object settings.

  • Study scope: SAM is evaluated across natural images, agriculture, manufacturing, remote sensing, and healthcare to examine its practical generalization and limitations.The study presents observations intended to guide future development of foundational vision models and robust segmentation algorithms.
  • Key observations: SAM demonstrates excellent generalization across common scenes and prompt modes, especially when target regions are clearly distinct from their surroundings.The authors attribute this observation to SAM’s promptable design and massive, diverse training data.
  • Practical constraints: Complex scenes may require many manual prompts with prior knowledge, while excessive interaction effort makes optimal prompting impractical.The paper also notes that prompt optimality cannot be guaranteed.
  • Key observations: SAM is less effective in low-contrast applications involving transparent or camouflaged objects embedded seamlessly in their surroundings.The authors identify robustness in complex low-contrast scenes as an area requiring further enhancement.
  • Key observations: SAM produces unsatisfactory results on professional medical and industrial data, particularly with box and everything modes.Click mode can still require domain-specific knowledge from both the user and the model.
  • Key observations: Small and irregular objects in remote sensing and agriculture can prevent SAM from producing complete segmentations.Designing effective strategies for these cases remains an open issue.

2. Qualitative Investigation

The qualitative investigation presents SAM’s visual results across diverse segmentation tasks and uses prompt modes to examine its practical behavior.

  • Qualitative investigation: The section analyzes SAM qualitatively across diverse segmentation tasks, focusing on observed advantages and limitations.The visual results are organized around practical applications rather than a single task.
  • Salient object segmentation: Figure 2 compares salient object segmentation results across Click, Box, and Everything prompt modes.SAM1, SAM2, and SAM3 denote Click, Box, and Everything modes, respectively.
  • Camouflaged object segmentation: Figure 3 compares camouflaged object segmentation results across Click, Box, and Everything prompt modes.The same SAM1, SAM2, and SAM3 notation is used for the three prompt modes.

2.1. Natural Image Scenes

SAM is effective at locating prominent objects in natural images but struggles with camouflage, transparency, shadows, clutter, and fine-grained boundaries.

  • Natural image subtasks: SAM accurately identifies salient object locations but captures fewer fine-grained details in richly detailed targets.Its performance is strong for prominent objects, while detailed boundary delineation remains weaker.
  • Natural image subtasks: SAM often fails to detect complete camouflaged objects when foreground and background are similar or scenes contain multiple cluttered objects.Increasing human interactions still leaves many false positives.
  • Natural image subtasks: SAM locates transparent objects effectively but still needs improvement in capturing their finer details.Transparent objects may have complex shapes and unpredictable refraction or reflection.
  • Natural image subtasks: SAM fails to detect shadows, illustrating the difficulty of segmenting targets that are hard to distinguish from their surroundings.The figure compares click, box, and everything prompt modes, with starred results denoting box-prompt outputs.
  • Section recap: Across natural-image tasks, SAM is strongest at finding object locations and weaker in noisy, low-contrast, or detail-demanding scenes.These limitations leave substantial opportunity for further exploration and improvement.

2.2. Agriculture

SAM shows appealing but imperfect performance in agriculture, with strong results in some pest and disease scenes but incomplete segmentation in challenging crop and pest cases.

  • Agricultural subtasks: SAM achieves relatively satisfactory crop segmentations in some agricultural scenes but performs worse in others.The authors associate this discrepancy with limited positive examples for crops in general segmentation training.
  • Agricultural subtasks: SAM exhibits excellent generalization for some pest and leaf disease monitoring scenes but struggles to detect entire pest bodies in harder scenes.The agricultural results are appealing overall, though not perfect.
  • Section recap: The authors suggest that sufficient agricultural prior knowledge could help SAM detect challenging agricultural objects more effectively.This is presented as a future expectation rather than an experimentally established result.

2.3. Manufacture

SAM performs acceptably in industrial scenarios, but practical manufacturing deployment depends substantially on expert prompts and domain-specific prior knowledge.

  • Manufacturing subtasks: SAM demonstrates powerful recognition ability on anomaly detection in the MVTec AD industrial dataset.The passage also introduces surface defect detection for products such as wood, textiles, and mineral materials.
  • Manufacturing subtasks: Human experts’ prior knowledge is crucial for successful AI model applications in real-world industrial settings.This requirement is especially relevant to practical surface-defect detection.
  • Section recap: SAM performs acceptably in industrial scenarios, while surface defect detection requires a significant amount of expert prior knowledge for deployment.The evidence supports acceptable performance alongside a substantial expertise requirement.

2.4. Remote Sensing

In remote sensing, SAM scales well to regular buildings and roads but faces difficulties with smaller, indistinguishable, and highly variable targets.

  • Remote-sensing subtasks: Remote-sensing targets vary substantially in shape and size, making accurate segmentation challenging.The evaluation focuses on extracting essential buildings and roads from aerial imagery.
  • Remote-sensing subtasks: SAM is proficient at segmenting regularly shaped remote-sensing objects and has good scalability for regular buildings and roads.This strength is limited when targets become smaller or less distinguishable.
  • Section recap: Smaller or less distinguishable remote-sensing objects may require task-specific properties and adaptation to achieve effective segmentation.Diverse shapes, sizes, and textures further challenge accurate segmentation.

2.5. Healthcare

In healthcare applications, SAM can produce satisfactory segmentations with sufficient human prompts, but box and automatic modes often underperform, especially when expert knowledge is needed.

  • Healthcare applications: Joint optical disc and cup segmentation separates the optic disc and optic cup to compute the clinically important cup-to-disc ratio for glaucoma screening.A higher cup-to-disc ratio indicates a larger optic cup and is often associated with higher glaucoma risk.
  • Healthcare applications: SAM shows severe shortcomings on joint optical disc and cup segmentation in retinal fundus images.The evaluation uses the RIGA benchmark.
  • Healthcare applications: SAM delivers satisfactory skin-lesion segmentation when provided with sufficient human prompts, but substantial room for improvement remains.Skin-lesion segmentation is difficult because of indistinct boundaries, variable contrast, and color differences.
  • Healthcare applications: SAM1 can produce appealing medical results, but effective prompts for tasks such as optic-disc and optic-cup segmentation often require substantial human prior knowledge.This requirement makes it difficult to obtain satisfactory results directly for some medical tasks.
  • Healthcare applications: In medical applications, box and automatic modes fall short, motivating dedicated SAM models tailored to healthcare scenarios.The paper contrasts these modes with SAM1, which can deliver appealing results when appropriately prompted.

3. Quantitative Investigation

The quantitative investigation evaluates SAM across eight benchmarks using mean absolute error and selects the candidate mask with the highest IoU against ground truth. Across these benchmarks, SAM remains substantially behind top-performing dedicated models, particularly on challenging scenarios.

  • 3.1. Datasets and Evaluation Metric: Eight benchmarks cover common, low-contrast, low-light, detailed-boundary, camouflaged-object, shadow, industrial-defect, and medical-polyp segmentation.The benchmarks include DUTS, COME15K-Diff, VT1000, DIS-TE4, COD10K, SBU, CDS2K, and ColonDB.
  • 3.1. Datasets and Evaluation Metric: Mean absolute error (M) is used for quantitative evaluation, with lower M indicating better model performance.
  • 3.2. Numerical Results: SAM2 generates N potential object masks, computes each mask’s IoU with the ground truth, and selects the mask with the highest IoU.This selection procedure chooses the candidate most aligned with the ground-truth mask for quantitative comparison.
  • 3.2. Numerical Results: 0.265 MAE is achieved by SAM with ViT-H on the industrial application in Table 1(g), 17.6% worse than DGNet.
  • 3.2. Numerical Results: The quantitative results show a considerable performance gap between SAM and the top-performing model across eight benchmarks, especially for challenging scenarios and applications.

4. Discussion and Outlook

The discussion proposes application-oriented models and datasets, richer prompt modes, revised pretraining, video SAM, and semi-supervised uses as directions for extending SAM beyond its current scope.

  • 4. Discussion and Outlook: Application-oriented SAM models and dedicated large-scale datasets are proposed because SAM performs unappealingly in healthcare, manufacturing, and remote sensing despite strong natural-image performance.
  • 4. Discussion and Outlook: Additional prompt modes, including voice and gesture, are identified as directions beyond click, box, and everything modes.
  • 4. Discussion and Outlook: The paper suggests using SA-1B as a pretraining resource for industrial and medical applications, potentially with metric learning to improve feature adaptability.
  • 4. Discussion and Outlook: Developing video SAM from an initial prompt in the first frame and leveraging SAM-generated pseudo-labels for semi-supervised segmentation remain open directions.

5. Conclusion

The paper presents a preliminary cross-application investigation of SAM, analyzing its benefits, limitations, challenges, and future directions. The authors intend to study SAM more deeply and develop learning strategies for broader scenarios.

  • 5. Conclusion: The study investigates SAM across natural images, agriculture, manufacturing, remote sensing, and healthcare.
  • 5. Conclusion: The analysis identifies SAM’s benefits, limitations, and potential challenges while suggesting future research directions.
  • 5. Conclusion: The authors plan deeper study and learning strategies to adapt SAM to a broader range of scenarios.
Loading 2304.05750v3…