Source-linked AI summary

Small Object Detection: A Comprehensive Survey on Challenges, Techniques and Real-World Applications

Mahya Nikouei, Bita Baroutian, Shahabedin Nabavi, Fateme Taraghi, Atefe Aghaei, Ayoob Sajedi, Mohsen Ebrahimi Moghaddam

arXiv:2503.20516v1cs.CV

TL;DR

Small object detection matters across surveillance, autonomous systems, medical imaging, and remote sensing, yet limited resolution, occlusion, background interference, and imbalance hinder reliable detection. This survey synthesizes Q1 Scopus-indexed deep-learning research from 2024–2025 across challenges, methods, datasets, metrics, applications, and trends. It concludes that architecture, feature fusion, attention, lightweight models, transformers, and knowledge distillation are central directions for improving SOD accuracy and efficiency.

  • Problem

    Small objects contain limited spatial and contextual information, while low resolution, occlusion, background interference, and class imbalance make reliable detection difficult.

  • Method

    The survey reviews Q1-journal articles indexed in Scopus during 2024–2025 and analyzes SOD challenges, techniques, datasets, evaluation metrics, applications, and research gaps.

  • Results

    Optimized backbone architectures represented 23.1% of reviewed trends, followed by attention mechanisms at 18.5% and feature extraction enhancement at 16.9%.

  • Takeaways & Limitations

    Recent SOD research emphasizes stronger architectures, improved feature fusion, and attention to important image regions, alongside lightweight models, transformers, and knowledge distillation.

  • Takeaways & Limitations

    Popular detectors are poorly suited to SOD and can be computationally intensive, while deployment is constrained by limited edge-device resources and domain shift in transfer learning.

Abstract

from arXiv · show

Small object detection (SOD) is a critical yet challenging task in computer vision, with applications like spanning surveillance, autonomous systems, medical imaging, and remote sensing. Unlike larger objects, small objects contain limited spatial and contextual information, making accurate detection difficult. Challenges such as low resolution, occlusion, background interference, and class imbalance further complicate the problem. This survey provides a comprehensive review of recent advancements in SOD using deep learning, focusing on articles published in Q1 journals during 2024-2025. We analyzed challenges, state-of-the-art techniques, datasets, evaluation metrics, and real-world applications. Recent advancements in deep learning have introduced innovative solutions, including multi-scale feature extraction, Super-Resolution (SR) techniques, attention mechanisms, and transformer-based architectures. Additionally, improvements in data augmentation, synthetic data generation, and transfer learning have addressed data scarcity and domain adaptation issues. Furthermore, emerging trends such as lightweight neural networks, knowledge distillation (KD), and self-supervised learning offer promising directions for improving detection efficiency, particularly in resource-constrained environments like Unmanned Aerial Vehicles (UAV)-based surveillance and edge computing. We also review widely used datasets, along with standard evaluation metrics such as mean Average Precision (mAP) and size-specific AP scores. The survey highlights real-world applications, including traffic monitoring, maritime surveillance, industrial defect detection, and precision agriculture. Finally, we discuss open research challenges and future directions, emphasizing the need for robust domain adaptation techniques, better feature fusion strategies, and real-time performance optimization.

1- Introduction

Small object detection is important across safety-critical and analytical applications but remains difficult because small objects provide limited information and are easily obscured. This survey reviews recent deep-learning advances and the practical constraints shaping SOD research.

  • Applications: SOD supports surveillance, medical imaging, autonomous vehicles, and remote sensing, where detecting small objects can affect safety, diagnosis, navigation, and disaster response.Applications include unattended bags, tumors, road signs, cyclists, animals, structures, and vehicles.
  • Core Challenges: Limited resolution, weak feature representation, background interference, occlusion, and class imbalance make small objects difficult to distinguish and detect reliably.Small objects may lack distinctive visual features and are vulnerable to lighting changes, motion blur, dust, and low illumination.
  • Survey Scope: The survey synthesizes Q1 Scopus-indexed studies from 2024–2025, covering SOD challenges, techniques, datasets, metrics, applications, and research gaps.It specifically examines recent deep-learning developments and compares their focus with limitations identified in earlier reviews.
  • Localization and Scale: Accurate localization is difficult because small objects have many possible positions, require high precision, and often mismatch anchor boxes or receptive fields across scales.These issues can produce imprecise bounding boxes and reduce detector reliability.
  • Feature Representation: Downsampling removes discriminative textures and edges, while FPNs trade shallow-layer localization precision against deep-layer semantic information.This shallow–deep imbalance contributes to suboptimal detection, especially in cluttered environments.
  • Model Limitations: Existing CNN-based detectors often preserve insufficient fine-grained detail and impose high computational demands, particularly on high-resolution inputs.Adapting them for SOD may also require costly architectural, training, or loss-function modifications.

4-1-1- Model Optimization and Lightweight Architectures

Recent SOD research prioritizes lightweight architectures and combined feature-processing strategies to balance detection accuracy with computational efficiency. The surveyed methods use approaches including motion processing, feature fusion, attention, deformable convolution, and knowledge distillation across diverse deployment settings.

  • Lightweight Architectures: Lightweight designs such as FFEDet and KDSMALL target scalable SOD for drones, mobile devices, and edge computing.These approaches use modular design, parameter sharing, and compression to reduce complexity while retaining detection capability.
  • Feature Fusion Optimization: Multi-scale feature extraction and fusion improve contextual understanding and precision across object scales.Attention mechanisms, transformer-based detectors, large-kernel designs, and multi-frame or cross-modal fusion are among the surveyed strategies.
  • Neural Network Architecture: Anchor-free and transformer-based architectures simplify prediction or capture long-range dependencies and multi-scale features for SOD.Hybrid CNN-transformer models combine the advantages of convolutional and transformer frameworks, while DETR-like methods can address NMS limitations.
  • Advanced Learning Strategies: Knowledge distillation improves SOD efficiency and real-time performance, while self-supervised learning reduces reliance on labeled data.These strategies are particularly relevant to unsupervised and semi-supervised training scenarios and lightweight UAV deployment.
  • Advanced Learning Strategies: Recent methods combine multiple components, including motion processing, deformable convolution, reinforcement learning, feature fusion, attention, and KD, to improve accuracy and reduce computation.Examples include MICPL, Bi-AFPN-P2 with DD-Head, and ScorePillar, which target spatial detail, multi-scale fusion, localization, or real-time LiDAR detection.

4-3- Clarity and Visual Information Improvement

Improving image clarity and visual information is presented as a central strategy for extracting fine-grained details, especially when detecting small or low-resolution objects. The section focuses on techniques that support object detection, segmentation, and image classification.

  • Clarity and Visual Information Improvement: Enhancing image clarity helps models extract fine-grained details needed for small or low-resolution object detection.The section frames visual-information improvement as relevant across object detection, segmentation, and image classification.

4-3-1- SR for Image Quality Enhancement

Super-resolution improves low-resolution imagery by recovering or generating finer details, supporting more accurate small-object localization and classification. The survey also notes that SR may be combined with denoising and deblurring to limit artifacts.

  • SR for Image Quality Enhancement: Super-resolution upscales low-resolution inputs into higher-resolution outputs using contextual information to recover fine details.Deep learning models such as CNNs learn this transformation from low-resolution imagery.
  • SR for Image Quality Enhancement: In SOD, SR helps models discern minute details and improves the localization and classification of tiny objects.GAN-based SR methods are also used to generate realistic high-resolution images.
  • SR for Image Quality Enhancement: SR is often coupled with denoising and deblurring so super-resolved images do not introduce artifacts that degrade model performance.The combined enhancement pipeline is intended to preserve the usefulness of the reconstructed imagery.

4-3-2- Utilizing Multi-Scale Information

Multi-scale processing captures both fine-grained details and broader context by extracting features at different resolutions or abstraction levels. Feature pyramids and related fusion structures address the challenge of combining these representations while balancing spatial and semantic information.

  • Utilizing Multi-Scale Information: Multi-scale information processes features at various scales to capture fine details and broader contextual information.This is useful because objects can appear at different sizes and require multiple abstraction levels for detection.
  • Utilizing Multi-Scale Information: Multi-resolution networks and feature pyramids merge multi-level feature maps to recognize small and large objects across scales.Early-layer fine-grained details can be combined with more abstract features from deeper layers.
  • Utilizing Multi-Scale Information: A key design challenge is balancing fine-grained spatial details with semantic information during multi-scale feature fusion.FPNs address this by combining high-resolution features with coarser contextual information from deeper layers.

4-3-3- Fusion of Information Across Network Layers

Fusion across network layers combines low-level spatial details with high-level semantic information, supporting small-object detection when objects are represented mainly in early layers. Attention and Feature Fusion Networks provide strategies for weighting or integrating these complementary features.

  • Layer Information Fusion: Layer fusion combines low-level edges and textures with high-level semantic information for a more comprehensive image representation.This is especially relevant when small objects are represented primarily in early network layers.
  • Fusion Strategies: Concatenation, addition, and attention mechanisms are strategies for fusing information from different network layers.Attention mechanisms dynamically weight important layers or image regions.
  • Feature Fusion Networks: Feature Fusion Networks combine outputs from different layers to capture semantic and spatial information for detecting poorly represented small objects.The resulting representation is intended to improve clarity and detection effectiveness.

4-4-1- Data Augmentation Techniques for SOD

Data augmentation expands SOD training data through transformations that simulate changes in scale, orientation, context, image quality, and occlusion. Advanced methods combine multiple images to create cluttered scenarios that improve discrimination of small objects.

  • Augmentation Overview: Data augmentation applies transformations to existing images so SOD models focus on small-scale features and simulated real-world conditions.The approach artificially expands the training dataset.
  • Geometric Transformations: Rotation and scaling expose models to varying orientations and sizes of small objects.These transformations are intended to improve generalization across spatial representations.
  • Geometric Transformations: Cropping and zooming keep small objects in frame while encouraging learning of fine-grained details.The transformations focus training on small-object appearance.
  • Contextual Transformations: Padding and contextual cropping add surrounding information and emphasize objects together with their immediate context.These operations are intended to improve contextual awareness.
  • Image-Quality Transformations: Brightness, contrast, and noise adjustments simulate lighting and image-quality variation in real-world conditions.The goal is to help models handle degraded or variable visual inputs.
  • Occlusion Simulation: Random erasing and occlusion simulation train models to detect small objects when portions of them are obscured.Partial erasure is used to simulate occlusion.
  • Advanced Augmentation: MixUp, CutMix, and Mosaic combine multiple images into composites containing complex backgrounds and small-object configurations.These methods are described as particularly impactful for SOD.

4-4-2- Synthetic Data Generation for SOD

Synthetic data generation creates new SOD samples when real-image collection is limited, while multi-task and transfer learning provide complementary feature and data-efficiency strategies. These methods support detection across applications but remain constrained by domain gaps, task conflicts, overfitting, and computational demands.

  • Synthetic Data Generation: Synthetic data generation creates new training samples when collecting real small-object images is impractical or expensive.It complements augmentation, which transforms existing data.
  • Synthetic Data Methods: GANs, CGI, and simulators generate synthetic small-object scenes with controllable appearances, positions, environments, lighting, and occlusions.Simulation platforms include Unity, CARLA, and Blender.
  • Synthetic Data Benefits: Synthetic datasets provide pixel-perfect bounding-box, segmentation-mask, and keypoint annotations while reducing manual labeling cost and time.The annotations support training detection models.
  • Synthetic Data Benefits: Augmented and synthetic data can improve generalization and robustness while addressing data scarcity, cost, and class imbalance.Underrepresented small-object classes can be oversampled through tailored generation and augmentation.
  • Challenges: Synthetic-data pipelines face domain gaps, over-augmentation, and scalability constraints requiring adaptation, careful tuning, or substantial resources.Visual differences from real images can reduce performance on real-world data.
  • Future Directions: Future directions include automated augmentation, diffusion-based realism improvements, and multimodal synthetic data for contextual learning.These directions are proposed for SOD research.
  • Multi-Task Learning: Multi-task learning shares representations across related tasks, using auxiliary objectives such as SR, edge detection, and depth estimation to support fine-grained SOD features.Task-specific heads can preserve distinct objectives while sharing a backbone.
  • Multi-Task Learning: Multi-task learning can enrich representations and reduce training resources, but task conflicts and architectural complexity require balancing and careful optimization.Dynamic loss weighting is identified as one task-balancing strategy.

5- Datasets and Evaluation Metrics

The survey organizes SOD datasets around challenges including size variation, occlusion, complex backgrounds, and real-world environmental conditions. The reviewed resources support surveillance, autonomous driving, aerial detection, tracking, and related applications.

  • Dataset Scope: SOD datasets address size variation, occlusion, and complex backgrounds across surveillance, autonomous driving, and real-time tracking tasks.Many are captured from UAVs and satellites to represent real-world detection conditions.
  • Dataset Summary: Table 2 summarizes datasets used for SOD.The supplied caption identifies the table as a summary of SOD datasets.
  • Aerial SOD Datasets: AI-TOD contains more than 2.5 million object instances and targets tiny objects affected by low resolution, scale variation, occlusion, and complex backgrounds.The dataset is intended to develop and evaluate small-object detection algorithms.
  • UAV and Aerial Datasets: UAV and aerial datasets cover small objects, crowd scenes, object detection, and multi-object tracking under occlusion and scale variation.Examples include SODA and other UAV-focused resources.
  • Driving Datasets: BDD100K contains 100,000 camera images covering driving environments, object categories, weather, lighting, and urban or rural conditions.The dataset supports autonomous-driving and real-world object-detection research.
  • Maritime Datasets: Water-surface datasets support detection of floating objects such as boats, logs, debris, and hazards under reflections and changing environmental conditions.Applications include waterway monitoring and environmental protection using UAVs.
  • Driving Datasets: KITTI supports autonomous-driving research through 3D annotations and detection, stereo-vision, and tracking scenarios.Its categories include cars, pedestrians, and cyclists in urban and related environments.

5-1-1- Aerial Image Datasets

Aerial imagery datasets are central SOD benchmarks because they combine varied object scales, occlusion, dense clustering, and complex capture conditions. Key resources differ in object coverage, scale, annotation format, and intended detection setting.

  • Aerial imagery datasets support SOD through challenging variation in object sizes, occlusion, and dense spatial clustering.
  • VisDrone: VisDrone contains 10 categories across over 10,209 images, with training, validation, and testing subsets totaling 6,471, 548, and 3,190 images.Its scenes include varied spatial density, occlusion, weather, lighting, and urban or rural backgrounds.
  • DIOR: DIOR provides 23,463 images and 192,472 annotated instances across 20 categories collected from multiple sensors and conditions.This variety supports evaluation of detection robustness in remote-sensing imagery.
  • DOTA: DOTA offers multiple versions with over 15 categories, including small vehicles, ships, and storage tanks annotated using rotated bounding boxes.DOTA-v1.0, v1.5, and v2.0 incrementally improve sample size and annotations.
  • VEDAI: VEDAI targets multi-class small-vehicle detection under varying resolutions and cluttered environments.

5-1-2- UAV-Based Datasets

UAV-based and related SOD datasets capture small objects under changing viewpoints, motion, weather, lighting, clutter, and occlusion. They span vehicle, pedestrian, maritime, autonomous-driving, and synthetic-data scenarios, while evaluation uses both general and size-aware metrics.

  • UAV imagery supports SOD datasets by providing views from varied altitudes and angles.
  • UAVDT: UAVDT contains over 100 annotated videos of vehicles across highways and parking lots, including motion blur, small objects, and changing weather.
  • SODA-D: SODA-D covers nine object classes with large, medium, small, and very small size labels under diverse weather and lighting conditions.
  • Crowd tracking: Crowd-tracking data target small human objects in dense urban scenes with substantial occlusion.
  • Autonomous driving: Autonomous-driving datasets include small pedestrians, road signs, bicycles, vehicles, and traffic lights across varied driving conditions.Berkeley DeepDrive contains 100,000 driving scenes, while KITTI emphasizes 3D detection and includes small pedestrians and bicycles.
  • TinyPerson: TinyPerson focuses on high-resolution annotations for small pedestrians, supporting surveillance and crowd analysis.
  • Maritime and tiny-object datasets: WSODD covers boats and buoys that may be occluded by waves, while AI-TOD addresses very small aerial objects such as vehicles and ships.
  • Synthetic datasets: Synthetic datasets address labeled-data scarcity by simulating realistic conditions, including motion blur and controlled industrial scenes.Motion-blur variation is intended to improve generalization to blurred real-world images.

5-2-2- Metrics for Size-Specific Detection

SOD evaluation extends general detection metrics with measures tailored to object size, tracking continuity, computational efficiency, image quality, and application-specific datasets. These metrics support comparisons across diverse real-world settings.

  • Size-specific metrics: Size-specific metrics evaluate detection separately for small, medium, and large objects, addressing performance differences across scales.COCO-style measures include APS_SS, APM_MM, and APL_LL.
  • Tiny-object metrics: APT_TT measures average precision specifically for tiny objects, which are smaller than COCO’s small-object category.
  • Tracking metrics: Tracking metrics T-AP10, T-AP15, and T-AP20 evaluate accuracy at 10%, 15%, and 20% IoU thresholds.T-mAP additionally accounts for temporal consistency across frames.
  • Foundational metrics: Precision, recall, and F1-score measure false-positive reduction, relevant-instance identification, and the balance between these objectives.
  • Efficiency and quality: Specialized measures capture computational efficiency, detection quality, and real-time capability in challenging SOD applications.Examples include PPN, SNR, CV, FPS, and inference time.
  • Dataset-specific evaluation: COCO, PASCAL VOC, VisDrone, TinyPerson, DIOR, and DOTA use combinations of AP, mAP, IoU-specific, and size-specific metrics suited to their object distributions and scenes.
  • Custom metrics: SODA-D and WSODD may use custom metrics such as mean Average Recall and image-quality measures for motion blur, varying object sizes, or maritime conditions.
  • Applications: SOD applications span remote sensing, UAV surveillance, autonomous navigation, industrial inspection, security, environmental management, healthcare, and marine monitoring.
Loading 2503.20516v1…