Source-linked AI summary
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
Tianhe Ren, Qing Jiang, Shilong Liu, Zhaoyang Zeng, Wenlong Liu, Han Gao, Hongjie Huang, Zhengyu Ma, Xiaoke Jiang, Yihao Chen, Yuda Xiong, Hao Zhang, Feng Li, Peijun Tang, Kent Yu, Lei Zhang
TL;DR
Open-set object detection needs models that generalize across diverse scenarios while also meeting edge-device speed and resource constraints. Grounding DINO 1.5 develops Pro and Edge variants using scaled modeling, extensive grounding data, and efficient feature processing. Pro reaches 54.3 AP on COCO and 55.7 AP on LVIS-minival, while Edge reaches 75.2 FPS and 36.2 AP on LVIS-minival under TensorRT optimization.
Problem
Open-set detection must support diverse object categories and applications, while deploying such models on edge devices is constrained by their computational cost and limited hardware resources.
Method
Grounding DINO 1.5 introduces Pro with a larger backbone and Grounding-20M training data, and Edge with high-level feature fusion for efficient inference.
Results
Grounding DINO 1.5 Pro achieves 54.3 AP on COCO detection zero-shot transfer and 55.7 AP on LVIS-minival zero-shot transfer, while Edge reaches 75.2 FPS and 36.2 AP on LVIS-minival under TensorRT optimization.
Takeaways & Limitations
The Pro and Edge models extend Grounding DINO toward stronger zero-shot detection and practical real-time edge deployment.
Abstract
from arXiv · showhide
This paper introduces Grounding DINO 1.5, a suite of advanced open-set object detection models developed by IDEA Research, which aims to advance the "Edge" of open-set object detection. The suite encompasses two models: Grounding DINO 1.5 Pro, a high-performance model designed for stronger generalization capability across a wide range of scenarios, and Grounding DINO 1.5 Edge, an efficient model optimized for faster speed demanded in many applications requiring edge deployment. The Grounding DINO 1.5 Pro model advances its predecessor by scaling up the model architecture, integrating an enhanced vision backbone, and expanding the training dataset to over 20 million images with grounding annotations, thereby achieving a richer semantic understanding. The Grounding DINO 1.5 Edge model, while designed for efficiency with reduced feature scales, maintains robust detection capabilities by being trained on the same comprehensive dataset. Empirical results demonstrate the effectiveness of Grounding DINO 1.5, with the Grounding DINO 1.5 Pro model attaining a 54.3 AP on the COCO detection benchmark and a 55.7 AP on the LVIS-minival zero-shot transfer benchmark, setting new records for open-set object detection. Furthermore, the Grounding DINO 1.5 Edge model, when optimized with TensorRT, achieves a speed of 75.2 FPS while attaining a zero-shot performance of 36.2 AP on the LVIS-minival benchmark, making it more suitable for edge computing scenarios. Model examples and demos with API will be released at https://github.com/IDEA-Research/Grounding-DINO-1.5-API
1 Introduction
Grounding DINO 1.5 extends open-set object detection with separate Pro and Edge models targeting stronger detection and faster inference. Pro scales model capacity and training data, while Edge reduces computational demands while retaining detection capabilities.
- Grounding DINO 1.5 introduces Pro and Edge models for stronger detection performance and faster inference speed, respectively.
- Pro incorporates a pre-trained ViT-L architecture and over 20 million grounding-annotated images from diverse sources.The expanded model capacity and data are intended to enrich semantic comprehension.
- Edge uses only high-level image features and trains on the same 20 million images as Pro to reduce computation while preserving context-aware detection.
- 54.3 AP on COCO detection zero-shot transfer and 55.7 AP on LVIS-minival zero-shot transfer are reported for Grounding DINO 1.5 Pro.The paper reports these results as new records on the respective benchmarks.
- Under TensorRT optimization, Grounding DINO 1.5 Edge reaches 75.2 FPS and 36.2 AP on LVIS-minival zero-shot transfer.These results support its suitability for edge computing scenarios.
2 Model Training
Grounding DINO 1.5 retains the dual-encoder-single-decoder framework while adapting architecture, feature enhancement, and training data for scalable open-set detection. Its Edge variant targets resource-constrained deployment by limiting multimodal fusion to high-level features and using an efficient enhancer.
- Model Architecture: The Grounding DINO 1.5 framework retains Grounding DINO’s dual-encoder-single-decoder structure for both Pro and Edge models.
- Grounding DINO 1.5 Pro: Pro uses a larger ViT-L vision backbone and deep early fusion between language and image features before decoding.The early-fusion design integrates the two modalities during feature extraction.
- Grounding DINO 1.5 Pro: Early fusion improves detection recall and bounding-box precision but can increase hallucinated predictions of absent objects.
- Grounding DINO 1.5 Pro: A more comprehensive sampling strategy increases negative samples during training to balance early fusion’s prediction benefits and robustness drawbacks.
- Grounding DINO 1.5 Edge: Edge addresses the computational gap between open-set detection models and limited edge-device resources by replacing multi-scale fusion with high-level-feature fusion.
- Training Dataset: Grounding DINO 1.5 is pre-trained on Grounding-20M, a dataset of over 20M publicly sourced grounding images processed through annotation and post-processing pipelines.
3 Model Evaluation
Grounding DINO 1.5 is evaluated across zero-shot transfer and fine-tuning settings on COCO, LVIS, and ODinW benchmarks. Grounding DINO 1.5 Pro reports strong results across these evaluations, while Grounding DINO 1.5 Edge combines zero-shot detection with edge-oriented speed.
- Zero-shot transfer: 54.3 AP on COCO zero-shot transfer improves upon Grounding DINO Swin-L by 1.8 AP.
- Zero-shot transfer: 55.7 AP on LVIS-minival and 47.6 AP on LVIS-val surpass DetCLIPv3 by 6.9 AP and 6.2 AP, respectively.
- Fine-tuning: 68.1 AP on LVIS-minival and 63.5 AP on LVIS-val after fine-tuning improve over the zero-shot setting by 12.4 AP and 15.9 AP, respectively.
- Fine-tuning: Fine-tuning on ODinW35 yields 70.6 AP across 35 datasets and 72.4 AP across 13 datasets on ODinW13.
- Edge model: 36.2 AP on LVIS-minival and 75.2 FPS under TensorRT characterize Grounding DINO 1.5 Edge’s zero-shot accuracy and optimized speed.
4 Case Analysis and Qualitative Visualization
Grounding DINO 1.5 Pro is visualized across common and long-tailed object scenarios, including challenging conditions and uncommon categories.
- Common Object Detection: The visualizations show robust performance on common objects, including challenging detection conditions.Examples cover monochromatic images, blurry objects, small objects, partial occlusion, and varied object sizes and shapes.
- Common Object Detection: The model detects objects in monochromatic images where color cues are minimal.The example indicates reliance on shape and texture for distinguishing objects.
- Common Object Detection: The model detects small and partially occluded objects in autonomous driving scenes.The visualizations also show accurate localization across objects of varying sizes and shapes.
- Long-tailed Object Detection: Grounding DINO 1.5 Pro recognizes diverse long-tailed categories that are uncommon in everyday settings.The examples include fungus, popsicle, and taco detection, illustrating specialized and contextual recognition.
4.3 Short Caption Grounding
Grounding DINO 1.5 Pro grounds textual phrases to visual objects across real and stylized domains, including captions longer than those used during pre-training.
- Short Caption Grounding: The model grounds objects across real-world images, cartoons, sketches, and animations.It aligns textual descriptions with visual features to identify and localize objects in stylized or abstract forms.
- Long Caption Grounding: Grounding DINO 1.5 Pro maps noun phrases in long captions to corresponding image objects.The paper presents this capability as extending beyond standard image-caption pairs toward deeper image understanding.
- Long Caption Grounding: The model processes longer caption contexts despite pre-training on shorter context windows.The paper describes this as generalization across different textual input lengths.
- Long Caption Grounding: The model detects terms absent from its training data, such as “fiat logo.”The paper presents this observation as evidence of strong generalization ability.
4.5 Dense Object Detection
Grounding DINO 1.5 Pro is presented as capable of detecting densely packed objects and recognizing both common and specialized object names, with consistent video boxes in offline processing.
- Dense Object Detection: The model detects objects in dense scenes where objects are closely positioned or overlapping.The paper illustrates this capability through visualizations in Figures 10 and 11.
- Dense Object Detection: Grounding DINO 1.5 Pro recognizes objects labeled with common names and specialized terminology.Examples include coin, tree, flower, land, kohlrabi, atlantic puffin, and oxalis purpurea.
- Dense Object Detection: The model produces consistent object bounding boxes in most offline-processed videos.The paper presents these results in Figure 12.
4.7 Side-by-side Comparison
Side-by-side visual comparisons report advantages for Grounding DINO 1.5 Pro over Grounding DINO in dense scenes, long-tailed detection, semantic understanding, and hallucination control.
- Side-by-side Comparison: Grounding DINO 1.5 Pro demonstrates superior performance over Grounding DINO in dense scene and long-tailed object detection.The comparison also reports higher accuracy of semantic understanding.
- Side-by-side Comparison: Grounding DINO 1.5 Pro has better accuracy and fewer object hallucinations than Grounding DINO.The comparison also highlights better context understanding in the final example.
4.8 Advanced Object Detection on Edge Devices
Grounding DINO 1.5 Edge is demonstrated in real-time applications on edge devices, while Grounding DINO 1.5 Pro is visualized across dense scenes, video detection, long-caption phrase grounding, and comparisons with Grounding DINO.
- Grounding DINO 1.5 Edge: Grounding DINO 1.5 Edge demonstrates practical, real-time application potential, particularly in office environments and robotics research.Figure 16 visualizes the model on NVIDIA Orin NX, displaying FPS, prompts, and a camera view of the recorded scene.
- Grounding DINO 1.5 Pro: Grounding DINO 1.5 Pro is visualized for phrase grounding on long captions.Figures 7–9 cover three parts of the long-caption phrase-grounding demonstrations.
- Grounding DINO 1.5 Pro: Grounding DINO 1.5 Pro is visualized on dense object scenarios and video object detection.Figures 10–12 present model predictions for these settings.
- Comparisons: Side-by-side visualizations compare Grounding DINO 1.5 Pro with Grounding DINO, including comparisons concerning object hallucinations.Figures 13–15 present the comparisons in multiple parts.
5 Conclusion
The paper concludes that Grounding DINO 1.5 advances open-set object detection through the Grounding DINO 1.5 Pro and Grounding DINO 1.5 Edge models. Pro establishes new COCO and LVIS zero-shot benchmark records, while Edge supports real-time detection across applications.
- Grounding DINO 1.5 comprises models intended to advance open-set object detection.
- Grounding DINO 1.5 Pro establishes new records on COCO and LVIS zero-shot benchmarks for detection accuracy and reliability.
- Grounding DINO 1.5 Edge enables real-time object detection across various applications.
6 Contributions and Acknowledgments
The acknowledgments recognize contributors to model design, infrastructure, data collection, training, evaluation, optimization, annotation-pipeline construction, and demo support.
- Technical contributions: Contributors worked on Grounding DINO 1.5 Pro model design, training infrastructure, data collection, model training, and model evaluation.
- Technical contributions: Contributors worked on Grounding DINO 1.5 Edge model design, training, evaluation, and runtime optimization.
- Technical contributions: Contributors constructed the Grounding-20M data collection and annotation pipeline.
- Demo support: The project also acknowledges application, product, frontend, backend, UX, testing, and demo-video support.
7 Appendix
The appendix reports detailed Grounding DINO 1.5 Pro zero-shot results on the ODinW35 benchmark.
- Table 6 reports detailed zero-shot results of Grounding DINO 1.5 Pro on the ODinW35 benchmark.