Substituting and Augmenting Real Imagery with Synthetic Data for Unmanned Aerial Systems-Based Multispectral Landmine Detection
DOI:
https://doi.org/10.13021/jssr2026.5590Abstract
In 2024, landmines were the cause of over 6,000 casualties (90% civilian), with children making up a large percentage of victims. Previous studies established Unmanned Aerial Systems (UAS) as a viable method for landmine detection, specifically through the use of multispectral imagery paired with advanced computer vision systems trained on real minefield imagery. Synthetic imagery can augment such datasets and reduce training cost and time across environments, but the impact of different ratios of real and synthetic data on model accuracy has yet to be explored in depth. To explore this, we trained 30 models, with 15 training sets of both YOLOv11 and RF-DETR, keeping each set at a constant 10,756 images and scoring all models on 1,320 real RGB-LWIR images across 264 stratified conditions. Compared to an all-real baseline of 71.4 and 73.1% mean Average Precision (mAP) at an Intersection over Union of 50% for YOLOv11 and RF-DETR, respectively, real imagery could be replaced with either blender-rendered or GAN-refined photographs at a 50:50 ratio at minimal accuracy cost of only 1.7-2.6 mAP@50 for both YOLOv11 and RF-DETR. However, substitution using GAN-refined renders resulted in a substantially higher cost of 17.3-19.9 points at the same ratio. Fully synthetic datasets reduced performance substantially for both architectures. No substituted set matched the all-real baseline, but RF-DETR maintained higher accuracy than YOLOv11 at 0% real imagery (72% vs. 48% of real-data performance retained), suggesting that the cost of substitution at a given ratio depends on both the synthetic data source and model used.


