Substituting and Augmenting Real Imagery with Blender-Generated Synthetic Data for Unmanned Aircraft System Based Multispectral Landmine Detection Using YOLOv11 and RF-DETR
DOI:
https://doi.org/10.13021/jssr2026.5585Abstract
In 2024, landmines were the cause of over 6,000 recorded casualties (90% civilian), with children making up a large percentage of victims. Previous studies established the usage of Unmanned Aerial Systems (UAS) as a viable method for landmine detection, specifically through the use of multispectral imagery paired with advanced computer vision systems trained on real minefield imagery. Synthetic imagery allows for augmenting such datasets, reducing the cost and time needed to train models for different environments, but the impact of different ratios of real and synthetic data on the accuracy of different models has yet to be explored in depth. We used Blender to generate RGB-LWIR imagery through the use of heatmaps, training 32 models, with 16 training sets of both YOLOv11 and RF-DETR, keeping 15 sets at a constant 10,756 images and scoring all models on 1,320 real RGB-LWIR images across 264 stratified conditions. A replacement of half of the real imagery with rendered imagery reduced mean Average Precision (mAP) at an Intersection over Union of 50% by 1.7–2.4 points for both YOLOv11 and RF-DETR. No size-matched training set beat the real-only set. However, the larger training set improved YOLOv11 mAP by 4.9 points, using synthetic imagery to augment rather than replace real data, but this set did not improve RF-DETR performance. Overall, up to half of a real multispectral training set is replaceable by rendered imagery at near-noise cost, while synthetic augmentation of the full real dataset yields the strongest performance improvements.


