Synthetic-to-Real Imagery Substitution and Augmentation for Multispectral Landmine Detection

Authors

  • Kabir Raghavan Department of Geography and Geoinformation Science, George Mason University, Fairfax, VA
  • Evan Li Department of Geography and Geoinformation Science, George Mason University, Fairfax, VA
  • Bhargav Mandakolathur Department of Geography and Geoinformation Science, George Mason University, Fairfax, VA
  • Aryan Ruia Department of Geography and Geoinformation Science, George Mason University, Fairfax, VA
  • James E. Gallagher Department of Geography and Geoinformation Science, George Mason University, Fairfax, VA
  • Edward Oughton Department of Geography and Geoinformation Science, George Mason University, Fairfax, VA

DOI:

https://doi.org/10.13021/jssr2026.5729

Abstract

Landmines caused more than 6,000 recorded casualties in 2024. 90% of these casualties were civilians, and nearly half of those civilians were children. Unmanned aerial systems (UAS) equipped with multispectral sensors and computer vision models can accelerate demining operations by detecting surface-laid mines. However, detection accuracy is constrained by scarce labeled landmine training data, which is difficult to collect and rarely matches the environment where models deploy. This research evaluates whether synthetic imagery can close this gap by testing sixteen training datasets spanning several real-to-synthetic ratios. Two computer vision models, YOLOv11 and RF-DETR, were trained on each unique dataset, producing 32 models. Every model was evaluated on the same held-out test set of 1,320 real RGB-LWIR images, covering 264 combinations of season, illumination, altitude, and sensor fusion level. Synthetic imagery came from three sources: Blender scenes constructed using a Vision Language Model (VLM), a Generative Adversarial Network (GAN) applied to real-world imagery, and the same GAN applied to the Blender-VLM renders. Replacing half of the real imagery with Blender renders or GAN-refined real images reduced mean average precision (mAP@0.5) by only 1.7 to 2.6 points, whereas the same substitution using GAN-refined Blender-VLM renders reduced performance by 17.3 to 19.9 points. No training dataset ratio mix outperformed real-only training. However, augmenting the complete real-world dataset with all synthetic sources increased YOLOv11 from 71.4% to 76.3% mAP, while RF-DETR performed best on real imagery alone. These results demonstrate that synthetic imagery can replace up to half of a real multispectral training dataset with negligible performance loss, providing practical guidance for reducing data collection costs in humanitarian demining research.

Published

2026-09-24

Issue

Section

College of Science: Department of Geography and Geoinformation Science