Evidence Grounded Vision Language Framework for Foreign Object Debris Analysis
DOI:
https://doi.org/10.13021/jssr2026.5614Abstract
Foreign Object Debris (FOD), including mechanical tools and sharp objects, poses a significant safety risk on airport
runways. Although current FOD monitoring systems improve situational awareness, they may lack the accuracy,
interpretability, and interactive support required for safety-critical decision-making. Furthermore, there is limited
research examining conversational interfaces for communicating FOD assessments to airport operators. This study
addresses this gap with an evidence-grounded framework for conversational FOD analysis using unmanned aerial
systems imagery. The framework integrates a binary FOD presence classifier, a multiclass FOD category classifier, a
centralized evidence-packet generator, and a pretrained vision language model (VLM). For each image, the generator
queries the classifiers and aggregates the FOD presence probability, category confidence scores, and image-attribution
outputs into a structured evidence packet. The VLM then processes the image, evidence packet, and analyst query to
generate an evidence-grounded response. Unlike task-specific VLM approaches, the proposed framework requires no
fine-tuning and relies only on few-shot prompting. We compare it with a baseline VLM supervised fine-tuned on a large
publicly available FOD dataset. Both approaches are evaluated on previously unseen FOD and non-FOD images using a
custom questionnaire designed to role play airport-analyst interactions. Compared with the fully fine-tuned VLM, our
framework improved FOD presence accuracy from 50% to 100%, category-identification accuracy from 45% to 90%, and
improved localization accuracy by 15%. These results show that centralized evidence grounding can improve response
reliability, traceability, and FOD-related decision support without task-specific VLM fine-tuning


