Evaluating LLM Failure in Disaster Recovery Regulatory Reasoning

Authors

  • Alice Li Sid and Reva Dewberry Department of Civil, Environmental, and Infrastructure Engineering, George Mason University, Fairfax, VA
  • Kavya Kumar Sid and Reva Dewberry Department of Civil, Environmental, and Infrastructure Engineering, George Mason University, Fairfax, VA
  • Dheeraj Nagula Sid and Reva Dewberry Department of Civil, Environmental, and Infrastructure Engineering, George Mason University, Fairfax, VA
  • Rohith Swaminathan Sid and Reva Dewberry Department of Civil, Environmental, and Infrastructure Engineering, George Mason University, Fairfax, VA
  • Catalina González-Dueñas Sid and Reva Dewberry Department of Civil, Environmental, and Infrastructure Engineering, George Mason University, Fairfax, VA

DOI:

https://doi.org/10.13021/jssr2026.5564

Abstract

When a disaster strikes, a community's ability to recover depends not only on the damage it suffers, but also on how accurately it can interpret FEMA’s complex funding rules. The Federal Emergency Management Agency (FEMA) administers billions of dollars in funding through Public Assistance (PA) and Individual Assistance (IA) programs to help communities and individuals recover from natural disasters. However, determining eligibility for funding requires interpreting dense, highly conditional FEMA regulations. FEMA has deployed a Large Language Model (LLM) chatbot to help answer policy questions, but LLMs are known to hallucinate and struggle with regulatory reasoning. This project establishes a framework to evaluate whether standard LLMs can process the dense, hierarchical, and cross-referenced structure of disaster and procurement regulations independently of any situational context. To evaluate this, we have defined six vulnerability classes: conditionality collapse, jurisdiction misattribution, temporal and state-dependent reasoning failure, cross-reference overconfidence, deontic language softening, and hallucination under specificity pressure. To create these scenarios, we have developed a pipeline that semantically chunks FEMA policy documents, identifies passages associated with the vulnerability classes, and generates benchmark scenarios containing a policy question, regulatory context, and ground-truth answer. The pipeline generated 1,379 benchmark scenarios across the six vulnerability classes, forming a structured test set for FEMA regulatory reasoning. Each scenario includes a vulnerability label, confidence score, explanation, scenario question, and supporting FEMA quote. Preliminary testing of the conditionality collapse and jurisdiction misattribution failure modes suggests the pipeline successfully identifies policy passages complex enough to become meaningful evaluation questions. These benchmarks provide a standardized way to evaluate LLM reasoning over FEMA Public Assistance policy and support the development of more reliable AI tools for disaster funding decisions.

Published

2026-09-24

Issue

Section

College of Engineering and Computing: Department of Civil, Environmental and Infrastructure Engineering