Computational Prediction of a Most Likely Binding Protein Given an Arbitrary DNA Sequence

Authors

  • Riva Dehadaraya Department of Chemistry and Biochemistry, George Mason University, Fairfax, VA
  • Kenneth Foreman Department of Chemistry and Biochemistry, George Mason University, Fairfax, VA

DOI:

https://doi.org/10.13021/jssr2026.5557

Abstract

Protein-DNA interactions play a crucial role in several biological functions such as gene expression, transcription, and DNA replication. Understanding these interactions can help facilitate the design of artificial DNA-binding proteins and can support the development of gene editing tools such as CRISPR. Previous studies in this field have focused more on predicting binding affinity and specificity based on sequence rather than physical determinants. To address this gap, we developed an algorithm to predict the most likely binding protein given an arbitrary DNA sequence. We generated a database of physical interaction maps from various RCSB Protein Data Bank protein-DNA crystal structures. The algorithm takes as input a query DNA sequence and converts that to a bit representation for comparison against the reference database. Then, the algorithm performs a bit match between the sequence and reference database, ranks candidate proteins, and outputs the most likely binding protein with a confidence score. Initial testing on known protein-DNA complexes suggests that the algorithm can successfully identify the correct binding proteins. These results indicate that our algorithm provides a computationally efficient approach to predict protein-DNA binding compared to more computationally expensive tools or those which need extensive experimental data. While advantageous, the algorithm currently makes predictions only for proteins in the RCSB DATA Bank, suggesting that extending its capabilities to all potential proteins remains as a next endeavor.

Published

2026-09-24

Issue

Section

College of Science: Department of Chemistry and Biochemistry