Abstract
MOTIVATION: Recent advancements in protein structure prediction methods have vastly increased the size of databases of protein structures, necessitating fast methods for protein structure comparison. Search methods that find structurally similar proteins can be applied to find remote homologs, study the functional relationships among proteins, and aid in protein engineering tasks. RESULTS: We design a "3Dn" structural alphabet that encodes the local neighborhoods around each amino acid in an interpretable way. In a search benchmark task, a combination of our alphabet and Foldseek's 3Di alphabet, outperforms each alphabet individually and ranks best among local search methods that do not require amino acid identity information. We provide software tools that enable the exploration of novel alphabets and combinations of alphabets for protein structure search. AVAILABILITY AND IMPLEMENTATION: The code is freely available at https://github.com/spetti/structure_comparison and at Zenodo https://doi.org/10.5281/zenodo.15734371.