IBB PAS Repository

GBSC: Graph-Based Sequence Clustering method for similar short tandem repeats in protein sequence

Jarnot, Patryk and Ziemska-Legiecka, Joanna and Grynberg, Marcin and Promponas, Vasilis J and Gruca, Aleksandra (2026) GBSC: Graph-Based Sequence Clustering method for similar short tandem repeats in protein sequence. Bioinformatics (Oxford, England) . btag378. ISSN 1367-4811

[img]
Preview
PDF
897kB

Official URL: https://academic.oup.com/bioinformatics/article/42...

Abstract

Motivation: Short tandem repeats (STRs) are abundant in protein sequences and play an important role in determining their structures and functions. Strikingly, the unusual compositional characteristics of tandem repeats break classical sequence analysis tools. Results: We here establish the first algorithm to effectively identify and cluster STRs: Graph-Based Sequence Clustering (GBSC) features linear time complexity, and clusters protein sequence fragments based on their STRs, while allowing for insertions and mutations, supporting an analysis of imperfect or cryptic repeats. Due to its computational efficacy, our algorithm can be used to systematically scan for patterns in large datasets. We compare our method both with state-of-the-art methods for identifying STRs in proteins and alternative clustering approaches. Unlike existing STR analysis methods, GBSC clusters repeat patterns rather than raw sequences, operating at the level of structural repeat identity, while tolerating biological variations and preventing erroneous merging of structurally and functionally distinct motifs. Whereas functional annotation is typically only available at the protein level, the functions of individual STRs and sequences of adjacent STRs remains largely unknown. On a challenging use-case we here demonstrate and discuss how our method can be used to associate previously unannotated repetitive protein fragments with similar ones, allowing a transfer of annotation by similarity. For the first time, GBSC offers a tool that systematically extends this fundamental bioinformatics principle to low-complexity regions across large datasets. Availability and implementation: GBSC is available at GitHub https://github.com/patryk-jarnot/GBSC and https://doi.org/10.5281/zenodo.18965247. The data and scripts to reproduce the analysis are available at https://doi.org/10.5281/zenodo.16906653. Keywords: Clustering; Functional Analysis; Identification; Protein Sequences; Short Tandem Repeats.

Item Type:Article
Subjects:Q Science > QH Natural history > QH301 Biology
ID Code:2650
Deposited By: Marcin Grynberg
Deposited On:05 Aug 2026 10:23
Last Modified:05 Aug 2026 10:23

Repository Staff Only: item control page