EFI - Enzyme Similarity Tool

This web resource is supported by a Research Resource from the National Institute of General Medical Sciences (R24GM141196-01).
The tools are available without charge or license to both academic and commercial users.
Reorganization of UniProtKB

With the current 2026_02 release, the UniProtKB database is reorganized to include an expanded number of Reference Proteomes to better capture biodiversity. This includes the removal of proteins from taxonomically unclassified organisms, i.e., those without a binomial species name (genus and species). The total number of accessions in UniProtKB has been reduced from 253,635,358 in the “legacy” 2025_03 release to 149,810,139 in the current 2026_02 release.

We are providing the option to select either the “legacy” 2025_03 database or the current UniProtKB database (now 2026_02) when generating SSNs. You can select the database in the “Database” accordion on the pages for the EFI-EST options, the EFI-GNT tool, and the Taxonomy Tool. We suggest that you compare the SSNs, GNNs, and GNDs generated from both databases as you explore the information you are seeking.

Because the “legacy” 2025_03 release contains UniProt IDs that are no longer active on the UniProt web site, we provide the Metadata Tool that provides access to the node attribute metadata for the UniProt IDs in the “legacy” 2025_03 release.

EFI-EST and Cytoscape Tutorials

References

1.Using sequence similarity networks for visualization of relationships across diverse protein superfamilies. Atkinson HJ, Morris JH, Ferrin TE, Babbitt PC (2009) PLoS One 4, e4345.

2. Pythoscape: a framework for generation of large protein similarity networks. Barber AE 2nd, Babbitt PC (2012) Bioinformatics 28, 2845-6.

3. Cytoscape: a software environment for integrated models of biomolecular interaction networks. Shannon P, Markiel A, Ozier O, Baliga NS, Wang JT, Ramage D, Amin N, Schwikowski B, Ideker T (2003) Genome Res 13, 2498-504.

4. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. Weizhong L, Godzik A (2006) Bioinformatics 22, 1658-9. 

5. CD-HIT: accelerated for clustering the next generation sequencing data. Fu L, Niu B, Zhu Z, Wu S, Li W (2012) Bioinformatics 28, 3150-2.

6. Enzyme Function Initiative-Enzyme Similarity Tool (EFI-EST): A web tool for generating protein sequence similarity networks. Gerlt JA, Bouvier JT, Davidson DB, Imker HJ, Sadkhin B, Slater DR, Whalen KL. (2015) Biochimica et Biophysica Acta (BBA) - Proteins and Proteomics doi:10.1016/j.bbapap.2015.04.015

Click here to contact us for help, reporting issues, or suggestions.