iTransaminase-PseAAC: Identification of Transaminase and Non-Transaminase Sites in Proteins Using Position-Relative Features and Statistical Moments

Available online July 1, 2025
PDF

Abstract

Accurately identifying transaminases from protein sequences helps in computational protein annotation. Transaminases are significant enzymes involved in amino-group transfer processes. This paper introduces iTransaminase-PseAAC, a sequence-based framework that uses statistical and position-relative properties to differentiate between transaminase and non-transaminase proteins. 3,077 transaminase and 2,500 non-transaminase protein sequences were included in the original dataset, which was gathered from UniProt. 1,002 positive and 2,129 negative sequences were obtained after using CD-HIT with a similarity threshold of 0.65 to remove sequence redundancy. Statistical moments were used in conjunction with the Position Relative Incidence Matrix (PRIM), Reverse Position Relative Incidence Matrix (RPRIM), Accumulative Absolute Position Incidence Vector (AAPIV), and Reverse Accumulative Absolute Position Incidence Vector (RAAPIV) to describe protein sequences. A Random Forest classifier was fed these characteristics, which capture amino acid composition, positional connections, and sequence orientation. Jackknife validation, 10-fold cross-validation, and self-consistency were used to assess the proposed model. For self-consistency testing and jackknife validation, the reported accuracies were 99.90\% and 99.93\%, respectively, with comparable excellent performance noted throughout the cross-validation folds. Additionally, a web interface for sequence-level prediction was created using Python. The results show that an efficient representation for transaminase classification on the reported benchmark may be obtained by combining statistical moments with position-relative sequence descriptors.

Keywords

Transaminase Protein sequence classification Random Forest PseAAC Statistical moments

Introduction

Sequence-based protein classification is a recurring problem in computational biology because experimentally determining protein function is costly and time-consuming. Pseudo-amino-acid-composition (PseAAC) representations and sequence-derived descriptors have therefore been used to encode information beyond simple residue frequencies [1]-[5]. The present work applies this general idea to the identification of transaminase proteins. Transaminases, also called aminotransferases, catalyze amino-group transfer between amino acids and keto acids. Their biochemical roles and substrate preferences have been studied experimentally and computationally [6]-[9]. Serum transaminase levels are also used clinically as biochemical indicators in several conditions, although the computational task addressed here is different: this paper predicts whether a protein sequence belongs to the transaminase class rather than diagnosing disease from patient measurements [10]-[13]. The problem is to derive a sequence representation that captures both composition and positional information, then use that representation for supervised classification. The proposed iTransaminase-PseAAC framework combines statistical moments with PRIM, RPRIM, AAPIV, and RAAPIV features and trains a Random Forest classifier. The evaluation uses self-consistency, leave-one-out (jackknife), and 10-fold cross-validation.

This paper’s primary contributions are:

  • A sequence representation that combines statistical moments with forward and reverse position-relative features

  • A Random Forest-based classifier evaluated using three validation procedures

  • A Python-based web interface for sequence-level prediction

  • A creation of a transaminase/non-transaminase benchmark from UniProt, followed by CD-HIT redundancy reduction at a 0.65 similarity cutoff.

The rest of the paper is organized as follows. The biological and computational background required to understand the sequence representation used in this investigation is given in Section 2. Related research on transaminases and sequence-based protein prediction is reviewed in Section 3. The Random Forest model, feature creation, and benchmark dataset are explained in Section 4. Section 5 presents the evaluation protocol. Section 6 reports and discusses the results, including the web interface and limitations. Section 7 concludes the paper.

Complete Article

The complete article, including all figures, tables, equations and algorithms, is available in the official publication PDF.

Conclusion

iTransaminase-PseAAC, a sequence-based computational framework for differentiating between transaminase and non-transaminase proteins, has been proposed in this paper. In order to express amino acid composition, relative residue locations, and sequence orientation inside the feature vector provided to a Random Forest classifier, the suggested representation integrates statistical moments with PRIM, RPRIM, AAPIV, and RAAPIV descriptors. A reduced benchmark with 1,002 positive and 2,129 negative sequences was created via CD-HIT redundancy reduction at a similarity threshold of 0.65 from the original sample, which had 3,077 transaminase and 2,500 non-transaminase sequences. Self-consistency, jackknife, and 10-fold cross-validation were used to assess the model. The fold-wise 10-fold cross-validation results gave an average accuracy of 99.91%, whereas the reported accuracies for self-consistency testing and jackknife validation were 99.90% and 99.93%, respectively. These results collectively show that the two protein classes on the presented benchmark may be successfully distinguished using the combined statistical and position-relative representation. The trained predictor was also made available for sequence-level categorization through the implementation of a Python-based web interface. Future work should therefore emphasize reproducible model settings, clearly documented feature dimensionality, and evaluation on an independent non-redundant protein benchmark. Such evaluation would provide stronger evidence of how well iTransaminase-PseAAC generalizes beyond the sequences used in the present paper.

References

  1. Chou, Kuo Chen. (2011). Some remarks on protein attribute prediction and pseudo amino acid composition. Journal of Theoretical Biology, 273(1), 236–247. https://doi.org/doi: 10.1016/j.jtbi.2010.12.024
  2. Akmal, M. A., Rasool, N., & Khan, Y. D. (2017). Prediction of N-linked glycosylation sites using position relative features and statistical moments. PLoS ONE, 12(8), 1–21. https://doi.org/doi: 10.1371/journal.pone.0181966
  3. Hussain, W., Khan, Y. D., Rasool, N., Khan, S. A., & Chou, K. C. (2019). SPalmitoylC-PseAAC: A sequence-based model developed via Chou’s 5-steps rule and general PseAAC for identifying S-palmitoylation sites in proteins. Analytical Biochemistry, 568(November 2018), 14–23. https://doi.org/doi: 10.1016/j.ab.2018.12.019
  4. Khan, Y. D., Jamil, M., Hussain, W., Rasool, N., Khan, S. A., & Chou, K. C. (2019). pSSbond-PseAAC: Prediction of disulfide bonding sites by integration of PseAAC and statistical moments. Journal of Theoretical Biology, 463, 47–55. https://doi.org/doi: 10.1016/j.jtbi.2018.12.015
  5. Khan, Y. D., Amin, N., Hussain, W., Rasool, N., Khan, S. A., & Chou, K. C. (2020). iProtease-PseAAC(2L): A two-layer predictor for identifying proteases and their types using Chou’s 5-step-rule and general PseAAC. Analytical Biochemistry, 588(February 2019), 113477. https://doi.org/doi: 10.1016/j.ab.2019.113477
  6. Tang, K., Yi, Y., Gao, Z., Jia, H., Li, Y., Cao, F., Wei, P. (2020). Identification , Heterologous Expression and Characterization of a Transaminase from Rhizobium sp . Catalysis Letters, (0123456789). https://doi.org/doi: 10.1007/s10562-020-03121-2
  7. Voß, M., Xiang, C., Esque, J., Nobili, A., Marian, J., André, I., Bornscheuer, U. T. (2020). within an # -amino acid transaminase scaffold Creation of ( R ) -amine transaminase activity within an $$-amino acid transaminase scaffold. https://doi.org/doi: 10.1021/acschembio.9b00888
  8. Cheng, F., Chen, X. L., Xiang, C., Liu, Z. Q., Wang, Y. J., & Zheng, Y. G. (2020). Fluorescence-based high-throughput screening system for R-$$-transaminase engineering and its substrate scope extension. Applied Microbiology and Biotechnology, 104(7), 2999–3009. https://doi.org/doi: 10.1007/s00253-020-10444-y
  9. Bowsher, R. R., & Henry, D. P. (2020). Neurochemistry International Purification , characterization and identification of rat brain cytosolic tyrosine transaminase as glutamine Transaminase-K. Neurochemistry International, 133(August 2019), 104653. https://doi.org/doi: 10.1016/j.neuint.2019.104653
  10. Zhou, X., Wang, Q., An, P., Du, Y., Zhao, J., & Song, A. (2019). Relationship between folate , vitamin B 12 , homocysteine , transaminase and mild cognitive impairment in China : a case-control study. International Journal of Food Sciences and Nutrition, 0(0), 1–10. https://doi.org/doi: 10.1080/09637486.2019.1648387
  11. Brumboiu, M. I., Brice, P., Ndemba, A., & Cazacu, I. (2019). Computing tools for analysing the pathological levels of serum transaminases, 41(1), 24–32.
  12. Dai, L., Ooi, V. V., Zhou, W., Ji, G., & Abd-Elsalam, S. (2020). Acupoint embedding therapy improves nonalcoholic fatty liver disease with abnormal transaminase: A PRISMA-compliant systematic review and meta-analysis. Medicine (United States), 99(3). https://doi.org/doi: 10.1097/MD.0000000000018775
  13. Yamamoto, J. M., Padro-Nuñez, S., Guarnizo-Poma, M., Lazaro-Alcantara, H., Paico-Palacios, S., Pantoja-Torres, B., Benites-Zapata, V. A. (2020). Association between serum transaminase levels and insulin resistance in euthyroid and non-diabetic adults: Serum transaminase levels and insulin resistance in healthy adults. Diabetes and Metabolic Syndrome: Clinical Research and Reviews, 14(1), 17–21. https://doi.org/doi: 10.1016/j.dsx.2019.11.013
  14. F. H. Crick. (1958). On protein synthesis. Symposia of the Society for Experimental Biology, 12, 138–163. https://doi.org/doi: 10.1038/227561a0
  15. Finkelstein, M., & Weissmann, G. (1978). The introduction of enzymes into cells by means of liposomes. Journal of Lipid Research, 19(3), 289–303.
  16. khanacademy. (2020). proteins-and-amino-acids/v/introduction-to-amino-acids.
  17. Ng, P. C., & Henikoff, S. (2003). SIFT: Predicting amino acid changes that affect protein function. Nucleic Acids Research, 31(13), 3812–3814. https://doi.org/doi: 10.1093/nar/gkg509
  18. Sheehan, J. C., & Hess, G. P. (1955). A new method of forming peptide bonds. Journal of the American Chemical Society, 77(4), 1067–1068. https://doi.org/doi: 10.1021/ja01609a099
  19. VELICK, S. F., & VAVRA, J. (1962). A kinetic and equilibrium analysis of the glutamic oxaloacetate transaminase mechanism. The Journal of Biological Chemistry, 237(7), 2109–2122.
  20. Oh, R. C., Hustead, T. R., Army, T., Medicine, F., & Program, R. (2011). Causes and Evaluation of Mildly Elevated Liver Transaminase Levels, 1003–1008.
  21. Matassa, C., Ormerod, D., Bornscheuer, U. T., Höhne, M., & Satyawali, Y. (2020). Three-liquid-phase Spinning Reactor for the Transaminase-catalyzed Synthesis and Recovery of a Chiral Amine. ChemCatChem, 12(5), 1288–1291. https://doi.org/doi: 10.1002/cctc.201902056
  22. Meng, Q., Capra, N., Palacio, C. M., Lanfranchi, E., Otzen, M., Van Schie, L. Z., Janssen, D. B. (2020). Robust $$-Transaminases by Computational Stabilization of the Subunit Interface. ACS Catalysis, 10(5), 2915–2928. https://doi.org/doi: 10.1021/acscatal.9b05223
  23. Matsumoto, T., Mori, Y., Tanaka, T., & Kondo, A. (2020). n-Butylamine production from glucose using a transaminase-mediated synthetic pathway in Escherichia coli. Journal of Bioscience and Bioengineering, 129(1), 99–103. https://doi.org/doi: 10.1016/j.jbiosc.2019.06.015
  24. Yoshida, T., Yamasaki, S., Kaneko, O., Taoka, N., Tomimoto, Y., Namatame, I., Lyssiotis, C. A. (2020). A covalent small molecule inhibitor of glutamate-oxaloacetate transaminase 1 impairs pancreatic cancer growth. Biochemical and Biophysical Research Communications, 522(3), 633–638. https://doi.org/doi: 10.1016/j.bbrc.2019.11.130
  25. Planchestainer, M., Hegarty, E., Gourlay, L. J., Paradisi, F., & Heckmann, C. M. (2019). Chemical Science Widely applicable background depletion step enables transaminase evolution through solid- phase screening , 5952–5958. https://doi.org/doi: 10.1039/c8sc05712e
  26. Hegde, A. U., Karnavat, P. K., Vyas, R., Dibacco, M. L., Grant, P. E., & Pearl, P. L. (2019). GABA Transaminase Deficiency With Survival Into Adulthood. https://doi.org/doi: 10.1177/0883073818823359
  27. Khan, Y. D., Rasool, N., Hussain, W., Khan, S. A., & Chou, K. C. (2018). iPhosY-PseAAC: identify phosphotyrosine sites by incorporating sequence statistical moments into PseAAC. Molecular Biology Reports, 45(6), 2501–2509. https://doi.org/doi: 10.1007/s11033-018-4417-z
  28. Dehzangi, A., Heffernan, R., Sharma, A., Lyons, J., Paliwal, K., & Sattar, A. (2015). Gram-positive and Gram-negative protein subcellular localization by incorporating evolutionary-based descriptors into Chou’s general PseAAC. Journal of Theoretical Biology, 364, 284–294. https://doi.org/doi: 10.1016/j.jtbi.2014.09.029
  29. Dou, Y., Yao, B., & Zhang, C. (2014). PhosphoSVM: Prediction of phosphorylation sites by integrating various protein sequence attributes with a support vector machine. Amino Acids, 46(6), 1459–1469. https://doi.org/doi: 10.1007/s00726-014-1711-5
  30. Feng, K. Y., Cai, Y. D., & Chou, K. C. (2005). Boosting classifier for predicting protein domain structural class. Biochemical and Biophysical Research Communications, 334(1), 213–217. https://doi.org/doi: 10.1016/j.bbrc.2005.06.075
  31. Kumar, R., Srivastava, A., Kumari, B., & Kumar, M. (2015). Prediction of $$-lactamase and its class by Chou’s pseudo-amino acid composition and support vector machine. Journal of Theoretical Biology, 365, 96–103. https://doi.org/doi: 10.1016/j.jtbi.2014.10.008
  32. Mondal, S., & Pai, P. P. (2014). Chou’s pseudo amino acid composition improves sequence-based antifreeze protein prediction. Journal of Theoretical Biology, 356, 30–35. https://doi.org/doi: 10.1016/j.jtbi.2014.04.006
  33. Nanni, L., Brahnam, S., & Lumini, A. (2014). Prediction of protein structure classes by incorporating different protein descriptors into general Chou’s pseudo amino acid composition. Journal of Theoretical Biology, 360, 109–116. https://doi.org/doi: 10.1016/j.jtbi.2014.07.003
  34. Fu, L., Niu, B., Zhu, Z., Wu, S., & Li, W. (2012). CD-HIT: Accelerated for clustering the next-generation sequencing data. Bioinformatics, 28(23), 3150–3152. https://doi.org/doi: 10.1093/bioinformatics/bts565
  35. Chou, K. C. (2001b). Using subsite coupling to predict signal peptides. Protein Engineering, 14(2), 75–79. https://doi.org/doi: 10.1093/protein/14.2.75
  36. Altay, G., & Emmert-Streib, F. (2010). Revealing differences in gene network inference algorithms on the network level by ensemble methods. Bioinformatics, 26(14), 1738–1744. https://doi.org/doi: 10.1093/bioinformatics/btq259
  37. Wan, S., Mak, M. W., & Kung, S. Y. (2016). Ensemble Linear Neighborhood Propagation for Predicting Subchloroplast Localization of Multi-Location Proteins. Journal of Proteome Research, 15(12), 4755–4762. https://doi.org/doi: 10.1021/acs.jproteome.6b00686
  38. Wan, S., Mak, M. W., & Kung, S. Y. (2017). Transductive Learning for Multi-Label Protein Subchloroplast Localization Prediction. IEEE/ACM Transactions on Computational Biology and Bioinformatics, 14(1), 212–224. https://doi.org/doi: 10.1109/TCBB.2016.2527657
  39. Khan, Y. D., Ahmad, F., & Anwar, M. W. (2012). A neuro-cognitive approach for Iris recognition using back propagation. World Applied Sciences Journal, 16(5), 678–685.
  40. Khan, Y. D., Khan, S. A., Ahmad, F., & Islam, S. (2014). Iris recognition using image moments and k-Means algorithm. The Scientific World Journal, 2014. https://doi.org/doi: 10.1155/2014/723595
  41. Wu, C. H. (1997). Artificial neural networks for molecular sequence analysis. Computers and Chemistry, 21(4), 237–256. https://doi.org/doi: 10.1016/S0097-8485(96)00038-1
  42. Bouaziz, A., Dartigues-Pallez, C., Da Costa Pereira, C., Precioso, F., & Lloret, P. (2014). Short text classification using semantic random forest. Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 8646 LNCS, 288–299. https://doi.org/doi: 10.1007/978-3-319-10160-6_26
  43. Pavlov, Y. L. (2019). Random forests. Random Forests, 1–122. https://doi.org/doi: 10.1201/9780367816377-11
  44. Bienvenido-huertas, D., Rubio-bellido, C., & P, J. L. (2020). Automation and optimization of in-situ assessment of wall thermal transmittance using a Random Forest algorithm, 168(August 2019). https://doi.org/doi: 10.1016/j.buildenv.2019.106479
  45. Liu, Z. P., Wu, L. Y., Wang, Y., Zhang, X. S., & Chen, L. (2010). Prediction of protein-RNA binding sites by a random forest method with combined features. Bioinformatics, 26(13), 1616–1622. https://doi.org/doi: 10.1093/bioinformatics/btq253
  46. Chou, K. C. (2001a). Prediction of signal peptides using scaled window. Peptides, 22(12), 1973–1979. https://doi.org/doi: 10.1016/S0196-9781(01)00540-X
  47. Feng, P., Ding, H., Chen, W., & Lin, H. (2013). Na ve Bayes Classifier with Feature Selection to Identify Phage Virion Proteins, 2013.
  48. Xu, Y., Shao, X., Wu, L., & Deng, N. (2013). iSNO-AAPair : incorporating amino acid pairwise coupling into PseAAC for predicting cysteine S -nitrosylation sites in proteins, 1–18. https://doi.org/doi: 10.7717/peerj.171
  49. Cai, Y. D., & Zhou, G. P. (2000). Prediction of protein structural classes by neural network. Biochimie, 82(8), 783–785. https://doi.org/doi: 10.1016/S0300-9084(00)01161-5
  50. Qiu, W. R., Jiang, S. Y., Xu, Z. C., Xiao, X., & Chou, K. C. (2017). iRNAm5C-PseDNC: Identifying RNA 5-methylcytosine sites by incorporating physical-chemical properties into pseudo dinucleotide composition. Oncotarget, 8(25), 41178–41188. https://doi.org/doi: 10.18632/oncotarget.17104
  51. Qiu, W. R., Sun, B. Q., Xiao, X., Xu, D., & Chou, K. C. (2017). iPhos-PseEvo: Identifying Human Phosphorylated Proteins by Incorporating Evolutionary Information into General PseAAC via Grey System Theory. Molecular Informatics, 36(5), 1–10. https://doi.org/doi: 10.1002/minf.201600010
  52. Chou, K.-C. (2015). Impacts of Bioinformatics to Medicinal Chemistry. Medicinal Chemistry, 11(3), 218–234. https://doi.org/doi: 10.2174/1573406411666141229162834
  53. Sigma, S. (2005). N Atural S P I, 1(2), 1–27.
  54. Cheng, F., Chen, X., Xiang, C., Liu, Z., & Wang, Y. (2020). Fluorescence-based high-throughput screening system for R - $$ -transaminase engineering and its substrate scope extension.
  55. Mussardo, G. (2019). No Title No Title. Statistical Field Theor, 53(9), 1689–1699. https://doi.org/doi: 10.1017/CBO9781107415324.004
  56. Wang, T., Zhao, B., & Foehr, E. D. (2019). Engineering a Long Lasting Tethered, Multimeric Human Growth Hormone Protein to Improve Pharmacokinetic Half-Life and Potency. Journal of Proteomics & Bioinformatics, 12(3), 56–60. https://doi.org/doi: 10.35248/0974-276x.19.12.497
10 6

Similar Articles

You may also start an advanced similarity search for this article.