Abstract
Keywords
Introduction
Sequence-based protein classification is a recurring problem in computational biology because experimentally determining protein function is costly and time-consuming. Pseudo-amino-acid-composition (PseAAC) representations and sequence-derived descriptors have therefore been used to encode information beyond simple residue frequencies [1]-[5]. The present work applies this general idea to the identification of transaminase proteins. Transaminases, also called aminotransferases, catalyze amino-group transfer between amino acids and keto acids. Their biochemical roles and substrate preferences have been studied experimentally and computationally [6]-[9]. Serum transaminase levels are also used clinically as biochemical indicators in several conditions, although the computational task addressed here is different: this paper predicts whether a protein sequence belongs to the transaminase class rather than diagnosing disease from patient measurements [10]-[13]. The problem is to derive a sequence representation that captures both composition and positional information, then use that representation for supervised classification. The proposed iTransaminase-PseAAC framework combines statistical moments with PRIM, RPRIM, AAPIV, and RAAPIV features and trains a Random Forest classifier. The evaluation uses self-consistency, leave-one-out (jackknife), and 10-fold cross-validation.
This paper’s primary contributions are:
A sequence representation that combines statistical moments with forward and reverse position-relative features
A Random Forest-based classifier evaluated using three validation procedures
A Python-based web interface for sequence-level prediction
A creation of a transaminase/non-transaminase benchmark from UniProt, followed by CD-HIT redundancy reduction at a 0.65 similarity cutoff.
The rest of the paper is organized as follows. The biological and computational background required to understand the sequence representation used in this investigation is given in Section 2. Related research on transaminases and sequence-based protein prediction is reviewed in Section 3. The Random Forest model, feature creation, and benchmark dataset are explained in Section 4. Section 5 presents the evaluation protocol. Section 6 reports and discusses the results, including the web interface and limitations. Section 7 concludes the paper.
Complete Article
The complete article, including all figures, tables, equations and algorithms, is available in the official publication PDF.
Conclusion
iTransaminase-PseAAC, a sequence-based computational framework for differentiating between transaminase and non-transaminase proteins, has been proposed in this paper. In order to express amino acid composition, relative residue locations, and sequence orientation inside the feature vector provided to a Random Forest classifier, the suggested representation integrates statistical moments with PRIM, RPRIM, AAPIV, and RAAPIV descriptors. A reduced benchmark with 1,002 positive and 2,129 negative sequences was created via CD-HIT redundancy reduction at a similarity threshold of 0.65 from the original sample, which had 3,077 transaminase and 2,500 non-transaminase sequences. Self-consistency, jackknife, and 10-fold cross-validation were used to assess the model. The fold-wise 10-fold cross-validation results gave an average accuracy of 99.91%, whereas the reported accuracies for self-consistency testing and jackknife validation were 99.90% and 99.93%, respectively. These results collectively show that the two protein classes on the presented benchmark may be successfully distinguished using the combined statistical and position-relative representation. The trained predictor was also made available for sequence-level categorization through the implementation of a Python-based web interface. Future work should therefore emphasize reproducible model settings, clearly documented feature dimensionality, and evaluation on an independent non-redundant protein benchmark. Such evaluation would provide stronger evidence of how well iTransaminase-PseAAC generalizes beyond the sequences used in the present paper.
References
- Chou, Kuo Chen. (2011). Some remarks on protein attribute prediction and pseudo amino acid composition. Journal of Theoretical Biology, 273(1), 236–247. https://doi.org/doi: 10.1016/j.jtbi.2010.12.024
- Akmal, M. A., Rasool, N., & Khan, Y. D. (2017). Prediction of N-linked glycosylation sites using position relative features and statistical moments. PLoS ONE, 12(8), 1–21. https://doi.org/doi: 10.1371/journal.pone.0181966
- Hussain, W., Khan, Y. D., Rasool, N., Khan, S. A., & Chou, K. C. (2019). SPalmitoylC-PseAAC: A sequence-based model developed via Chou’s 5-steps rule and general PseAAC for identifying S-palmitoylation sites in proteins. Analytical Biochemistry, 568(November 2018), 14–23. https://doi.org/doi: 10.1016/j.ab.2018.12.019
- Khan, Y. D., Jamil, M., Hussain, W., Rasool, N., Khan, S. A., & Chou, K. C. (2019). pSSbond-PseAAC: Prediction of disulfide bonding sites by integration of PseAAC and statistical moments. Journal of Theoretical Biology, 463, 47–55. https://doi.org/doi: 10.1016/j.jtbi.2018.12.015
- Khan, Y. D., Amin, N., Hussain, W., Rasool, N., Khan, S. A., & Chou, K. C. (2020). iProtease-PseAAC(2L): A two-layer predictor for identifying proteases and their types using Chou’s 5-step-rule and general PseAAC. Analytical Biochemistry, 588(February 2019), 113477. https://doi.org/doi: 10.1016/j.ab.2019.113477
- Tang, K., Yi, Y., Gao, Z., Jia, H., Li, Y., Cao, F., Wei, P. (2020). Identification , Heterologous Expression and Characterization of a Transaminase from Rhizobium sp . Catalysis Letters, (0123456789). https://doi.org/doi: 10.1007/s10562-020-03121-2
- Voß, M., Xiang, C., Esque, J., Nobili, A., Marian, J., André, I., Bornscheuer, U. T. (2020). within an # -amino acid transaminase scaffold Creation of ( R ) -amine transaminase activity within an $$-amino acid transaminase scaffold. https://doi.org/doi: 10.1021/acschembio.9b00888
- Cheng, F., Chen, X. L., Xiang, C., Liu, Z. Q., Wang, Y. J., & Zheng, Y. G. (2020). Fluorescence-based high-throughput screening system for R-$$-transaminase engineering and its substrate scope extension. Applied Microbiology and Biotechnology, 104(7), 2999–3009. https://doi.org/doi: 10.1007/s00253-020-10444-y
- Bowsher, R. R., & Henry, D. P. (2020). Neurochemistry International Purification , characterization and identification of rat brain cytosolic tyrosine transaminase as glutamine Transaminase-K. Neurochemistry International, 133(August 2019), 104653. https://doi.org/doi: 10.1016/j.neuint.2019.104653
- Zhou, X., Wang, Q., An, P., Du, Y., Zhao, J., & Song, A. (2019). Relationship between folate , vitamin B 12 , homocysteine , transaminase and mild cognitive impairment in China : a case-control study. International Journal of Food Sciences and Nutrition, 0(0), 1–10. https://doi.org/doi: 10.1080/09637486.2019.1648387
- Brumboiu, M. I., Brice, P., Ndemba, A., & Cazacu, I. (2019). Computing tools for analysing the pathological levels of serum transaminases, 41(1), 24–32.
- Dai, L., Ooi, V. V., Zhou, W., Ji, G., & Abd-Elsalam, S. (2020). Acupoint embedding therapy improves nonalcoholic fatty liver disease with abnormal transaminase: A PRISMA-compliant systematic review and meta-analysis. Medicine (United States), 99(3). https://doi.org/doi: 10.1097/MD.0000000000018775
- Yamamoto, J. M., Padro-Nuñez, S., Guarnizo-Poma, M., Lazaro-Alcantara, H., Paico-Palacios, S., Pantoja-Torres, B., Benites-Zapata, V. A. (2020). Association between serum transaminase levels and insulin resistance in euthyroid and non-diabetic adults: Serum transaminase levels and insulin resistance in healthy adults. Diabetes and Metabolic Syndrome: Clinical Research and Reviews, 14(1), 17–21. https://doi.org/doi: 10.1016/j.dsx.2019.11.013
- F. H. Crick. (1958). On protein synthesis. Symposia of the Society for Experimental Biology, 12, 138–163. https://doi.org/doi: 10.1038/227561a0
- Finkelstein, M., & Weissmann, G. (1978). The introduction of enzymes into cells by means of liposomes. Journal of Lipid Research, 19(3), 289–303.
- khanacademy. (2020). proteins-and-amino-acids/v/introduction-to-amino-acids.
- Ng, P. C., & Henikoff, S. (2003). SIFT: Predicting amino acid changes that affect protein function. Nucleic Acids Research, 31(13), 3812–3814. https://doi.org/doi: 10.1093/nar/gkg509
- Sheehan, J. C., & Hess, G. P. (1955). A new method of forming peptide bonds. Journal of the American Chemical Society, 77(4), 1067–1068. https://doi.org/doi: 10.1021/ja01609a099
- VELICK, S. F., & VAVRA, J. (1962). A kinetic and equilibrium analysis of the glutamic oxaloacetate transaminase mechanism. The Journal of Biological Chemistry, 237(7), 2109–2122.
- Oh, R. C., Hustead, T. R., Army, T., Medicine, F., & Program, R. (2011). Causes and Evaluation of Mildly Elevated Liver Transaminase Levels, 1003–1008.
- Matassa, C., Ormerod, D., Bornscheuer, U. T., Höhne, M., & Satyawali, Y. (2020). Three-liquid-phase Spinning Reactor for the Transaminase-catalyzed Synthesis and Recovery of a Chiral Amine. ChemCatChem, 12(5), 1288–1291. https://doi.org/doi: 10.1002/cctc.201902056
- Meng, Q., Capra, N., Palacio, C. M., Lanfranchi, E., Otzen, M., Van Schie, L. Z., Janssen, D. B. (2020). Robust $$-Transaminases by Computational Stabilization of the Subunit Interface. ACS Catalysis, 10(5), 2915–2928. https://doi.org/doi: 10.1021/acscatal.9b05223
- Matsumoto, T., Mori, Y., Tanaka, T., & Kondo, A. (2020). n-Butylamine production from glucose using a transaminase-mediated synthetic pathway in Escherichia coli. Journal of Bioscience and Bioengineering, 129(1), 99–103. https://doi.org/doi: 10.1016/j.jbiosc.2019.06.015
- Yoshida, T., Yamasaki, S., Kaneko, O., Taoka, N., Tomimoto, Y., Namatame, I., Lyssiotis, C. A. (2020). A covalent small molecule inhibitor of glutamate-oxaloacetate transaminase 1 impairs pancreatic cancer growth. Biochemical and Biophysical Research Communications, 522(3), 633–638. https://doi.org/doi: 10.1016/j.bbrc.2019.11.130
- Planchestainer, M., Hegarty, E., Gourlay, L. J., Paradisi, F., & Heckmann, C. M. (2019). Chemical Science Widely applicable background depletion step enables transaminase evolution through solid- phase screening , 5952–5958. https://doi.org/doi: 10.1039/c8sc05712e
- Hegde, A. U., Karnavat, P. K., Vyas, R., Dibacco, M. L., Grant, P. E., & Pearl, P. L. (2019). GABA Transaminase Deficiency With Survival Into Adulthood. https://doi.org/doi: 10.1177/0883073818823359
- Khan, Y. D., Rasool, N., Hussain, W., Khan, S. A., & Chou, K. C. (2018). iPhosY-PseAAC: identify phosphotyrosine sites by incorporating sequence statistical moments into PseAAC. Molecular Biology Reports, 45(6), 2501–2509. https://doi.org/doi: 10.1007/s11033-018-4417-z
- Dehzangi, A., Heffernan, R., Sharma, A., Lyons, J., Paliwal, K., & Sattar, A. (2015). Gram-positive and Gram-negative protein subcellular localization by incorporating evolutionary-based descriptors into Chou’s general PseAAC. Journal of Theoretical Biology, 364, 284–294. https://doi.org/doi: 10.1016/j.jtbi.2014.09.029
- Dou, Y., Yao, B., & Zhang, C. (2014). PhosphoSVM: Prediction of phosphorylation sites by integrating various protein sequence attributes with a support vector machine. Amino Acids, 46(6), 1459–1469. https://doi.org/doi: 10.1007/s00726-014-1711-5
- Feng, K. Y., Cai, Y. D., & Chou, K. C. (2005). Boosting classifier for predicting protein domain structural class. Biochemical and Biophysical Research Communications, 334(1), 213–217. https://doi.org/doi: 10.1016/j.bbrc.2005.06.075
- Kumar, R., Srivastava, A., Kumari, B., & Kumar, M. (2015). Prediction of $$-lactamase and its class by Chou’s pseudo-amino acid composition and support vector machine. Journal of Theoretical Biology, 365, 96–103. https://doi.org/doi: 10.1016/j.jtbi.2014.10.008
- Mondal, S., & Pai, P. P. (2014). Chou’s pseudo amino acid composition improves sequence-based antifreeze protein prediction. Journal of Theoretical Biology, 356, 30–35. https://doi.org/doi: 10.1016/j.jtbi.2014.04.006
- Nanni, L., Brahnam, S., & Lumini, A. (2014). Prediction of protein structure classes by incorporating different protein descriptors into general Chou’s pseudo amino acid composition. Journal of Theoretical Biology, 360, 109–116. https://doi.org/doi: 10.1016/j.jtbi.2014.07.003
- Fu, L., Niu, B., Zhu, Z., Wu, S., & Li, W. (2012). CD-HIT: Accelerated for clustering the next-generation sequencing data. Bioinformatics, 28(23), 3150–3152. https://doi.org/doi: 10.1093/bioinformatics/bts565
- Chou, K. C. (2001b). Using subsite coupling to predict signal peptides. Protein Engineering, 14(2), 75–79. https://doi.org/doi: 10.1093/protein/14.2.75
- Altay, G., & Emmert-Streib, F. (2010). Revealing differences in gene network inference algorithms on the network level by ensemble methods. Bioinformatics, 26(14), 1738–1744. https://doi.org/doi: 10.1093/bioinformatics/btq259
- Wan, S., Mak, M. W., & Kung, S. Y. (2016). Ensemble Linear Neighborhood Propagation for Predicting Subchloroplast Localization of Multi-Location Proteins. Journal of Proteome Research, 15(12), 4755–4762. https://doi.org/doi: 10.1021/acs.jproteome.6b00686
- Wan, S., Mak, M. W., & Kung, S. Y. (2017). Transductive Learning for Multi-Label Protein Subchloroplast Localization Prediction. IEEE/ACM Transactions on Computational Biology and Bioinformatics, 14(1), 212–224. https://doi.org/doi: 10.1109/TCBB.2016.2527657
- Khan, Y. D., Ahmad, F., & Anwar, M. W. (2012). A neuro-cognitive approach for Iris recognition using back propagation. World Applied Sciences Journal, 16(5), 678–685.
- Khan, Y. D., Khan, S. A., Ahmad, F., & Islam, S. (2014). Iris recognition using image moments and k-Means algorithm. The Scientific World Journal, 2014. https://doi.org/doi: 10.1155/2014/723595
- Wu, C. H. (1997). Artificial neural networks for molecular sequence analysis. Computers and Chemistry, 21(4), 237–256. https://doi.org/doi: 10.1016/S0097-8485(96)00038-1
- Bouaziz, A., Dartigues-Pallez, C., Da Costa Pereira, C., Precioso, F., & Lloret, P. (2014). Short text classification using semantic random forest. Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 8646 LNCS, 288–299. https://doi.org/doi: 10.1007/978-3-319-10160-6_26
- Pavlov, Y. L. (2019). Random forests. Random Forests, 1–122. https://doi.org/doi: 10.1201/9780367816377-11
- Bienvenido-huertas, D., Rubio-bellido, C., & P, J. L. (2020). Automation and optimization of in-situ assessment of wall thermal transmittance using a Random Forest algorithm, 168(August 2019). https://doi.org/doi: 10.1016/j.buildenv.2019.106479
- Liu, Z. P., Wu, L. Y., Wang, Y., Zhang, X. S., & Chen, L. (2010). Prediction of protein-RNA binding sites by a random forest method with combined features. Bioinformatics, 26(13), 1616–1622. https://doi.org/doi: 10.1093/bioinformatics/btq253
- Chou, K. C. (2001a). Prediction of signal peptides using scaled window. Peptides, 22(12), 1973–1979. https://doi.org/doi: 10.1016/S0196-9781(01)00540-X
- Feng, P., Ding, H., Chen, W., & Lin, H. (2013). Na ve Bayes Classifier with Feature Selection to Identify Phage Virion Proteins, 2013.
- Xu, Y., Shao, X., Wu, L., & Deng, N. (2013). iSNO-AAPair : incorporating amino acid pairwise coupling into PseAAC for predicting cysteine S -nitrosylation sites in proteins, 1–18. https://doi.org/doi: 10.7717/peerj.171
- Cai, Y. D., & Zhou, G. P. (2000). Prediction of protein structural classes by neural network. Biochimie, 82(8), 783–785. https://doi.org/doi: 10.1016/S0300-9084(00)01161-5
- Qiu, W. R., Jiang, S. Y., Xu, Z. C., Xiao, X., & Chou, K. C. (2017). iRNAm5C-PseDNC: Identifying RNA 5-methylcytosine sites by incorporating physical-chemical properties into pseudo dinucleotide composition. Oncotarget, 8(25), 41178–41188. https://doi.org/doi: 10.18632/oncotarget.17104
- Qiu, W. R., Sun, B. Q., Xiao, X., Xu, D., & Chou, K. C. (2017). iPhos-PseEvo: Identifying Human Phosphorylated Proteins by Incorporating Evolutionary Information into General PseAAC via Grey System Theory. Molecular Informatics, 36(5), 1–10. https://doi.org/doi: 10.1002/minf.201600010
- Chou, K.-C. (2015). Impacts of Bioinformatics to Medicinal Chemistry. Medicinal Chemistry, 11(3), 218–234. https://doi.org/doi: 10.2174/1573406411666141229162834
- Sigma, S. (2005). N Atural S P I, 1(2), 1–27.
- Cheng, F., Chen, X., Xiang, C., Liu, Z., & Wang, Y. (2020). Fluorescence-based high-throughput screening system for R - $$ -transaminase engineering and its substrate scope extension.
- Mussardo, G. (2019). No Title No Title. Statistical Field Theor, 53(9), 1689–1699. https://doi.org/doi: 10.1017/CBO9781107415324.004
- Wang, T., Zhao, B., & Foehr, E. D. (2019). Engineering a Long Lasting Tethered, Multimeric Human Growth Hormone Protein to Improve Pharmacokinetic Half-Life and Potency. Journal of Proteomics & Bioinformatics, 12(3), 56–60. https://doi.org/doi: 10.35248/0974-276x.19.12.497