您好,欢迎访问湖南省农业科学院 机构知识库!

A pipeline for improved QSAR analysis of peptides: physiochemical property parameter selection via BMSF, near-neighbor sample selection via semivariogram, and weighted SVR regression and prediction

文献类型: 外文期刊

作者: Dai, Zhijun 2 ; Wang, Lifeng 1 ; Chen, Yuan 1 ; Wang, Haiyan 3 ; Bai, Lianyang 2 ; Yuan, Zheming 1 ;

作者机构: 1.Hunan Agr Univ, Hunan Prov Key Lab Germplasm Innovat & Utilizat C, Changsha, Hunan, Peoples R China

2.Hunan Agr Univ, Hunan Prov Key Lab Biol & Control Plant Dis & Ins, Changsha, Hunan, Peoples R China

3.Kansas State Univ, Dept Stat, Manhattan, KS 66506 USA

4.Hunan Acad Agr Sci, Changsha, Hunan, Peoples R China

关键词: Peptides;Quantitative structure-activity regression;Feature selection;Semivariogram;Support vector regression

期刊名称:AMINO ACIDS ( 影响因子:3.52; 五年影响因子:3.6 )

ISSN:

年卷期:

页码:

收录情况: SCI

摘要: In this paper, we present a pipeline to perform improved QSAR analysis of peptides. The modeling involves a double selection procedure that first performs feature selection and then conducts sample selection before the final regression analysis. Five hundred and thirty-one physicochemical property parameters of amino acids were used as descriptors to characterize the structure of peptides. These high-dimensional descriptors then go through a feature selection process given by the binary matrix shuffling filter (BMSF) to obtain a set of important low-dimensional features. Each descriptor that passes the BMSF filtering also receives a weight defined through its contribution to reduce the estimation error. These selected features served as the predictors for subsequent sample selection and modeling. Based on the weighted Euclidean distances between samples, a common range was determined with high-dimensional semivariogram and then used as a threshold to select the near-neighbor samples from the training set. For each sample to be predicted, the QSAR model was established using SVR with the weighted, selected features based on the exclusive set of near-neighbor training samples. Prediction was conducted for each test sample accordingly. The performances of this pipeline are tested with the QSAR analysis of angiotensin-converting enzyme inhibitors and HLA-A*0201 data sets. Improved prediction accuracy was obtained in both applications. This pipeline can optimize the QSAR modeling from both the feature selection and sample selection perspectives. This leads to improved accuracy over single selection methods. We expect this pipeline to have extensive application prospect in the field of regression prediction.

  • 相关文献

[1]QSAR modeling of E. coli promoters with parameters selected by binary matrix shuffling filter. Wang, Li-Feng,Dai, Zhi-Jun,Yuan, Zhe-Ming,Wang, Kai,Wang, Li-Feng,Dai, Zhi-Jun,Yuan, Zhe-Ming,Bai, Lian-Yang.

[2]Quantitative Sequence- Activity Model Analysis of Oligopeptides Coupling an Improved High- Dimension Feature Selection Method with Support Vector Regression. Dai, Zhijun,Zhang, Hongyan,Yuan, Zheming,Wang, Lifeng,Dai, Zhijun,Zhang, Hongyan,Bai, Lianyang,Yuan, Zheming,Bai, Lianyang. 2014

作者其他论文 更多>>