2018年10月11日
Existing algorithms generate one solution for a biomarker detection dataset. This protocol demonstrates the existence of multiple similarly effective solutions and presents a user-friendly software to help biomedical researchers investigate their datasets for the proposed challenge. Computer scientists may also provide this feature in their biomarker detection algorithms.
This measure can help answer key questions in the biomedical detection field about generating multiple solutions. The main advantage of this technique is that it provides a user-friendly graphical user interface for assisting biomedical researchers in the detection of multiple feature subsides. Begin by loading the data Matrix and class labels into the software.
Click load data matrix to select the user specified data metrics file and load class labels to select the corresponding class label file. To determine the class labels in the number of top-ranked features, select the names of the positive and negative classes in the appropriate drop down boxes and select 10 as the number of top-ranked features in the top X drop down box for a comprehensive screen of the feature subset. To tune the system parameters for different performances, select the performance measurement accuracy as the accuracy balanced accuracy drop down box for the selected extreme learning machine classifier.
Then, select a cut-off value of 0.7 for the specified performance measurement in the performance cutoff input box. To run the pipeline, click analyze and select 0.7 as the default value of the performance measurement cut off. And, 10 as the default number of the best feature subsets.
Then collect and interpret the features detected by the software. To generate a 3D scatter plot of the top 10 features of the subsets with the best classification performances detected by the software, click analyze and sort the three features in a feature subset in ascending order of their ranks, using the ranks of the three features as the F1, F2, and F3 axes. Change the performance cut off value to 0.7 and click analyze to generate a 3D scatter plot of the feature subsets with a greater than or equal to performance cut off performance measurement value.
Then click 3D tuning to open a new window for manual tuning of the viewing angles of the 3D scatter plot and reduce to reduce the redundancy of the detected feature subsets. To annotate a gene in both the DNA and protein sequence levels, open the David database web page and click on the gene ID conversion link to input the feature IDs of the first biomarker subset of the prepared data set. Click the gene list link and click submit list to retrieve the annotations of interest, and show Gene list to obtain the list of Gene symbols.
Next, open the GeneCards database web page and enter the name of the gene of interest into the database query input box to find the annotations of this gene. Open the Online Mendelian Inheritance in Man database and search for the gene to find the annotations of this gene from the database. To annotate the encoded proteins, open the UniProt knowledge base database page and search for the annotations of the gene from this database.
Open the group based prediction system, or GPS web server, and retrieve the protein sequence encoded by the biomarker gene from the UniProt knowledge base database and use the online GPS tool to predict the proteins post-transitional modification residues. To annotate the protein-protein interactions and there enriched functional modules, open the string web server page and use the string database to search the lift for the genes of interest to find their orchestrated properties. To export the detected biomarker subsets for further analysis, click export the table and select the appropriate text format for saving the files.
Then, export the visualization plots as individual image files, clicking save under each plot and selecting the appropriate image format for saving each file. In this representative experiment, two data sets were formatted as CSV files and loaded into the software as demonstrated. In the first data set, 128 samples with 12, 625 features and individual class labels were loaded with the final data Matrix containing 95 negative samples and 33 positive samples.
Similar operations were also conducted for the second difficult data set. Searching for a user-specific keyword in the feature names reveals a histogram of the features for each data set. After executing the pipeline algorithm for each data set, 120 qualified biomarker subsets were detected for the easy to discriminate data set, with 57 triplet biomarker subsets demonstrating a 100%accuracy.
Only 76 biomarker subsets where detected for the difficult data set, however. And, with a lower biomarkers subset accuracy suggesting that biomarkers are phenotype specific, another major challenge in biomarker detection. While using this procedure, it's important to remember that a future selection problem has multiple solutions.
Read the SIM best of performance. After its development, this technique paved the way for biomedical researchers to explore biomedical detection with multiple solutions.
查看完整文字稿并访问数千部科学视频
本方案展示了生物标志物检测数据集存在多种有效的解决方案。它提供了一个用户友好的软件界面,以协助生物医学研究人员开展研究工作。
生物标志物检测通常仅产生一个经过优化的解决方案,可能忽略了其他具有相似预测性能的替代子集。本方案表明,多个生物标志物子集均可实现同样有效的二分类,有助于生成更可靠的假设,并减少对单一且可能存在偏倚的标志物的依赖。通过识别多种有效的解决方案,研发团队可在发现早期获得更充分的靶点验证信心和机制层面的风险评估。
该方法可整合到早期发现工作流程中,通过定量检测和可视化生物标志物亚群,支持假设验证和先导化合物的鉴定。