方法文章

TRPM4为例的TCGA数据与单细胞数据联合分析

DOI:

10.3791/69304

2025年12月5日

* These authors contributed equally

本文内容

摘要

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

本文介绍了一种基于转录组分析和单细胞分析,并结合101种机器学习算法,系统研究单个基因在膀胱癌(BLCA)中作用的实验方案,旨在为该单个基因构建预后预测模型。

摘要

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

本文介绍了一种基于公共转录组数据集和单细胞数据集的分析方法,可用于全面描述单个基因在肿瘤中的作用,包括塑造肿瘤免疫微环境、决定肿瘤分子亚型以及预测肿瘤患者的预后。同时,引入单基因数据不仅能够避免单一转录组分析所带来的随机性和异质性问题,还能深入探究该基因在哪些特定细胞簇中表达,并进一步研究该基因在信号通路中所发挥的功能。考虑到许多研究人员可能不熟悉单细胞分析,本文介绍了一个名为 TISCH2 的在线网站(http://tisch.compbio.cn/),以帮助用户完成单细胞数据分析。此外,101 种机器学习方法的应用在构建最精确的预后模型中发挥了不可或缺的作用。综上所述,这种整合了生物信息学分析、机器学习和单细胞分析的单基因综合分析方法,将在研究单个基因在肿瘤进展中的功能以及基因在通路中作用的研究中发挥不可替代且至关重要的作用。

引言

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

众所周知,膀胱癌(BLCA)是全球最具侵袭性和转移性的恶性肿瘤之一,而目前膀胱癌的生物标志物仍存在诸如准确性不足等问题1。为了明确BLCA患者的预后并预测其临床结局,寻找膀胱癌的生物标志物并建立预后模型具有重要意义。尽管人们已开发出一些用于生物标志物研究的方法,但目前大多数方法仍局限于转录组学层面,这不可避免地导致样本异质性2。此外,仅研究转录组数据往往难以探究不同细胞亚群的作用,以及基因在不同通路中的单细胞水平功能,从而使以往的研究精确性和有效性受限3。考虑到单细胞分析对初学者可能存在挑战,已有在线平台被推出以支持单细胞分析,帮助用户快速掌握相关技能4。最后但同样重要的是,即使通过转录组分析和单细胞分析发现某个单基因可作为合适的生物标志物,也不能保证该标志物在所有队列中均适用,因此有必要构建与该生物标志物相关的预后模型,以增强结论的普适性5。101种机器学习算法是指结合10种不同的机器学习算法构建101个预后模型,旨在筛选出最优的预后模型。由于纳入了随机森林、XGBoost和SVM等在处理复杂数据方面能力更强且稳定性更高的算法,该组合算法表现出显著的稳定性。为了使该预后模型更准确且与生物标志物更相关,首先对所有基因与该生物标志物进行相关性分析,然后根据需求筛选约20至30个基因,再利用101种机器学习算法构建预后模型,随后通过一系列分析最终确定结果。

与过去基于简单转录组学分析预测的生物标志物相比6,单细胞转录组学分析能够无偏倚地将组织或样本分解为其基本的细胞组成。它可清晰区分不同类型的细胞(如T细胞、B细胞和巨噬细胞),发现新的细胞亚群(如耗竭性T细胞和调节性T细胞),并捕捉连续的动态过程(如细胞分化轨迹)4。这类似于将水果和牛奶分别归类,从而清晰地展示每种成分及其数量。与单队列Cox模型相比7,采用101种机器学习方法构建的预后模型也更加准确且科学合理,因为它能够从数据中自动学习变量之间复杂的非线性相互作用。例如,某种特定基因突变的影响可能仅在特定年龄和肿瘤大小的患者中才具有显著性。随机森林和神经网络等机器学习模型能够自动捕捉这些高阶交互作用,而无需人工预先设定8。该方法的关键适用条件包括:用于分析的数据集必须包含绝大多数基因,基因总数不宜过低,而用于构建预后模型的基因数量也不宜过高,通常维持在20至30个左右为最佳。单细胞数据集的质量控制标准应满足:细胞数量大于1000,每个细胞的UMI数大于1000,每个细胞检测到的基因数大于5009

本文以TRPM4在膀胱尿路上皮癌(BLCA)中的作用为例,系统性地提供了一种从公共转录组数据集和单细胞数据集中鉴定新型生物标志物的逐步研究方法。通过整合多种研究手段及多样化数据集——包括癌症基因组图谱-膀胱癌(TCGA-BLCA)数据集、单细胞数据集GSE145281,以及BLCA数据集GSE32894和GSE31684——从多个角度提升肿瘤学领域生物标志物的精确性与适用性,并推动TRPM4在BLCA中的深入研究。TCGA-BLCA数据集来源于加州大学圣克鲁兹分校Xena(UCSC Xena)网站;GSE32894和GSE31684数据集从基因表达综合数据库(GEO)下载;单细胞数据GSE145281来源于肿瘤免疫单细胞中心2(TISCH2)。此外,BLCA分子亚型及治疗反应相关数据提取自相关研究论文的补充电子表格文件。随后开展生物信息学分析、单细胞分析及机器学习,构建了一套整合性的单基因研究方法体系。

访问受限。请登录或开始试用以查看此内容。

方案

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

NOTE: All the codes used in this article can be found on the website https://github.com/YaoGeng-nmu/Analysis-of-TRPM4-based-on-TCGA-data-and-single-cell-data-in-BLCA/blob/main/code.

1. Preparation of transcriptomics data

  1. Preparation of the TCGA-BLCA dataset
    1. Download all the TCGA datasets from the website UCSC Xena (https://xenabrowser.net/datapages/)10. The datasets downloaded are pre-processed, eliminating the need for additional work such as gene annotation11.
    2. Select the TCGA dataset and then click on the TCGA-BLCA dataset. Then, go to the page of the TCGA-BLCA dataset, click on Gene Expression RNA-seq, and download the clinical data and gene expression profiles from the TCGA-BLCA dataset. In this dataset, ensure that there are 400 BLCA patients' mRNA expression data.
    3. To differentiate tumor samples from adjacent normal tissue samples, separate the samples with the endings 01 and 11 as samples with 01 indicate tumor samples, while the samples with 11 indicate cancer-adjacent tissues (normal samples).
  2. Preparation of GSE32894 and GSE31684
    1. Click on the GEO Website (https://www.ncbi.nlm.nih.gov) in order to download BLCA dataset GSE32894 and GSE3168412,13.
    2. After clicking on the Respective Entries in the upper-right corner, make sure to move to the pages for datasets GSE32894 and GSE31684.
    3. Next, download the respective mRNA data, clinical data, and survival data for GSE32894. Click on the http button of the GSE32894 dataset. Also, download the GPL6947 platform data using the platform button. Click on the Series Matrix File(s) button and download GSE32894 after clicking on the GSE32894_series_matrix.txt.gz button when downloading mRNA data, clinical data, and survival data for GSE32894.
    4. Download the respective mRNA data, clinical data, and survival data for GSE31684. Click on the http button of the GSE31684 dataset. Also, download the GPL570 platform data using the platform button. Click on the Series Matrix File(s) button and download GSE31684 after clicking on the GSE31684_series_matrix.txt.gz button when downloading mRNA data, clinical data, and survival data for GSE31684.
    5. In the filtering step, remove meaningless genes, such as those with an expression level of zero or close to zero. Employ the ComBat method, which is based on the sva package14, for batch correction.

2. Preparation of single-cell analysis

  1. Preparation of GSE145281
    1. To conduct research at the single-cell level, find the BLCA single-cell dataset, considering that transcriptome data alone is insufficient to determine which cell cluster expresses TRPM4 most significantly.
    2. Here, go to TISCH2 (http://tisch.compbio.cn/), an online website for single-cell analysis of tumors. 
    3. According to experimental needs, select the Cancer Type needed, using BLCA as an example in this case.
    4. Click on the BLCA button, choosing the dataset called GSE145281.

3. Preparation of BLCA's molecular subtypes and response to therapeutic choices data

  1. Download the data for BLCA's molecular subtypes and response to therapeutic choices from the article referenced here15.
  2. Open the PubMed website (https://pubmed.ncbi.nlm.nih.gov/). Search for the article and click on Supplementary Tables. There is a total of 18 files in this compressed file package.

4. BLCA immune microenvironment analysis of TRPM4

  1. Heatmap plot of the expression of 133 immunomodulators in different TRPM4 groups
    1. According to the expression of TRPM4, divide all BLCA samples into the high-TRPM4 group and the low-TRPM4 group.
    2. Copy all the sample names and TRPM4 expression data separately, then sort them in descending order of TRPM4 expression, and categorize the first half of the samples into the High-TRPM4 group and the latter half into the Low-TRPM4 Group in the Group column according to TRPM4 expression. All TCGA-BLCA patients' IDs are listed in Supplementary Table 1.
    3. Install R packages required for the transcriptomic data analysis. Make slight modifications to the provided code, such as changing the file location, and then run the code. Carefully check if the gene IDs and the IDs of the immunomodulators are correctly matched, ensuring that there are no spaces or other inconspicuous factors after the gene IDs, as these are the primary causes of failed heat map generation.
  2. Stromal and immunological analysis of tumor samples using the ESTIMATE R package
    1. Make sure that all R packages required for the transcriptomic data analysis are installed. According to the code requirements, name the required data with the corresponding names, and then click the Run button.
      ​NOTE: It is important to remember that this website also offers online analysis capabilities. While it is possible to obtain the desired results by visiting the website (https://bioinformatics.mdanderson.org/estimate/rpackage.html), this method is not suitable for analyzing data that has been self-collected.
    2. Click on the Disease button and choose the Bladder Urothelial Carcinoma option. Then, select the RNA-seq-V2 platform button. Download all samples' stromal score, immune score, and estimate score.
  3. Violin plot of the expression between different TRPM4 groups
    1. Install all R packages required for transcriptomic data analysis. All the R packages needed are listed in Supplementary Table 2. According to the code requirements, name the required data with the corresponding names, and then click the Run button.
    2. Make sure to carefully verify that the gene IDs are entered correctly, ensuring there are no spaces or other inconspicuous factors at the end of the gene IDs, as these are the primary causes of failures in the creation of violin plots.
  4. Triangle heatmap plot between different TRPM4 groups
    1. Install all R packages required for the transcriptomic data analysis. Make sure to prepare the expression data of several immune checkpoint genes, such as TIGIT and CD80. Open the TCGA-BLCA transcriptome data and then search through each entry individually, given that the number of immune checkpoint genes is not extensive.
    2. According to the code requirements, name the required data with the corresponding names, and then click the Run button.
  5. Correlation plot between different genes and TRPM4
    1. Install all R packages required for the transcriptomic data analysis. Fill in TRPM4 as the first gene and add the gene being researched as the second gene, which is CD3E in this case.
    2. After clicking the Run button, make sure to change CD3E to other genes that need to be analyzed, such as CTLA4.

5. Single-cell analysis of TRPM4 in the GSE145281 dataset

  1. BLCA GSE145281 dataset analysis using single-cell analysis
    1. Enter the website and click on BLCA. Single-cell analysis allows for a more in-depth exploration of which cell cluster exhibits high expression of TRPM4 compared to transcriptomics analysis. Make sure that the single-cell analysis data is collected from GSE130001 and check against the quality control standard that the cell number per dataset should be greater than 1000, the UMI count per cell should be greater than 1000, and the gene number per cell should be greater than 500. The differentially expressed genes of each cluster compared to all other cells are identified based on the log-transformed fold change (|logFC| >= 0.25), and the clustering resolution should be 0.5.
  2. Cell annotation for different cell clusters
    1. Make sure that all cell clusters are annotated. For example, the B cells can be annotated by CD19, MS4A1, CD27, IGHD, IGHM, TCL1A, and FCRL5. All the genes used for cell annotation can be found in Supplementary Table 3.
    2. Then, choose GSE145281 and download all results of TRPM4. Click on the Overall button and download the UMAP plot of GSE145281 of different clusters and cell types.
    3. Click on the Gene button and choose the single gene that is going to be analyzed. Click on the GSEA button and observe the functions of this single gene across various pathways, such as the KEGG pathway and Hallmark pathway16,17.
    4. Click on the CCI button and find the cell chat plot and cell chat bubble plot of the cluster highly expressing a single gene.
    5. Click on the TF enrichment button to see the enriched Transcript Factor for each cluster.

6. 101 Machine learning to construct a prognostic model relating to TRPM4

  1. Install all R packages required for transcriptomic data analysis. Considering that a single machine learning approach might not always be the optimal method for predicting prognosis, a prognosis model can be constructed using 101 machine learning techniques. Conduct a correlation analysis between all the genes in the TCGA-BLCA cohort and GSE32894 and GSE31684 cohorts and TRPM4.
  2. Based on the number of genes strongly associated with TRPM4, select all genes with a correlation greater than or equal to 0.6 or less than or equal to -0.6 with TRPM4. Select all genes strongly associated with TRPM4 in all cohorts.
  3. Combine all the data from GSE32894 and GSE31684 cohorts to form a validation set and use TCGA-BLCA as the training set.
  4. According to the code requirements, name the required data with the corresponding names, and then click the Run button.

访问受限。请登录或开始试用以查看此内容。

结果

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

众所周知,利用转录组学分析免疫微环境在单个基因水平上的核心优势在于,它能够建立单个基因与免疫细胞及分子在基因表达水平上动态相互作用的直接关联,从而直接观察免疫特征,并识别在单个基因高/低表达组之间浸润的免疫细胞比例差异,或免疫检查点分子表达变化的情况18。生物信息学分析的代表性结果包括与免疫微环境和免疫分子分析相关的结果,例如小提琴图等。单细胞分析的代表性结果包括每个细胞簇中目标基因的表达水平,以及每个细胞簇的通路热图分析。101 种机器学习模型的代表性结果则是用于构建预后模型的算法热图。

首先,热图显示有133种免疫调节分子在低表达组中表达水平更高TRPM4 组(图1A)。根据 ESTIMATE R 软件包的计算结果,很明显,高风险组中的基质评分(stromalscore)、免疫评分(immunescore)和 ESTIMATE 评分(estimatescore)均较低TRPM4 组,从而表明在高表达样本中 TRPM4 ...

访问受限。请登录或开始试用以查看此内容。

讨论

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

在以往的研究中,单个基因的生物信息学分析常遇到浅显、不准确及适用性有限等问题19。此外,以往关于单个基因的研究通常局限于转录组数据。而当仅关注转录组数据时,可能会出现样本异质性和随机性等问题,导致难以识别普遍规律20。因此,为解决上述问题,已引入单细胞数据以更深入地理解单个基因在通路水平上的关键作用。本文的关键步骤是引入了一个名为 TISCH2(http://tisch.compbio.cn/)的在线网站用于单细胞分析4。质量控制所采用的标准为:每个数据集的细胞数量超过 1000,每个细胞的 UMI 数量超过 1000,每个细胞检测到的基因数量超过 500。此外,101 种机器学习方法也是关键步骤之一,用于寻找最优预后模型,以预测所有患者的预后情况和临床结局21。最后但同样重要的是,务必应用多种生物信息学分析图来描述单个基因的研究结果。

在绘制图表和进行分析的过程中,难免会遇到一些问题。如果发...

访问受限。请登录或开始试用以查看此内容。

披露

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

作者声明不存在利益冲突。

致谢

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

我们感谢BioBean生物信息学联盟开发了智能分析框架(可在 http://www.sxdyc.com/ 获取)。他们创新的计算基础设施通过精准分析和自动化数据解读模块,显著加快了研究工作流程。

访问受限。请登录或开始试用以查看此内容。

材料

本文使用的材料清单
姓名公司目录编号评论
R 4.3.3

参考文献

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,
  1. Zhang, C., et al. Identification of multicohort-based predictive signature for NMIBC recurrence reveals SDCBP as a novel oncogene in bladder cancer. Ann Med. 57 (1), 2458211(2025).
  2. Han, M. H., et al. Plasma GFAP and Amyloid Pathology Predict Cognitive Response to Multidomain Interventions in MCI. Aging Dis. , (2025).
  3. Xie, S., et al. Towards Precision Aging Biology: Single-Cell Multi-Omics and Advanced AI-Driven Strategies. Aging Dis. , (2025).
  4. Han, Y., et al. TISCH2: expanded datasets and new tools for single-cell transcriptome analyses of the tumor microenvironment. Nucleic Acids Res. 51 (D1), D1425-D1431 (2023).
  5. Yao, Y., et al. Advances in prognostic models for osteosarcoma risk. Heliyon. 10 (7), e28493(2024).
  6. Ding, X., et al. Glutamine metabolism reprogramming promotes bladder cancer progression via PYCR1: a multi-omics and functional validation study. J Transl Med. 23 (1), 1277(2025).
  7. Zhu, T., et al. Methylparaben and propylparaben promote bladder cancer invasion via MMP2 and PPARG modulation. Ecotoxicol Environ Saf. 306, 119383(2025).
  8. Xie, J. H., et al. Deciphering cutaneous melanoma prognosis through LDL metabolism: Single-cell transcriptomics analysis via 101 machine learning algorithms. Exp Dermatol. 33 (4), e15070(2024).
  9. Sun, D., et al. TISCH: a comprehensive web resource enabling interactive single-cell transcriptome visualization of tumor microenvironment. Nucleic Acids Res. 49 (D1), D1420-D1430 (2021).
  10. Cao, X., et al. D-Mannose Upregulates Testin via the NF-κB Pathway to Inhibit Breast Cancer Proliferation. J Biochem Mol Toxicol. 39 (8), e70398(2025).
  11. Goldman, M. J., et al. Visualizing and interpreting cancer genomics data via the Xena platform. Nat Biotechnol. 38 (6), 675-678 (2020).
  12. Yu, Q., et al. GREM1 may be a biological indicator and potential target of bladder cancer. Sci Rep. 14 (1), 23280(2024).
  13. Wu, Q., et al. Membrane palmitoylated protein MPP1 inhibits immune escape by regulating the USP12/ CCL5 axis in urothelial carcinoma. Int Immunopharmacol. 146, 113802(2025).
  14. Johnson, W. E., Li, C., Rabinovic, A. Adjusting batch effects in microarray expression data using empirical Bayes methods. Biostatistics. 8 (1), 118-127 (2007).
  15. Hu, J., et al. Siglec15 shapes a non-inflamed tumor microenvironment and predicts the molecular subtype in bladder cancer. Theranostics. 11 (7), 3089-3108 (2021).
  16. Kanehisa, M., Sato, Y., Morishima, K. BlastKOALA and GhostKOALA: KEGG Tools for Functional Characterization of Genome and Metagenome Sequences. J Mol Biol. 428 (4), 726-731 (2016).
  17. Munkley, J., et al. Hallmarks of glycosylation in cancer. Oncotarget. 7 (23), 35478-35489 (2016).
  18. Da, Y., et al. A high stroma-tumor ratio is associated with an immunosuppressive tumor microenvironment and a poor prognosis in bladder cancer. Front Oncol. 15, 1604609(2025).
  19. Chen, C., et al. Bioinformatics Methods for Mass Spectrometry-Based Proteomics Data Analysis. Int J Mol Sci. 21 (8), 2873(2020).
  20. Jonauskaite, D., et al. Universal Patterns in Color-Emotion Associations Are Further Shaped by Linguistic and Geographic Proximity. Psychol Sci. 31 (10), 1245-1260 (2020).
  21. Zhu, W., et al. Integrated machine learning identifies epithelial cell marker genes for improving outcomes and immunotherapy in prostate cancer. J Transl Med. 21 (1), 782(2023).
  22. Zhao, J., et al. Bioinformatics prediction and experimental verification of key biomarkers for diabetic kidney disease based on transcriptome sequencing in mice. PeerJ. 10, e13932(2022).

访问受限。请登录或开始试用以查看此内容。

重印与许可

申请许可以重复使用本 JoVE 文章的文本或图表

申请许可

标签

TRPM4 TISCH2

相关文章