$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
NOTE: All the codes used in this article can be found on the website https://github.com/YaoGeng-nmu/Analysis-of-TRPM4-based-on-TCGA-data-and-single-cell-data-in-BLCA/blob/main/code.
1. Preparation of transcriptomics data
- Preparation of the TCGA-BLCA dataset
- Download all the TCGA datasets from the website UCSC Xena (https://xenabrowser.net/datapages/)10. The datasets downloaded are pre-processed, eliminating the need for additional work such as gene annotation11.
- Select the TCGA dataset and then click on the TCGA-BLCA dataset. Then, go to the page of the TCGA-BLCA dataset, click on Gene Expression RNA-seq, and download the clinical data and gene expression profiles from the TCGA-BLCA dataset. In this dataset, ensure that there are 400 BLCA patients' mRNA expression data.
- To differentiate tumor samples from adjacent normal tissue samples, separate the samples with the endings 01 and 11 as samples with 01 indicate tumor samples, while the samples with 11 indicate cancer-adjacent tissues (normal samples).
- Preparation of GSE32894 and GSE31684
- Click on the GEO Website (https://www.ncbi.nlm.nih.gov) in order to download BLCA dataset GSE32894 and GSE3168412,13.
- After clicking on the Respective Entries in the upper-right corner, make sure to move to the pages for datasets GSE32894 and GSE31684.
- Next, download the respective mRNA data, clinical data, and survival data for GSE32894. Click on the http button of the GSE32894 dataset. Also, download the GPL6947 platform data using the platform button. Click on the Series Matrix File(s) button and download GSE32894 after clicking on the GSE32894_series_matrix.txt.gz button when downloading mRNA data, clinical data, and survival data for GSE32894.
- Download the respective mRNA data, clinical data, and survival data for GSE31684. Click on the http button of the GSE31684 dataset. Also, download the GPL570 platform data using the platform button. Click on the Series Matrix File(s) button and download GSE31684 after clicking on the GSE31684_series_matrix.txt.gz button when downloading mRNA data, clinical data, and survival data for GSE31684.
- In the filtering step, remove meaningless genes, such as those with an expression level of zero or close to zero. Employ the ComBat method, which is based on the sva package14, for batch correction.
2. Preparation of single-cell analysis
- Preparation of GSE145281
- To conduct research at the single-cell level, find the BLCA single-cell dataset, considering that transcriptome data alone is insufficient to determine which cell cluster expresses TRPM4 most significantly.
- Here, go to TISCH2 (http://tisch.compbio.cn/), an online website for single-cell analysis of tumors.
- According to experimental needs, select the Cancer Type needed, using BLCA as an example in this case.
- Click on the BLCA button, choosing the dataset called GSE145281.
3. Preparation of BLCA's molecular subtypes and response to therapeutic choices data
- Download the data for BLCA's molecular subtypes and response to therapeutic choices from the article referenced here15.
- Open the PubMed website (https://pubmed.ncbi.nlm.nih.gov/). Search for the article and click on Supplementary Tables. There is a total of 18 files in this compressed file package.
4. BLCA immune microenvironment analysis of TRPM4
- Heatmap plot of the expression of 133 immunomodulators in different TRPM4 groups
- According to the expression of TRPM4, divide all BLCA samples into the high-TRPM4 group and the low-TRPM4 group.
- Copy all the sample names and TRPM4 expression data separately, then sort them in descending order of TRPM4 expression, and categorize the first half of the samples into the High-TRPM4 group and the latter half into the Low-TRPM4 Group in the Group column according to TRPM4 expression. All TCGA-BLCA patients' IDs are listed in Supplementary Table 1.
- Install R packages required for the transcriptomic data analysis. Make slight modifications to the provided code, such as changing the file location, and then run the code. Carefully check if the gene IDs and the IDs of the immunomodulators are correctly matched, ensuring that there are no spaces or other inconspicuous factors after the gene IDs, as these are the primary causes of failed heat map generation.
- Stromal and immunological analysis of tumor samples using the ESTIMATE R package
- Make sure that all R packages required for the transcriptomic data analysis are installed. According to the code requirements, name the required data with the corresponding names, and then click the Run button.
NOTE: It is important to remember that this website also offers online analysis capabilities. While it is possible to obtain the desired results by visiting the website (https://bioinformatics.mdanderson.org/estimate/rpackage.html), this method is not suitable for analyzing data that has been self-collected.
- Click on the Disease button and choose the Bladder Urothelial Carcinoma option. Then, select the RNA-seq-V2 platform button. Download all samples' stromal score, immune score, and estimate score.
- Violin plot of the expression between different TRPM4 groups
- Install all R packages required for transcriptomic data analysis. All the R packages needed are listed in Supplementary Table 2. According to the code requirements, name the required data with the corresponding names, and then click the Run button.
- Make sure to carefully verify that the gene IDs are entered correctly, ensuring there are no spaces or other inconspicuous factors at the end of the gene IDs, as these are the primary causes of failures in the creation of violin plots.
- Triangle heatmap plot between different TRPM4 groups
- Install all R packages required for the transcriptomic data analysis. Make sure to prepare the expression data of several immune checkpoint genes, such as TIGIT and CD80. Open the TCGA-BLCA transcriptome data and then search through each entry individually, given that the number of immune checkpoint genes is not extensive.
- According to the code requirements, name the required data with the corresponding names, and then click the Run button.
- Correlation plot between different genes and TRPM4
- Install all R packages required for the transcriptomic data analysis. Fill in TRPM4 as the first gene and add the gene being researched as the second gene, which is CD3E in this case.
- After clicking the Run button, make sure to change CD3E to other genes that need to be analyzed, such as CTLA4.
5. Single-cell analysis of TRPM4 in the GSE145281 dataset
- BLCA GSE145281 dataset analysis using single-cell analysis
- Enter the website and click on BLCA. Single-cell analysis allows for a more in-depth exploration of which cell cluster exhibits high expression of TRPM4 compared to transcriptomics analysis. Make sure that the single-cell analysis data is collected from GSE130001 and check against the quality control standard that the cell number per dataset should be greater than 1000, the UMI count per cell should be greater than 1000, and the gene number per cell should be greater than 500. The differentially expressed genes of each cluster compared to all other cells are identified based on the log-transformed fold change (|logFC| >= 0.25), and the clustering resolution should be 0.5.
- Cell annotation for different cell clusters
- Make sure that all cell clusters are annotated. For example, the B cells can be annotated by CD19, MS4A1, CD27, IGHD, IGHM, TCL1A, and FCRL5. All the genes used for cell annotation can be found in Supplementary Table 3.
- Then, choose GSE145281 and download all results of TRPM4. Click on the Overall button and download the UMAP plot of GSE145281 of different clusters and cell types.
- Click on the Gene button and choose the single gene that is going to be analyzed. Click on the GSEA button and observe the functions of this single gene across various pathways, such as the KEGG pathway and Hallmark pathway16,17.
- Click on the CCI button and find the cell chat plot and cell chat bubble plot of the cluster highly expressing a single gene.
- Click on the TF enrichment button to see the enriched Transcript Factor for each cluster.
6. 101 Machine learning to construct a prognostic model relating to TRPM4
- Install all R packages required for transcriptomic data analysis. Considering that a single machine learning approach might not always be the optimal method for predicting prognosis, a prognosis model can be constructed using 101 machine learning techniques. Conduct a correlation analysis between all the genes in the TCGA-BLCA cohort and GSE32894 and GSE31684 cohorts and TRPM4.
- Based on the number of genes strongly associated with TRPM4, select all genes with a correlation greater than or equal to 0.6 or less than or equal to -0.6 with TRPM4. Select all genes strongly associated with TRPM4 in all cohorts.
- Combine all the data from GSE32894 and GSE31684 cohorts to form a validation set and use TCGA-BLCA as the training set.
- According to the code requirements, name the required data with the corresponding names, and then click the Run button.