August 21st, 2026
This protocol presents a reproducible workflow for analyzing spatial transcriptomics data, guiding users from public data acquisition and Seurat-based quality control through integration, spatial feature detection, cell-type deconvolution, region-of-interest annotation, and cell–cell communication analysis, with practical checkpoints that support transparent execution.
Hello, everyone. In this video, we will walk through a practical spatial transcriptomics data analysis pipeline from data acquisition and loading to basic exploration, and then finally, advanced analysis. Overall, the workflow consists of three main steps.
First, downloading the data, second, obtaining the analysis code, and third, running the pipeline to generate results. Step one, data acquisition and the directory structure preparation. First, obtain public spatial transcriptomic data sets.
Download the raw data archive. And extract the archive. Organize the files into a standardized directive structure.
First, create a main data directory, and then create a dedicated sub-directory for each sample. Transfer the following essential files for each sample into the respective sub-directory. And then create a spatial sub-folder within each sample directory.
Place the following files into the spatial sub-folder. Place the filtered feature PC metrics S1 file into the main sub-directory of the sample. Decompress the gzip files in the spatial folder.
Ensure the original file names remain exactly as required by the load 10X spatial function. Step two, software environment setup. Here we skip the installation of R language and the data spill begins from obtaining the analysis script in the GitHub repository.
Install the required R packages from Graham in the file conductor by executing the script setup R.Install the auto suit by executing the installation commands provided in the official document sheet. Navigate to the official installation URL to retrieve setup scripts. Initialize the required Python environment and system dependencies according to the setup page instructions.
Obtain the custom tool by navigating to the GitHub repository and download the source code. Navigate into the TOS directory and install the Python dependencies. Step three, spatial data loading and the quality control.
Read the spatial data into Seurat object. Use read 10X image to manually load the high resolution tissue image. Specifying the image there and the image name.
Use the load 10X spatial with image parameter set to the image object created in the previous step and create the Seurat object. Calculate quality control metrics. Compute the percentage of mitochondrial reads using percentage feature set with the pattern mt.
Visualize and interpret the data based on QC metrics. Generate violin plots of nCount_Spatial, nFeature_Spatial, and percent. mt using violin plot function.
Create spatial feature plots of these metrics using spatial feature plots. And identify spots outside the tissue area. Optional, apply filters to remove low quality spots.
After running the script, you can get these results, including QC metrics and the feature spatial plot. Step four, data pre-processing, integration, and clustering. Normalize in the pre-process individual samples.
Apply as a SC transform normalization to each sample separately with assay Spatial. Integrate multiple samples. Prepare the list of SCTransform normalized objects for integration.
Ensure each object has a RNA assay by copying the spatial assay. Use select integration features in the PREP SCT integration to identify shared variable features. Find the integration anchors using find integration anchors with normalization method SCT.
Integrate the data using IntegrateData. Perform dimensionality reduction in the clustering on the integrated assay. Run PCA on the integrated data using runPCA.
Determine the optimal number of the principle components for the downstream analysis by calculating the cumulative variance explained. Identify the elbow point program medically. Run UMap using the determined number of PCs.
Cluster the cells using FindNeighbors and FindClusters. Specify the determined PCs in the set resolution 0.5. Perform differential expression analysis between target groups using the FindWorkers function.
Identify spatially variable genes. For each original sample, run find spatially variable features using the Moran's I method on the SCT assay to compute spatial autocorrection. After running this script, you can get the elbow plot, the UMap plot, and the cluster plot, the cluster markers heat map, volcano plot, spatial feature differentially expressed genes, spatially reliable genes, and the colon layer markers.
And the colon layer markers dot plot, along with the markers in the spatial feature plot. Step five, single-cell reference data pre-processing. Read the single-cell RNA-seq count matrix using read 10X, and create a Seurat object.
Perform standard QC normalization and declustering. Calculate the percentage of mitochondrial reads in the filter cells. Normalize data using SC transform.
Set a vara.to. regress percent.mt. Long PCA UMap and cluster the cells using the dynamic PC selection method described in previous steps, then annotate cell types.
Circulate module scores for canonical cell type marker genes using AddModuleScore. Annotate clusters based on the module scores and known biology. Alternatively import pre-computed annotations from metadata.
After learning this script, you can get the QC metrics. They have upload the UMap by cluster, Umap by sample, cell type scores. Step six, reference guided deconvolution with SPOTlight.
First, prepare data for SPOTlight. Convert the annotated single-cell Seurat object and the spatial Seurat object to single-cell experiment object. Log normalize the single-cell data using LogMoreCounts.
Then run SPOTlight deconvolution. First, identify high-variable genes on the single-cell data using ModelGeneVar. Compute cell type markers using score markers and filter for high quality markers.
Downsample the single-cell reference for each cell type to a manageable number to reduce computational time. Execute the deconvolution using the SPOTlight function, providing the single-cell reference, the spatial data, marker list, and HVGs, then we can visualize and export the results. You can get the deconvolution result like this, shown as a scatter pipe route.
Step seven, unsupervised deconvolution with Stdeconvolve. First, prepare the spatial data. Extract the row count metrics from the spatial Seurat object using GetAssayData with the slot counts.
Remove low quality spots and genes using clean counts from STdeconvolve. Identify the latent cell types for the filters that corpus four genes expressed in a minimum fraction of spots using restrict strict LDA by reach that allocation model across a range of potential topic numbers using fitLDA. Select the optimal model based on minimum complexity using optimal model with opt min.
Analyze and visualize the results. Extract the serotype proportion, Theta, and the gene profiles Beta for optimal model using getBetaTheta. To add in the biological interpretation of the corrosion topics, import region of interest annotations generated by the select spatial spot tool.
Use these annotations as group's parameter in the with all topics function. To your project the deconvolved cell type proportions back onto the spatial coordinates and color code the spots by their ROI. After running the script, you will get a result like this, a scale's pipe load, like those produced by the SPOTlight.
Step eight, spatial cell-cell communication using Giotto. First, convert the Seurat object to a Giotto object. Use the createGiottoObject function, providing low-count metrics and the spatial coordinates.
Preprocess the Giotto object and add deconvolution result. Normalize the data using normalized Giotto, add as a serotype annotations. Choose cell metadata using addCellmetadata.
Create spatial network using createSpatialNetwork. Load a ligand receptor database into our environment. Run explore CellCellcom to identify significant liquid receptor interactions between cell types that are in spatial proximity.
You can get the cell-cell communication dot plot like this. Step nine is optional. Interactive spot selection with SelectSpatialSpot.
Prepare the data for interactive tool using the script six. Extract the spatial coordinates from the Seurat object using GetTissueCoordinates. Format and export the data.
Export the formative data frame to a CSV file. Then perform region of interest analysis. Launch the custom.
Select the spatial spots dash applications and load the CSV file. Interactively select the spots based on the spatial location. Then export the list of select spots and their assigned group, labeled as a new CSV file.
We just check what we can get after running the script one-by-one. All the results are saved in the results five-fold. And as you can see, we can get the QC metrics, and divide.
And you can check the spatial feature plot here. Moreover, the deconvolution process are done in the way of supervised and unsupervised both. The SPOTlight results are here.
As you can see, all the spots contains the proportion information. The unsupervised deconvolution results by the STdeconvolve are also generated here. And this is Seurat clusters result of the spot in the spatial data.
You can also visualize them in the spatial team plot, like this. And moreover, you can also get the spot-spot communication result by the Giotto. This entire workflow is 100%open source.
Spotting analysis from the expression metrics to advanced spatial modeling. All the steps runs on a machine with 16GB RAM. The code base is modular, one script for each task.
Note that this pipeline does not cover upstream processing from the FASTQ price. Focus is on the 2D spatial data, and currently includes only Visium's downloaders. That's all, thank you for watching.
View the full transcript and gain access to thousands of scientific videos
This article presents a comprehensive computational workflow for analyzing spatial transcriptomics (ST) datasets using R. The protocol addresses common challenges in ST analysis, such as data import, quality control, integration, deconvolution, spatial statistics, and visualization, by providing a streamlined, script-based approach. The workflow is adaptable to standard array-based ST datasets and emphasizes reproducibility and parameter transparency.
Spatial transcriptomics data analysis is pivotal for understanding tissue architecture and microenvironmental biology in early discovery and translational research. This workflow enables biopharma teams to integrate, deconvolve, and interpret spatial gene expression data with reproducibility and parameter transparency. By standardizing computational steps, it supports robust target validation and risk-adjusted portfolio decisions.
This workflow bridges early discovery, lead identification, and translational research by providing a reproducible computational backbone for spatial transcriptomics analysis.