$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Workflow implementation and data integration illustrate major tissue features
The computational workflow was applied to mouse colon spatial transcriptomics data to illustrate the expected outputs across the analytical stages. As depicted in the workflow schematic (Figure 1), the pipeline began with data acquisition and quality control, where spatial feature plots delineated tissue boundaries (Figure 2A,B). Seurat's anchor-based integration workflow was then used to reduce technical batch effects while preserving interpretable biological variation. UMAP visualizations showed sample alignment and spatial clustering patterns after integration (Figure 2C,D). The quantitative, dynamic selection of principal components (PCs) based on cumulative variance was implemented to guide dimensional reduction and downstream clustering (see Supplementary Figure 1). Marker gene heatmap analysis showed distinct transcriptional profiles underlying the spatial clusters (Figure 2E).
To assess whether the computational clusters were consistent with the known anatomical architecture of colon histology, the expression profiles of canonical layer-specific marker genes were evaluated. The mucosal epithelium layer showed expression of epithelial cell markers, including Epcam and Krt8, alongside the goblet cell marker Muc2. Mesenchymal and stromal markers such as Col1a1 and Vim marked lamina propria and submucosal regions, whereas the outer muscularis propria layer was indicated by smooth muscle structural genes such as Acta2 and Tagln. The spatial restriction of these lineage-associated markers supports the interpretation that the integration and clustering workflow preserved major histological laminations of the colonic tissue along the mucosal-to-muscularis axis (see Supplementary Figure 2).
Following cluster validation, downstream differential expression analysis was performed to identify differentially expressed genes (DEGs) between experimental conditions (Figure 2F,G). Furthermore, spatially variable genes were identified using Moran's I statistic, highlighting genes with significant non-random spatial distribution across the tissue (Figure 2H).
Cellular deconvolution and spatial interaction networks reveal tissue microorganization
Processing of the single-cell RNA-seq reference data yielded annotations supported by QC filtering (Figure 3A), unsupervised clustering (Figure 3B), marker gene validation (Figure 3C), and concordance with independent annotations (Figure 3D). The cellular composition (Figure 3E) informed the downsampling strategy for deconvolution. SPOTlight estimated reference-guided cell-type proportions across spatial spots (Figure 4A,B), whereas STdeconvolve provided an unsupervised topic-modeling view of spatial cellular patterns (Figure 5B). The custom Select Spatial Spots tool provided histological context for these patterns (Figure 5A). Finally, using deconvoluted cell-type assignments, spatial communication analysis identified ligand-receptor interactions between spatially proximal cell-type groups (Figure 6A,B).
Troubleshooting observations from protocol optimization
During protocol optimization, several issues were identified that informed practical checkpoints. Suboptimal deconvolution results occurred when single-cell references were poorly matched to the tissue context, indicating the need to use tissue- and species-matched scRNA-seq data when available. Initial clustering attempts with default parameters did not always resolve expected biological structures; inspecting the PC selection, clustering resolution, and marker gene coherence helped identify spatially interpretable domains aligned with tissue anatomy. These observations provide practical examples of how users can diagnose common analytical problems during execution of the workflow.

Figure 1: Workflow for integrated spatial transcriptomics analysis. Schematic representation of the analytical pipeline, from data acquisition and pre-processing to advanced spatial analyses. Key steps include: (1) Data loading, quality control, and multi-sample integration using Seurat; (2) Spatial clustering and detection of spatially variable genes; (3) Cell-type deconvolution via reference-based (SPOTlight) and unsupervised (STdeconvolve) methods; (4) Spatial cell-cell communication analysis with Giotto and interactive region-of-interest selection using a custom tool, Select Spatial Spots. Results from all modules are synthesized to derive biological insights into tissue architecture and cellular microenvironment. Please click here to view a larger version of this figure.

Figure 2: Data integration, clustering, and differential expression analysis. (A,B) Quality control metrics for spatial samples A1 and B1, showing distributions of gene counts, UMI counts, and mitochondrial gene percentages. (C) UMAP visualization of integrated spatial transcriptomics data colored by sample origin (left) and clustering identity (right). (D) Spatial projection of cluster identities onto tissue sections. (E) Heatmap of top marker genes for each spatial cluster. (F) Volcano plot displaying differentially expressed genes between conditions A1_colon_d0 and B1_colon_d14. (G) Spatial expression patterns of representative differentially expressed genes across tissue sections. (H) Spatial expression maps of top spatially variable genes identified via Moran's I statistic, with the left two panels displaying genes from sample A1_colon_d0 and the right two panels displaying genes from sample B1_colon_d14. Please click here to view a larger version of this figure.

Figure 3: Single-cell reference data processing and annotation. (A) Quality control metrics for scRNA-seq reference data before and after filtering. (B) UMAP visualization of scRNA-seq data colored by unsupervised clusters. (C) Dot plot showing expression scores of canonical cell-type marker genes across clusters. (D) Annotated UMAP visualization of scRNA-seq data with major cell types labeled. (E) Cellular composition of the scRNA-seq reference dataset. The dashed red line indicates the downsampling threshold (n = 50 cells per type) applied during SPOTlight deconvolution to balance computational efficiency and cell-type representation. Please click here to view a larger version of this figure.

Figure 4: Spatial deconvolution of cellular heterogeneity. (A,B) Spatial scatterpie plots from SPOTlight deconvolution showing the proportional composition of major cell types at each spot for samples A1 (A) and B1 (B). (C) Representative spatial distribution of B cells in samples A1 (left) and B1 (right), demonstrating the spatially resolved localization patterns of a specific immune cell population identified through deconvolution. Please click here to view a larger version of this figure.

Figure 5: Interactive region-of-interest analysis and unsupervised deconvolution comparison. (A) Interface of the custom "Select Spatial Spots" tool showing interactive selection of regions corresponding to proximal colon, distal colon, and other tissue domains. (B) Spatial scatterpie visualization of unsupervised deconvolution results (STdeconvolve) for sample A1, with spots colored according to the manually annotated regions from (A), illustrating correspondence between histology-based annotation and computationally derived cell topic distributions. Please click here to view a larger version of this figure.

Figure 6: Spatially informed cell-cell communication networks. (A,B) Ligand-receptor interaction networks inferred by Giotto for samples A1 (A) and B1 (B). Nodes represent cell types, edges represent significant ligand-receptor pairs (FDR < 0.05), and edge thickness corresponds to interaction strength. To ensure comparability and visualization clarity, a uniform significance threshold (FDR < 0.05) was applied across all samples, and the top 20 interactions ranked by log2FC are displayed for each condition. Networks highlight cell-type-specific communication patterns within the spatial context of colon tissue. Please click here to view a larger version of this figure.
Supplementary Figure 1: Quantitative evaluation of parameter optimization for dimensionality reduction. The elbow plot demonstrates the workflow's programmatic approach to dynamically selecting the optimal number of principal components (PCs). The selection is calculated based on cumulative standard deviation and marginal variance thresholds, represented by the red vertical line, to capture biological variance while mitigating technical noise prior to downstream clustering.Please click here to download this file.
Supplementary Figure 2: Validation of spatial clustering using canonical colonic layer-specific markers. (A) Dot plot showing the enriched expression of epithelial, stromal, and smooth muscle markers across computational clusters. (B) Spatial feature plots mapping representative markers (Epcam, Col1a1, Acta2) back to the tissue coordinates.Please click here to download this file.