Visium HD human colorectal cancer Landscape + Clustergram + Enrich¶
This tutorial explores a Visium HD human colorectal cancer (CRC) sample that was already segmented into cells and clustered with Scanpy (220,703 cells, 15 Leiden clusters). It shows the handoff from Scanpy to Celldega:
- view the clustered cells in their tissue context with a Landscape,
- organize the clusters with a SetCollection: pseudobulk expression, the fraction of cells expressing each gene, and marker genes ranked per cluster,
- explore the result in a Clustergram linked to the Landscape and to Enrich (Enrichr gene-set enrichment), and
- annotate clusters and genes interactively.
Data source: broadinstitute/Celldega_Supporting_Data.
Files are downloaded once through direct Hugging Face /resolve/main/ URLs, so
huggingface_hub is not a notebook dependency.
from importlib.metadata import version
from pathlib import Path
import shutil
from urllib.request import Request, urlopen
import warnings
import celldega as dega
from pandas.errors import PerformanceWarning
import scanpy as sc
print(f"celldega {version('celldega')}")
print(f"scanpy {version('scanpy')}")
celldega 0.26.0 scanpy 1.10.4
Data¶
Two files are used, each downloaded on first run and reused afterwards:
crc_cluster.h5ad(~1.6 GB): the Scanpy analysis, with log-normalized expression inX, Leiden clusters inobs["leiden"], and a UMAP.crc_raw_feature_cell_matrix.h5(~280 MB): Space Ranger's cell-segmented count matrix for the same sample, which supplies the raw counts the analysis file does not keep.
HF_RESOLVE_BASE = (
"https://huggingface.co/datasets/"
"broadinstitute/Celldega_Supporting_Data/resolve/main"
)
CLUSTERED_FILENAME = "crc_cluster.h5ad"
RAW_COUNTS_FILENAME = "crc_raw_feature_cell_matrix.h5"
# Store data in the repository's notebooks/data folder when run from a checkout,
# otherwise beside the notebook.
working_dir = Path.cwd().resolve()
project_root = next(
(
d
for d in (working_dir, *working_dir.parents)
if (d / "pyproject.toml").is_file() and (d / "notebooks").is_dir()
),
None,
)
data_dir = (project_root / "notebooks" if project_root else working_dir) / "data"
data_dir = data_dir / "visium-hd_data" / "Visium_HD_Human_Colon_Cancer"
def download_if_missing(remote_filename):
"""Download one public Hub file into data_dir, with no HF client dependency."""
destination = data_dir / remote_filename
if destination.is_file():
return destination
destination.parent.mkdir(parents=True, exist_ok=True)
partial = destination.with_name(f"{destination.name}.part")
url = f"{HF_RESOLVE_BASE}/{remote_filename}?download=true"
request = Request(url, headers={"User-Agent": "celldega-example-notebook"})
print(f"Downloading {url}")
try:
with urlopen(request) as response, partial.open("wb") as output:
shutil.copyfileobj(response, output, length=8 * 1024 * 1024)
partial.replace(destination)
except Exception:
partial.unlink(missing_ok=True)
raise
return destination
adata = sc.read_h5ad(download_if_missing(CLUSTERED_FILENAME))
# Visium HD's cell-segmented matrix uses the same cell barcodes
# ("cellid_000000003-1", ...). It lists every probe-set gene, some under
# duplicate symbols, so make names unique the same way Scanpy did before
# aligning to the analyzed cells and genes.
with warnings.catch_warnings():
# Expected: fixed by var_names_make_unique() on the next line.
warnings.filterwarnings("ignore", message="Variable names are not unique")
raw = sc.read_10x_h5(download_if_missing(RAW_COUNTS_FILENAME))
raw.var_names_make_unique()
adata.layers["counts"] = raw[adata.obs_names, adata.var_names].X
del raw
adata
AnnData object with n_obs × n_vars = 220703 × 18132
obs: 'clusters', 'orig_clusters', 'leiden'
var: 'gene_ids', 'feature_types', 'genome'
uns: 'clusters', 'clusters_colors', 'log1p', 'neighbors', 'pca', 'umap'
obsm: 'X_pca', 'X_umap'
varm: 'PCs'
layers: 'counts'
obsp: 'connectivities', 'distances'
Plotting the UMAP also stores Scanpy's Leiden palette in
adata.uns["leiden_colors"]. The Landscape, SetCollection, and Clustergram below
all read that palette, so every view colors a cluster the same way.
sc.pl.umap(adata, color="leiden")
Landscape¶
The Landscape loads pre-tiled image and segmentation data (LandscapeFiles) from
base_url and takes cell metadata, including the Leiden clusters, from adata.
base_url = (
"https://raw.githubusercontent.com/SharkieJones/celldega_Visium-HD_hCRC/"
"refs/heads/main/Visium_HD_Human_Colon_Cancer_2025-10-25_2um"
)
landscape = dega.viz.Landscape(
technology="Visium-HD",
base_url=base_url,
height=600,
max_tiles_to_view=10,
adata=adata,
)
/var/folders/8d/jxpy9rd10j7fp2rcj_s5sz3c0000gq/T/ipykernel_42987/3932876498.py:6: UserWarning: Transformation matrix not found at https://raw.githubusercontent.com/SharkieJones/celldega_Visium-HD_hCRC/refs/heads/main/Visium_HD_Human_Colon_Cancer_2025-10-25_2um/micron_to_image_transform.csv. Using identity. landscape = dega.viz.Landscape(
SetCollection: cluster signatures and marker genes¶
A SetCollection treats each Leiden cluster as a set of cells and calculates
set-level data from the cell-level AnnData:
expression: a log1p-CPM pseudobulk signature from each cluster's summed counts,expression.layers["fraction_expressing"]: the fraction of each cluster's cells with nonzero counts for each gene (the Clustergram's dot size),uns["rank_genes_groups"]: marker genes ranked per cluster with Scanpy, on log-normalized counts as Scanpy recommends, andvar["leiden_marker"]: each gene's best-scoring cluster. Having clustered the cells, this defines matching gene clusters.
calc_signature prints the expression source behind each result.
setc = dega.SetCollection(adata, set_col="leiden", name="leiden")
with warnings.catch_warnings():
warnings.filterwarnings(
"ignore",
category=PerformanceWarning,
module=r"scanpy\.tools\._rank_genes_groups",
)
setc.calc_signature(
adata,
modality_name="expression",
layer="counts",
aggregate="sum",
normalization="log1p_cpm",
fraction_expressing_layer="fraction_expressing",
rank_genes_groups=True,
rank_genes_groups_kwargs={"method": "t-test"},
)
setc.mod["expression"]
calc_signature('expression'): sum of adata.layers['counts'] (normalization='log1p_cpm')
layers['fraction_expressing']: fraction of cells with adata.layers['counts'] > 0.0
uns['rank_genes_groups']: t-test on log1p(normalize_total(adata.layers['counts'])), grouped by 'leiden'
AnnData object with n_obs × n_vars = 15 × 18132
obs: 'leiden', 'n_cells', 'set_source', 'color'
var: 'gene_ids', 'feature_types', 'genome', 'gene', 'leiden_marker', 'entity_type'
uns: 'feature_type', 'aggregate', 'normalization', 'layer', 'expr_threshold', 'fraction_expressing_layer', 'rank_genes_groups', 'axis_entities'
layers: 'fraction_expressing'
Clustergram¶
The Matrix colors columns by cluster (col_attr=["leiden"]) and rows by each
gene's best-scoring cluster (row_attr=["leiden_marker"]), using the same
palette as the Landscape. All genes are hierarchically clustered; the
marker-ranked views add DIM slider levels with each cluster's top 1, 2, 3, ...
markers, each re-biclustered ahead of time. rank_dim=3 opens on the top 3
markers per cluster.
marker_levels = [1, 2, 3, 5, 10, 25, 50, 75]
mat = dega.clust.Matrix(
collection=setc,
color_by="expression",
size_by_layer="fraction_expressing",
col_attr=["leiden"],
row_attr=["leiden_marker"],
)
# Transform only the Matrix copy; the SetCollection keeps the original values.
mat.norm("row", by="zscore")
mat.cluster(view="rank_genes_groups", levels=marker_levels)
for view in mat.views:
print(f"{view['level']:>3} markers/cluster -> {view['n_rows']} genes")
cgm = dega.viz.Clustergram(
matrix=mat,
width=450,
height=500,
manual_col_cat="cell_type",
manual_row_cat="info",
rank_dim=3,
)
/Users/feni/Documents/celldega/src/celldega/clust/matrix.py:759: UserWarning: Found 17 constant features. Replacing zero std with small value to avoid inf/NaN. zscore_normalize_inplace(data_values, axis=0)
1 markers/cluster -> 13 genes 2 markers/cluster -> 21 genes 3 markers/cluster -> 33 genes 5 markers/cluster -> 52 genes 10 markers/cluster -> 92 genes 25 markers/cluster -> 231 genes 50 markers/cluster -> 456 genes 75 markers/cluster -> 681 genes
Linked Landscape, Clustergram, and Enrich¶
- Click a column (cluster) label to highlight that cluster in the Landscape and send its top marker genes to Enrich.
- Click a gene (row) label to show that gene's expression in the Landscape.
- Select a dendrogram cluster of genes to run enrichment on that gene set.
- Drag the DIM slider to move between marker-ranked views and all genes.
- Use the ROW / COL dropdowns in the control panel to switch which category the bar plots summarize, and hover a category to locate it in the matrix.
- Annotate: select columns or rows (e.g. from a dendrogram) and assign a
cell_typeorinfocategory. Annotations are colored in the matrix and summarized in the category bars.
linked_widgets = dega.viz.spatial_clustergram(
landscape,
cgm,
enrich=True,
)
linked_widgets
Working with the result in Python¶
The DIM slider is two-way. Setting rank_dim moves it, and requests snap to the
nearest precomputed level:
cgm.rank_dim = 10 # each cluster's top 10 markers
cgm.rank_dim = 0 # all genes
Interactive annotations sync back to Python, where they can be saved:
cgm.col_manual_df # cluster -> cell_type
cgm.row_manual_df # gene -> info
cgm.col_manual_colors_df # cell_type -> color
cgm.col_manual_df.to_parquet("crc_cluster_annotations.parquet")