Skip to content

Clust Module API Reference

This module provides the main Matrix class for hierarchical data clustering and visualization.

AxisEntity

Bases: TypedDict

Describes what entity a clustergram axis represents.

Attributes:

Name Type Description
entity str

The type of entity (cell, gene, nbhd, cluster, etc.)

attr str

The attribute of that entity (name, leiden, custom_column, etc.) - For cells: 'leiden' means cell clusters, 'name' means individual cells - For nbhd: 'name' means specific neighborhoods - For genes: typically 'name'

Examples:

Clustergram with cell clusters on rows (cells grouped by leiden)

{"entity": "cell", "attr": "leiden"}

Clustergram with specific cells on columns

{"entity": "cell", "attr": "name"}

Clustergram with neighborhoods on columns

{"entity": "nbhd", "attr": "name"}

Clustergram with genes on rows

{"entity": "gene", "attr": "name"}

Hextile neighborhoods by cell clusters

row: {"entity": "cell", "attr": "leiden"} col: {"entity": "hextile", "attr": "nbhd_cluster"}

EntityType

Bases: Enum

Entity types that can be represented in clustergram rows/columns.

Matrix

High-performance matrix class for single-cell genomics data processing.

Features automatic processing pipeline, hierarchical clustering, and visualization export. Uses intelligent caching for performance with large datasets.

Examples:

Basic usage - applies norm_col='total', norm_row='zscore'

mat = Matrix(adata) viz_data = mat.cluster()

Custom processing with colors

mat = Matrix(adata, filter_genes=5000, norm_row='qn', global_colors={"high": "red", "low": "blue"})

No processing

mat = Matrix(adata, disable_processing=True)

dat property

Lazy dat structure with intelligent caching.

__init__(data=None, meta_col=None, meta_row=None, col_attr=None, row_attr=None, row_entity='gene', col_entity='cell_cluster', filter_genes=None, norm_col='total', norm_row='zscore', disable_processing=True, global_colors=None, name=None, *, collection=None, color_by=None, size_by=None, dot_plot=None)

Create Matrix with automatic processing unless disabled.

Parameters:

Name Type Description Default
data DataFrame | AnnData | None

DataFrame or AnnData object

None
meta_col DataFrame | None

Column metadata (for DataFrame input)

None
meta_row DataFrame | None

Row metadata (for DataFrame input)

None
col_attr list[str] | None

Column attribute names (categorical or numeric)

None
row_attr list[str] | None

Row attribute names (categorical or numeric)

None
row_entity str | dict | AxisEntity | None

Entity specification for rows. Accepted formats: - str: Shorthand with implicit attr mapping: - "gene" → {"entity": "gene", "attr": "name"} - "nbhd" → {"entity": "nbhd", "attr": "name"} - "cell" → {"entity": "cell", "attr": "name"} - "hextile" → {"entity": "hextile", "attr": "name"} - "cell_cluster" or "cluster" → {"entity": "cell", "attr": "leiden"} - tuple: Compact format, e.g., ("nbhd", "name") - dict: Full format, e.g., {"entity": "nbhd", "attr": "name"}

'gene'
col_entity str | dict | AxisEntity | None

Entity specification for columns (same formats as row_entity)

'cell_cluster'
filter_genes int | None

Number of top variable genes to keep (None = no filtering)

None
norm_col str | None

Column normalization ('total', 'zscore', 'qn', None)

'total'
norm_row str | None

Row normalization ('total', 'zscore', 'qn', None)

'zscore'
disable_processing bool

Skip automatic processing (default: False)

True
global_colors dict[str, str] | DataFrame | None

Global category color mapping (dict or DataFrame with 'color' column)

None
name str | None

Name for the matrix (default: None)

None
collection Any

A Celldega collection (SetCollection/DatasetCollection/etc.) to build data from, in place of passing data directly — equivalent to data=collection.mod[color_by]. Requires color_by.

None
color_by str | None

Modality key on collection driving color/opacity (the main matrix).

None
size_by str | None

Optional modality key on collection driving the secondary size channel (dot-plot size, e.g. fraction of cells expressing — but any per-cell magnitude works, e.g. significance) — equivalent to calling .set_dot_matrix(collection.mod[size_by]) after construction. Omit for a matrix with no size channel.

None
dot_plot str | None

Alias for size_by (pass either, not both).

None

Examples:

mat = Matrix(adata) # Applies norm_col='total', norm_row='zscore'

Custom processing with colors

colors = {"Cancer": "#ff0000", "Normal": "#0000ff"} mat = Matrix(adata, filter_genes=5000, norm_row='qn', global_colors=colors)

No processing

mat = Matrix(adata, disable_processing=True)

Dot plot directly from a SetCollection, no manual DataFrame wrangling

setc.calc_signature(adata, modality_name="expression") setc.calc_signature(adata, modality_name="fraction_expressing", aggregate="fraction") mat = Matrix(collection=setc, color_by="expression", size_by="fraction_expressing")

Raw matrix without data

mat = Matrix() # Empty matrix for manual loading

With entity specifications for widget interaction:

Genes (rows) by cell clusters (columns) - typical gene expression heatmap

mat = Matrix(df, row_entity="gene", col_entity="cell_cluster")

Or equivalently with new format:

mat = Matrix(df, row_entity={"entity": "gene", "attr": "name"}, col_entity={"entity": "cell", "attr": "leiden"})

Neighborhoods by cell types

mat = Matrix(df, row_entity={"entity": "cell", "attr": "leiden"}, col_entity={"entity": "nbhd", "attr": "name"})

add_category(axis, name, data)

Add category to metadata.

Parameters:

Name Type Description Default
axis AxisInput

'row'/'col', 0/1, or Axis enum (0/ROW=rows, 1/COL=columns)

required
name str

Category name

required
data Series

Category values (must match axis length)

required

add_cats(axis, cat_data)

Add multiple categories to metadata.

Parameters:

Name Type Description Default
axis AxisInput

'row'/'col', 0/1, or Axis enum

required
cat_data dict[str, Any]

Dict with category name as key, values as list/Series/dict

required

Examples:

Add multiple categories at once

mat.add_cats('col', { 'cell_type': ['T-cell', 'B-cell', 'NK-cell'], 'treatment': ['control', 'treated', 'control'] })

From existing metadata

mat.add_cats('col', meta_df.to_dict('series'))

clust(dist_type='cosine', linkage_type='average', force=False)

Perform hierarchical clustering.

Parameters:

Name Type Description Default
dist_type DistanceType

Distance metric ('cosine', 'euclidean', 'correlation')

'cosine'
linkage_type LinkageType

Linkage method ('average', 'complete', 'ward')

'average'
force bool

Override size limits for large matrices

False

cluster(**cluster_kwargs)

Perform clustering and return visualization data.

Parameters:

Name Type Description Default
**cluster_kwargs Any

Clustering parameters (dist_type, linkage_type, force)

{}

Returns:

Name Type Description
dict dict[str, Any]

Visualization-ready JSON structure

Examples:

mat = Matrix(adata) viz_data = mat.cluster() # Use defaults viz_data = mat.cluster(dist_type='euclidean', linkage_type='ward')

downsample_to(category='leiden', axis='col', propagate_metadata=False)

Downsample data by aggregating categories using scanpy.get.aggregate.

Parameters:

Name Type Description Default
category str

Metadata column to aggregate by

'leiden'
axis AxisInput

Which axis to aggregate ('col'/1/COL for cells, 'row'/0/ROW for genes)

'col'
propagate_metadata bool | list[str]

Whether to propagate other metadata columns to the aggregated result using the modal (most frequent) value per group. - False: Skip metadata propagation (fast, default) - True: Propagate all metadata columns (slow for large datasets) - list[str]: Propagate only specified columns

False
Requires

scanpy for aggregation functionality

Note

Uses scanpy.get.aggregate under the hood for fast mean aggregation. See: https://scanpy.readthedocs.io/en/stable/generated/scanpy.get.aggregate.html

export_viz_json()

Export visualization as JSON dict.

.. deprecated:: 0.10 Use :meth:export_viz_parquet instead.

export_viz_json_string()

Export visualization as JSON string.

.. deprecated:: 0.10 Use :meth:export_viz_parquet instead.

export_viz_parquet()

Export visualization using Parquet encoded tables.

export_viz_to_widget(which_viz='viz')

Export visualization for widget.

.. deprecated:: 0.10 Use :class:celldega.viz.Clustergram with matrix instead.

filter(axis, by, num)

Filter features by specified metric.

Parameters:

Name Type Description Default
axis AxisInput

'row'/'col', 0/1, or Axis enum (0/ROW=rows, 1/COL=columns)

required
by FilterType

Metric ('var' for variance, 'mean' for mean)

required
num int

Number of top features to keep

required

load_adata(adata, col_attr=None, row_attr=None)

Load AnnData object.

Parameters:

Name Type Description Default
adata AnnData

AnnData object (will be transposed to genes x cells)

required

load_df(df, meta_col=None, meta_row=None, col_attr=None, row_attr=None)

Load DataFrame with metadata.

Parameters:

Name Type Description Default
df DataFrame

Data matrix

required
meta_col DataFrame | None

Column metadata (must match df.columns)

None
meta_row DataFrame | None

Row metadata (must match df.index)

None
col_attr list[str] | None

Column attribute names for viz (categorical or numeric)

None
row_attr list[str] | None

Row attribute names for viz (categorical or numeric)

None

make_viz()

Generate visualization data structure.

norm(axis, by)

Normalize data along specified axis.

Parameters:

Name Type Description Default
axis AxisInput

'row'/'col', 0/1, or Axis enum (0/ROW=rows, 1/COL=columns)

required
by NormType

Normalization method ('total', 'zscore', 'qn')

required

process(filter_genes=None, norm_col='total', norm_row='zscore')

Apply processing pipeline to the matrix.

Parameters:

Name Type Description Default
filter_genes int | None

Number of top variable genes to keep

None
norm_col str | None

Column normalization method ('total', 'zscore', 'qn', None)

'total'
norm_row str | None

Row normalization method ('total', 'zscore', 'qn', None)

'zscore'

Examples:

mat = Matrix(adata, disable_processing=True) # Raw data mat.process(filter_genes=5000, norm_row='qn') # Custom processing

random_subsample(axis, num, seed=42)

Randomly subsample features.

Parameters:

Name Type Description Default
axis AxisInput

'row'/'col', 0/1, or Axis enum (0/ROW=rows, 1/COL=columns)

required
num int

Number of features to sample

required
seed int

Random seed for reproducibility

42

set_cat_color(axis, cat_index, cat_name, color)

Set color for specific category value in a specific category column.

Parameters:

Name Type Description Default
axis AxisInput

'row'/'col', 0/1, or Axis enum

required
cat_index int

Category column index (1-based, like original Network)

required
cat_name str

Category value name to color

required
color str

Hex color string or named color

required
Example

Set color for 'Cancer' in the first column category

mat.set_cat_color('col', 1, 'Cancer', '#ff0000')

set_cat_colors(axis, cat_index, color_mapping)

Set colors for multiple category values in a specific category column.

Parameters:

Name Type Description Default
axis AxisInput

'row'/'col', 0/1, or Axis enum

required
cat_index int

Category column index (1-based)

required
color_mapping dict[str, str]

Dict mapping category values to colors

required
Example

Set colors for multiple values in tissue type category

mat.set_cat_colors('col', 1, { 'Liver': '#00ff00', 'Brain': '#ffff00', 'Heart': '#ff00ff' })

set_dot_matrix(dot)

Attach a secondary matrix that drives dot-plot size encoding.

The values are interpreted as a per-cell size channel (typically the fraction of cells in a cluster expressing a gene, in [0, 1]) that is rendered as square/dot size in a dot-plot :class:~celldega.viz.Clustergram, independently of the main matrix which continues to drive color/opacity and the clustering order.

The dot matrix is aligned to the main matrix by row and column name at export time, so it does not need to share the clustered ordering — only the same labels. Missing entries become 0 (no dot).

Parameters:

Name Type Description Default
dot DataFrame | ndarray | AnnData | Matrix

A DataFrame indexed by row names with columns of column names, a 2D array matching the main matrix shape/order, an AnnData (X with obs_names rows / var_names cols), or another Matrix.

required

Returns:

Type Description
Matrix

self (to allow chaining, e.g. Matrix(df).set_dot_matrix(frac)).

set_global_cat_colors(color_mapping=None)

Set global category color mapping that applies across all categories.

Parameters:

Name Type Description Default
color_mapping dict[str, str] | DataFrame | None

Dict mapping category values to colors, DataFrame with 'color' column, or None to auto-generate

None
Note

If metadata has a 'color' column, those colors will be used automatically.

set_matrix_colors(pos='red', neg='blue')

Set matrix color scheme for positive and negative values.

Parameters:

Name Type Description Default
pos str

Color for positive values (hex or named color)

'red'
neg str

Color for negative values (hex or named color)

'blue'
Example

mat.set_matrix_colors(pos="#ff0000", neg="#0000ff")

subset(axis, by)

Subset data by feature list.

Parameters:

Name Type Description Default
axis AxisInput

'row'/'col', 0/1, or Axis enum (0/ROW=rows, 1/COL=columns)

required
by list[str]

List of feature names to keep

required

to_adata()

Convert to AnnData object.

to_cluster(axis='row', n_clusters=None, threshold=None, criterion=None)

Cut the dendrogram into flat cluster labels.

Cuts the hierarchical clustering linkage (computed by :meth:clust) along one axis and returns a label per row/column. Use n_clusters to request a fixed number of clusters (fcluster with criterion="maxclust") or threshold to cut at a linkage distance (criterion="distance"). Exactly one of the two must be given unless criterion is set explicitly.

This is the programmatic counterpart to the Clustergram's interactive dendrogram slider: it turns a clustered Matrix into discrete groups (e.g. consensus domains, meta-clusters) that can be attached back to a collection's obs / var.

Parameters:

Name Type Description Default
axis AxisInput

"row"/"col" (or 0/1) — which dendrogram to cut.

'row'
n_clusters int | None

Target number of flat clusters (maxclust criterion).

None
threshold float | None

Linkage-distance cutoff (distance criterion).

None
criterion str | None

Explicit scipy fcluster criterion; overrides the n_clusters/threshold inference when given (paired with whichever of the two is supplied as t).

None

Returns:

Type Description
Series

A pd.Series of integer cluster labels indexed by the axis names

Series

(data.index for rows, data.columns for columns), in data order.

Raises:

Type Description
ValueError

If the matrix is unclustered, the linkage is empty, or neither n_clusters nor threshold is provided.

to_df()

Return DataFrame copy of data.

write_dega_files(path, name=None)

Write Clustergram visualization data to a DegaFiles directory.

This creates a cgm/ subdirectory containing the parquet files needed to load the Clustergram in JavaScript without a Python backend.

Parameters

path : str or Path Path to the DegaFiles directory (the same directory used for Landscape and Yearbook data). name : str, optional Name for this Clustergram. If provided, files are saved to cgm/{name}/. If None, uses the matrix's name attribute, or "default" if no name is set.

Examples

mat = Matrix(adata) mat.clust() mat.write_dega_files("./my_dega_files", name="skin_cancer_clusters")

JavaScript can then load from:

base_url + '/cgm/skin_cancer_clusters/'

Notes

The following files are created: - mat.parquet: The matrix data - row_nodes.parquet: Row node information - col_nodes.parquet: Column node information - row_linkage.parquet: Row dendrogram linkage - col_linkage.parquet: Column dendrogram linkage - meta.json: Metadata including colors and config

normalize_axis_entity(value)

Normalize an axis entity specification to the AxisEntity format.

Handles backwards compatibility with string-only entity values and supports compact tuple format.

Parameters:

Name Type Description Default
value str | tuple | dict | AxisEntity | None

Entity specification - can be: - str: Shorthand format with implicit attr (see mapping below) - tuple: Compact format (entity, attr) e.g., ("nbhd", "name") - dict/AxisEntity: Full format with entity and attr keys - None: Returns default {"entity": "gene", "attr": "name"}

required
String Shorthand Mapping

When a string is provided, the following implicit attr values are used: - "gene" → {"entity": "gene", "attr": "name"} - "nbhd" → {"entity": "nbhd", "attr": "name"} - "cell" → {"entity": "cell", "attr": "name"} - "hextile" → {"entity": "hextile", "attr": "name"} - "cell_cluster" or "cluster" → {"entity": "cell", "attr": "leiden"} - any other string → {"entity": , "attr": "name"}

Returns:

Type Description
AxisEntity

AxisEntity with entity and attr keys

Examples:

String shorthand (attr is implicit)

>>> normalize_axis_entity("gene")
{"entity": "gene", "attr": "name"}
>>> normalize_axis_entity("nbhd")
{"entity": "nbhd", "attr": "name"}
>>> normalize_axis_entity("cell_cluster")
{"entity": "cell", "attr": "leiden"}

Tuple format (explicit entity and attr)

>>> normalize_axis_entity(("nbhd", "name"))
{"entity": "nbhd", "attr": "name"}
>>> normalize_axis_entity(("cell", "leiden"))
{"entity": "cell", "attr": "leiden"}

Dict format (most explicit)

>>> normalize_axis_entity({"entity": "cell", "attr": "leiden"})
{"entity": "cell", "attr": "leiden"}