From a CellProfiler run#

CellProfiler’s ExportToSpreadsheet writes one CSV per object plus an Image.csv. mantispy joins those into a single AnnData and parses what every column means: which object it describes, which measurement family it belongs to and which channel it was measured in. Later steps, such as selecting intensity features or building feature sets, read this annotation instead of re-parsing names.

jump_export() downloads one real export, the directory the JUMP pipeline wrote for one field of view of plate BR00121438.

import tempfile
from pathlib import Path

import mantispy as mt

directory = mt.ds.jump_export()
sorted(path.name for path in directory.iterdir())
['Cells.csv', 'Cytoplasm.csv', 'Experiment.csv', 'Image.csv', 'Nuclei.csv']

Cells.csv, Cytoplasm.csv and Nuclei.csv hold one row per object. Image.csv holds one row per field of view: the plate, well and site it came from, the file names of each channel, and the MeasureImageQuality statistics. Experiment.csv records the pipeline itself.

Reading#

primary_object sets what a row of the result is, and the other objects are joined onto it. The JUMP pipeline gives Cytoplasm a parent cell and a parent nucleus and gives Cells no parent among the exported objects, so the cytoplasm is the object that joins all three tables here, and one row is one cell.

adata = mt.io.read_profiles(directory, primary_object="Cytoplasm")
adata
AnnData object with n_obs × n_vars = 239 × 5839
    obs: 'Metadata_ImageNumber', 'Metadata_ObjectNumber', 'Metadata_AbsPositionZ', 'Metadata_AbsTime', 'Metadata_BinningX', 'Metadata_BinningY', 'Metadata_ChannelID', 'Metadata_ChannelName', 'Metadata_Col', 'Metadata_ExposureTime', 'Metadata_FieldID', 'Metadata_ImageResolutionX', 'Metadata_ImageResolutionY', 'Metadata_ImageSizeX', 'Metadata_ImageSizeY', 'Metadata_MainEmissionWavelength', 'Metadata_MainExcitationWavelength', 'Metadata_MaxIntensity', 'Metadata_ObjectiveMagnification', 'Metadata_ObjectiveNA', 'Metadata_PlaneID', 'Metadata_Plate', 'Metadata_PositionX', 'Metadata_PositionY', 'Metadata_PositionZ', 'Metadata_Row', 'Metadata_Site', 'Metadata_Well', 'Metadata_Center_X', 'Metadata_Center_Y'
    var: 'object', 'feature_group', 'feature', 'channel', 'scale', 'angle', 'gray_levels', 'radial_bin', 'params', 'is_feature'
    uns: 'mantispy'
    layers: None (.X)

What landed where#

X holds the measurements, one row per cell, as float32.

adata.X.shape, adata.X.dtype
((239, 5839), dtype('float32'))

obs holds the columns that identify where a cell came from, all with the Metadata_ prefix. Among them are the centroid, Metadata_Center_X and Metadata_Center_Y. It is not a profile feature, but QC uses it to find cells on the edge of a field and crowding needs it, so it is kept here.

adata.obs[["Metadata_Plate", "Metadata_Well", "Metadata_Site", "Metadata_Center_X", "Metadata_Center_Y"]].head()
Metadata_Plate Metadata_Well Metadata_Site Metadata_Center_X Metadata_Center_Y
0 BR00121438 J04 1 138.503697 21.141035
1 BR00121438 J04 1 522.719343 31.477202
2 BR00121438 J04 1 954.107582 36.398345
3 BR00121438 J04 1 730.240582 25.867151
4 BR00121438 J04 1 626.728038 33.968995

var is the parsed annotation, with one row per feature describing what it measures.

adata.var.head()
object feature_group feature channel scale angle gray_levels radial_bin params is_feature
Cytoplasm_AreaShape_Area Cytoplasm AreaShape Area NaN NaN NaN NaN NaN NaN True
Cytoplasm_AreaShape_BoundingBoxArea Cytoplasm AreaShape BoundingBoxArea NaN NaN NaN NaN NaN NaN True
Cytoplasm_AreaShape_Compactness Cytoplasm AreaShape Compactness NaN NaN NaN NaN NaN NaN True
Cytoplasm_AreaShape_Eccentricity Cytoplasm AreaShape Eccentricity NaN NaN NaN NaN NaN NaN True
Cytoplasm_AreaShape_EquivalentDiameter Cytoplasm AreaShape EquivalentDiameter NaN NaN NaN NaN NaN NaN True

Channel names are not hardcoded. CellProfiler writes one Intensity_MeanIntensity_<channel> measurement per channel it measured, and the names are read from those. The file names in Image.csv can differ: this run loaded OrigDNA, corrected it with IllumDNA and measured the result as DNA. A two-channel assay with channels called Hoechst and GFP parses the same way as a Cell Painting run [Bray et al., 2016]. This one measured five fluorescence channels and three brightfield planes.

adata.uns["mantispy"]["channels"]
['AGP', 'BFHigh', 'BFLow', 'Brightfield', 'DNA', 'ER', 'Mito', 'RNA']

Columns that are not features#

CellProfiler also writes columns that are not measurements: object numbers, parent links, child counts and positions. The parser marks them is_feature = False and keeps them out of X. CellProfiler 4 writes the centroid and the bounding box under AreaShape, beside the shape measurements, and those are positions as well.

from mantispy._core.features import parse_feature_names

parse_feature_names(
    [
        "Cytoplasm_AreaShape_Area",
        "Cells_AreaShape_Center_X",
        "Cells_AreaShape_BoundingBoxMaximum_Y",
        "Nuclei_Location_CenterMassIntensity_X_DNA",
        "Cytoplasm_Parent_Cells",
    ]
)[["object", "feature_group", "feature", "is_feature"]]
object feature_group feature is_feature
Cytoplasm_AreaShape_Area Cytoplasm AreaShape Area True
Cells_AreaShape_Center_X Cells AreaShape NaN False
Cells_AreaShape_BoundingBoxMaximum_Y Cells AreaShape NaN False
Nuclei_Location_CenterMassIntensity_X_DNA Nuclei Location NaN False
Cytoplasm_Parent_Cells Cytoplasm Parent NaN False

How objects are joined#

Cells and Nuclei measurements are joined onto the cytoplasm they belong to rather than added as extra rows, so one row still means one cell and Nuclei_AreaShape_Area sits next to Cells_AreaShape_Area. The link is read from the Parent_ columns. Matching on object number instead would pair unrelated objects that happen to share a number, without any warning, so a missing link raises an error, and by default the join must be one-to-one.

sorted(name for name in adata.var_names if name.startswith("Nuclei"))[:4]
['Nuclei_AreaShape_Area',
 'Nuclei_AreaShape_BoundingBoxArea',
 'Nuclei_AreaShape_Compactness',
 'Nuclei_AreaShape_Eccentricity']

Image quality#

MeasureImageQuality describes a field of view rather than a cell, so it does not belong in X. It is stored separately, keyed by image, together with the plate and well, so that image QC can set thresholds per plate.

adata.uns["mantispy"]["image_table"].filter(like="FocusScore")
ImageQuality_FocusScore_OrigAGP ImageQuality_FocusScore_OrigBrightfield ImageQuality_FocusScore_OrigBrightfield_H ImageQuality_FocusScore_OrigBrightfield_L ImageQuality_FocusScore_OrigDNA ImageQuality_FocusScore_OrigER ImageQuality_FocusScore_OrigMito ImageQuality_FocusScore_OrigRNA ImageQuality_LocalFocusScore_OrigAGP_10 ImageQuality_LocalFocusScore_OrigAGP_20 ... ImageQuality_LocalFocusScore_OrigER_5 ImageQuality_LocalFocusScore_OrigER_50 ImageQuality_LocalFocusScore_OrigMito_10 ImageQuality_LocalFocusScore_OrigMito_20 ImageQuality_LocalFocusScore_OrigMito_5 ImageQuality_LocalFocusScore_OrigMito_50 ImageQuality_LocalFocusScore_OrigRNA_10 ImageQuality_LocalFocusScore_OrigRNA_20 ImageQuality_LocalFocusScore_OrigRNA_5 ImageQuality_LocalFocusScore_OrigRNA_50
ImageNumber
1972 0.020928 0.000115 0.000331 0.000209 0.103867 0.059147 0.003104 0.058702 0.03778 0.038886 ... 0.018465 0.056921 0.022152 0.013422 0.026218 0.005306 0.04627 0.074392 0.029123 0.070208

1 rows × 40 columns

What each well received#

An export knows its plate, well and site, and nothing about what the well was treated with. For a JUMP plate the annotation is a join, once the laboratory is named:

adata.obs["Metadata_Source"] = "source_4"
mt.pp.annotate_jump(adata)
adata.obs[["Metadata_Well", "Metadata_Perturbation", "Metadata_Control"]].drop_duplicates()
Metadata_Well Metadata_Perturbation Metadata_Control
0 J04 JCP2022_033924 True

For your own screen, read_profiles(directory, platemap="platemap.csv") joins a table mapping wells to treatments, and mt.pp.annotate_controls(adata, negcon=("DMSO",)) marks the negative controls, which several later steps use.

Checking the contract#

validate returns a report instead of raising, so you can inspect a partly formed object.

report = mt.io.validate(adata)
report.ok
True
broken = adata.copy()
broken.obs = broken.obs.drop(columns="Metadata_Plate")
print(mt.io.validate(broken))
ERROR: obs is missing required column 'Metadata_Plate' (resolution 'cell')

Writing#

write validates first, then writes h5ad (or zarr for a .zarr suffix). The schema version is stored in the file and checked on read, so a file written by a future incompatible version raises an error instead of loading incorrectly.

path = Path(tempfile.mkdtemp()) / "cells.h5ad"
mt.io.write(adata, path)
mt.io.read(path)
AnnData object with n_obs × n_vars = 239 × 5839
    obs: 'Metadata_ImageNumber', 'Metadata_ObjectNumber', 'Metadata_AbsPositionZ', 'Metadata_AbsTime', 'Metadata_BinningX', 'Metadata_BinningY', 'Metadata_ChannelID', 'Metadata_ChannelName', 'Metadata_Col', 'Metadata_ExposureTime', 'Metadata_FieldID', 'Metadata_ImageResolutionX', 'Metadata_ImageResolutionY', 'Metadata_ImageSizeX', 'Metadata_ImageSizeY', 'Metadata_MainEmissionWavelength', 'Metadata_MainExcitationWavelength', 'Metadata_MaxIntensity', 'Metadata_ObjectiveMagnification', 'Metadata_ObjectiveNA', 'Metadata_PlaneID', 'Metadata_Plate', 'Metadata_PositionX', 'Metadata_PositionY', 'Metadata_PositionZ', 'Metadata_Row', 'Metadata_Site', 'Metadata_Well', 'Metadata_Center_X', 'Metadata_Center_Y', 'Metadata_Source', 'Metadata_JCP2022', 'Metadata_InChIKey', 'Metadata_Perturbation', 'Metadata_Control'
    var: 'object', 'feature_group', 'feature', 'channel', 'scale', 'angle', 'gray_levels', 'radial_bin', 'params', 'is_feature'
    uns: 'mantispy'
    layers: None (.X)

Already have profiles?#

Published profiles and JUMP reads well-level tables from pycytominer [Serrano et al., 2025], CytoTable, the Cell Painting Gallery [Weisbart et al., 2024] and JUMP.