mantispy.ds.cp_posh

Contents

mantispy.ds.cp_posh#

mantispy.ds.cp_posh(cache_dir=None, *, aggregated=False, feature_selected=False)[source]#

Single cells of insitro cp-POSH, a broad-morphology pooled CRISPR Cell Painting screen.

The 124-gene proof-of-concept dataset from insitro/cp-posh: A549 cells carrying a pooled CRISPR-knockout library, stained with a six-channel Cell Painting panel (WGA, a mitochondrial probe, phalloidin, concanavalin A, DAPI and a marker round) and read by in-situ sequencing of the guide barcodes. Each cell gets a broad, untargeted morphology profile of about 1,278 CellStats features rather than the handful of hand-picked readouts a targeted screen keeps, so it is the broad-morphology complement to scallops_arv471().

The features are already well-normalized by the authors, so normalize() is not needed before analysis; a per-plate control normalization would re-do work already done.

The base and its variants are pre-built by scripts/build_staged_datasets.py from the raw parquet and rehosted on scverse-exampledata, so the loader fetches a single h5ad rather than reassembling the base on every call. The two flags select the variant:

  • both False: the cells with every feature.

  • feature_selected=True: the cells after pycytominer-default feature selection.

  • aggregated=True: one median profile per (Metadata_Gene, Metadata_sgRNA) guide, with Metadata_CellCount.

  • aggregated=True, feature_selected=True: that same guide aggregate on the feature-selected block.

Parameters:
  • cache_dir (str | Path | None (default: None)) – Where to keep the download. Defaults to mantispy.settings.cache_dir.

  • aggregated (bool (default: False)) – Return the guide-level median aggregate instead of the cells.

  • feature_selected (bool (default: False)) – Return the feature-selected block instead of all features.

Returns:

Metadata_Gene: the gene the cell’s guide targets, taken from the upstream gene_id. The two control classes keep their upstream spellings, "nontargeting" (the non-targeting guides) and "intergenic" (guides against intergenic regions); "nontargeting" is the spelling the analysis functions read.

Metadata_sgRNA: the guide, the upstream barcode.

Metadata_Perturbation: the guide again, so each guide is its own perturbation, matching scallops_arv471().

Metadata_Perturbation_Type: "crispr", the kind of screen this is.

Metadata_Plate: the plate, the upstream plate ("EL37").

Metadata_Well: the physical well, such as "B04", taken from the upstream plate_well ("EL37_B04") by dropping the plate prefix so the well vocabulary can parse it.

Metadata_Control: True for the non-targeting and intergenic guides, the reference hit_calling() and the control normalization test against.

The upstream treatment column is a constant (no small molecule) and is dropped, and the ID becomes the observation index. The known-mechanism genes, whose knockout moves cells away from the controls, are KIF18A, the proteasome (PSMB1, PSMD4), the mitochondrial ribosome (MRPL43, MRPS5), the ARP2/3 complex (ARPC4, ACTR6) and COPI (COPE, ARCN1), scored against nontargeting and intergenic.

Return type:

AnnData

Notes

The CellStats feature names are insitro’s own, not CellProfiler’s <Object>_<Group>_<Feature>_<Channel>, so the annotation columns of var are supplied empty rather than parsed. Left to the parser, a name such as nucleus_mask_height would read as the mask feature group of a nucleus object and invent feature families that are not there, the same reason the learned embeddings of jump_lite() carry an empty annotation. Anything that reads var["feature_group"] or var["channel"] has nothing to work with here.

A cell carries no count. The aggregated variant is one median profile per (Metadata_Gene, Metadata_sgRNA) guide, which writes Metadata_CellCount.

Raises:

ValueError – aggregated or feature_selected is not a bool.