mantispy.ds.scallops_arv471

mantispy.ds.scallops_arv471#

mantispy.ds.scallops_arv471(cache_dir=None, *, aggregated=False)[source]#

SCALLOPS ARV-471, single cells of an optical pooled screen under an estrogen-receptor degrader.

The drug arm of a genome-scale optical pooled CRISPR screen from Genentech/scallops-manuscript, its Figure 3 table. Cells express a guide library and are treated with ARV-471 (vepdegestrant), a PROTAC that recruits the CRL4-CRBN E3 ligase to the estrogen receptor and drives its degradation, then read by in-situ sequencing of the guide barcodes and a phenotype round that stains DNA and the estrogen receptor and counts ESR1, CCND1 and GREB1 transcripts. A guide that knocks out a gene the drug needs rescues the receptor, so cells carrying it keep the phenotype of an untreated cell. The genes with that known mechanism are the members of the ligase the PROTAC hijacks, CRBN, DDB1, CUL4A and CUL4B, and ESR1 itself, the drug’s target.

This loads only the ARV-471 condition, at single-cell resolution, so a hit is a guide whose cells sit away from the non-targeting cells in the phenotype space. The matched DMSO condition and the barcode-calling columns are left in the upstream file.

The base and its guide-level aggregate are pre-built by scripts/build_staged_datasets.py from the raw parquet and rehosted on scverse-exampledata, so the loader fetches a single h5ad rather than filtering and reassembling the cells on every call.

Parameters:
  • cache_dir (str | Path | None (default: None)) – Where to keep the download. Defaults to mantispy.settings.cache_dir.

  • aggregated (bool (default: False)) – Return one median profile per (Metadata_Gene, Metadata_sgRNA) guide instead of the cells, with Metadata_CellCount.

Returns:

Metadata_Gene: the gene the cell’s guide targets, with the non-targeting guides written as "nontargeting" (the upstream NTC), the spelling the analysis functions read.

Metadata_sgRNA: the guide identifier.

Metadata_Perturbation: the guide, so each guide is its own perturbation.

Metadata_Perturbation_Type: "crispr", the kind of screen this is.

Metadata_Control_Type: the schema’s reserved control-type column, carrying the upstream type, one of "target" (a screened gene), "ntc" (a non-targeting guide) or "neg" (a guide against an olfactory-receptor gene, a targeting negative control). The raw classes are kept rather than folded onto the reserved negcon/poscon/trt vocabulary, none of which fits the targeting negative cleanly.

Metadata_Control: True for the non-targeting guides, the reference hit_calling() and normalization test against. The olfactory-receptor negatives are not flagged, so they can be scored as perturbations that should not move.

Metadata_Plate: the plate, A or B.

Metadata_Well: the physical well, written as W03. The raw well is an integer that the well vocabulary cannot parse, so it is padded and prefixed. The ARV-471 arm sits in one well per plate, so this is constant.

The nine features are the ER and the two DAPI median intensities and the ESR1, CCND1 and GREB1 spot counts in the nucleus and the whole cell. The barcode, geometry (nucleus centers) and quality columns of the upstream table are dropped.

Return type:

AnnData

Notes

Cells whose segmentation touches the field boundary (Cells_Location_IntersectsBoundary_IF) are cut off, so their intensities and spot counts undercount, and are dropped. Cells missing any phenotype feature are dropped too, so every returned cell has a full feature vector.

A cell carries no count. The aggregated variant is one median profile per (Metadata_Gene, Metadata_sgRNA) guide, which writes Metadata_CellCount.

Raises:

ValueError – aggregated is not a bool.