mantispy.tl.ora

Contents

mantispy.tl.ora#

mantispy.tl.ora(adata, groupby='cluster', net=None, gene_key='Metadata_Gene', source='source', target='target', tmin=5, key_added='ora', copy=False, *, padj_by='all', min_overlap=1)[source]#

Test each group’s genes for over-representation of gene sets.

Parameters:
  • adata (AnnData) – Object whose obs carries the grouping and the gene symbols, for example after cluster().

  • groupby (str (default: 'cluster')) – obs column defining the groups whose genes are tested.

  • net (DataFrame | None (default: None)) – A gene-set network with source and target columns, such as mantispy.ds.gene_sets() returns.

  • gene_key (str (default: 'Metadata_Gene')) – obs column holding the gene symbol.

  • source (str (default: 'source')) – Column of net naming the set. Renamed to source internally.

  • target (str (default: 'target')) – Column of net naming the gene. Renamed to target internally.

  • tmin (int (default: 5)) – Smallest number of a set’s genes that must be in the universe for the set to be tested.

  • key_added (str (default: 'ora')) – Name for the output table.

  • copy (bool (default: False)) – Return a modified copy instead of mutating in place.

  • padj_by (Literal['all', 'group'] (default: 'all')) – Scope of the Benjamini-Hochberg correction. "all" (default) corrects once across every group and set. "group" corrects within each group’s tests, so a group’s modest enrichment is not penalized by unrelated groups (use this when many groups are tested at once). Because the scope is chosen per call, q-values from an "all" run and a "group" run are not directly comparable, so keep one scope within a single comparison.

  • min_overlap (int (default: 1)) – Smallest number of a group’s genes a set must contain to be tested for that group. 1 (default) drops only the sets a group’s genes do not hit at all (a == 0), so a set that shares no gene with the group is not tested and does not enlarge the correction denominator. The test is two-tailed, so a retained set may come out over- OR under-represented (read the sign of odds_ratio); min_overlap bounds only how many of the group’s genes a set must contain, not the direction of the result. Values >= 2 require stronger overlap before a set is tested.

Return type:

AnnData | None

Returns:

None, or the modified copy. Writes uns["mantispy"][key_added] with group, source (the set), n (the group’s genes in that set), odds_ratio (the Haldane-Anscombe log odds ratio), pvalue (a two-tailed Fisher exact test) and qvalue (Benjamini-Hochberg corrected), sorted by q. Only sets with at least min_overlap of a group’s genes appear in that group’s rows.

Raises:
  • KeyError – obs has no groupby or no gene_key.

  • ValueError – net is not given, or its source/target columns are missing.

Notes

The universe is the set of distinct genes in obs[gene_key], so a set is tested only on its genes that the screen measured, and sets with fewer than tmin measured genes are skipped. Each group and set is tested with a two-tailed Fisher exact test over that universe.

min_overlap shrinks the set of tested hypotheses to the sets a group’s genes actually hit (overlap >= min_overlap), rather than the whole collection. Sets a group does not touch are never tested, so they do not weigh on the correction under either scope: with padj_by="group" they stay out of that group’s own Benjamini-Hochberg family, and with padj_by="all" (default) they are absent from the single pooled family shared across groups.