mantispy.pp.feature_select

Contents

mantispy.pp.feature_select#

mantispy.pp.feature_select(adata, operations=('drop_degenerate', 'variance_threshold', 'correlation_threshold', 'drop_na_columns', 'blocklist'), min_variance=1e-06, freq_cut=0.05, unique_cut=0.01, corr_threshold=0.9, corr_method='pearson', corr_window=None, corr_stride=None, na_cutoff=0.05, outlier_cutoff=500.0, blocklist='default', noise_removal_perturb_groups='Metadata_Perturbation', noise_removal_stdev_cutoff=0.8, corr_absolute=False, corr_iterative=False, decorrelate=False, decorr_threshold=0.99, decorr_method='pearson', key_added='selected', copy=False)[source]#

Flag the features worth keeping.

Parameters:
  • adata (AnnData) – Object to select features on. Usually well-level profiles.

  • operations (Sequence[str] (default: ('drop_degenerate', 'variance_threshold', 'correlation_threshold', 'drop_na_columns', 'blocklist'))) – Which operations to run, from OPERATIONS. The default is pycytominer’s own, which omits frequency_threshold, drop_outliers and noise_removal, plus drop_degenerate, which removes nothing from an object normalize() did not flag.

  • min_variance (float (default: 1e-06)) – variance_threshold: keep features with variance above this.

  • freq_cut (float (default: 0.05)) – frequency_threshold: drop a feature when the count of its second most common value divided by the count of its most common is below this. Either this rule or unique_cut drops a feature.

  • unique_cut (float (default: 0.01)) – frequency_threshold: drop a feature when its share of distinct values is below this.

  • corr_threshold (float (default: 0.9)) – correlation_threshold: drop one member of every pair correlated above this.

  • corr_method (str (default: 'pearson')) – correlation_threshold: "pearson" or "spearman".

  • corr_window (int | None (default: None)) – correlation_threshold: None runs the exact pass (the default, matching pycytominer). An int switches to the two-pass fast path (prune redundancy within name-sorted windows of that size, then run the exact pass on the survivors); 500 is a good default, several times faster on large screens. It keeps a different set of features (a different member of each correlated group) but preserves the information and the downstream signal.

  • corr_stride (int | None (default: None)) – correlation_threshold: step between windows, default half the window.

  • corr_absolute (bool (default: False)) – correlation_threshold: threshold |r| rather than the signed correlation, so a strongly anti-correlated pair is also reduced. Off by default, matching pycytominer.

  • corr_iterative (bool (default: False)) – correlation_threshold: drop the most-connected feature one at a time until no pair is left, keeping at least as many features as the single pass. Off by default. With corr_absolute this approximates cytominer’s R path (caret::findCorrelation).

  • decorrelate (bool (default: False)) – Run an extra, experimental redundancy step after the operations, on the features they keep. Unlike correlation_threshold it removes features that are a linear combination of several others, not just pairwise duplicates, by a rank-revealing QR. Off by default, and not part of pycytominer.

  • decorr_threshold (float (default: 0.99)) – decorrelate: drop a feature once its multiple correlation with the kept set reaches this. The default 0.99 removes only near-collinear features, so the kept set spans almost the same space; lower it towards corr_threshold for a smaller, more aggressive set.

  • decorr_method (str (default: 'pearson')) – decorrelate: "pearson" or "spearman".

  • na_cutoff (float (default: 0.05)) – drop_na_columns: drop features missing in more than this fraction of rows.

  • outlier_cutoff (float (default: 500.0)) – drop_outliers: drop features whose absolute value exceeds this.

  • blocklist (str | Sequence[str] (default: 'default')) – blocklist: "default" for the bundled list, or explicit names. Matched against the current names and against var["original_name"], so it works either side of standardize_feature_names().

  • noise_removal_perturb_groups (str (default: 'Metadata_Perturbation')) – noise_removal: obs column grouping replicates.

  • noise_removal_stdev_cutoff (float (default: 0.8)) – noise_removal: drop features whose within-group standard deviation, averaged over groups, is above this. An absolute threshold on the scale normalize() left the values on, so it is only meaningful next to the normalization that produced them; pycytominer’s default assumes whole-plate standardization.

  • key_added (str (default: 'selected')) – Name of the boolean var column to write.

  • copy (bool (default: False)) – Return a modified copy instead of mutating in place.

Return type:

AnnData | None

Returns:

None, or the modified copy. Writes var[key_added] and a per-operation count of removals to uns["mantispy"]["feature_select"], each count being what that operation removes on its own. Nothing is dropped; use subset_features() for that.

Raises:
  • ValueError – If operations names an operation that is not in OPERATIONS.

  • KeyError – If noise_removal is requested but noise_removal_perturb_groups is not an obs column.

Notes

drop_degenerate runs first, and the other operations judge only the features it keeps: a feature normalize() could not scale can hold values large enough to decide the correlation ranking of every feature it is compared with. Every other operation judges that whole set, so each count in uns["mantispy"]["feature_select"] says what that operation alone would remove and is the same whatever order operations runs in. The counts therefore overlap: a feature that is both constant and mostly missing is counted by variance_threshold and by drop_na_columns, and the counts sum to more than the number of features actually removed, which is n_vars minus var[key_added].sum(). decorrelate is the exception: it runs last, on the features the operations kept, so its count is what it removes from those survivors.

correlation_threshold is the most expensive operation. pycytominer uses pandas.DataFrame.corr, one Cython pass per column pair. Here the pairs come from chunked matrix products over blocks of columns, so no n_vars ** 2 array is held in memory.