Metadata-Version: 2.5
Name: cpdrift
Version: 1.0.0
Summary: Inductive evaluation and representation-drift measurement for batch correction in image-based profiling
Project-URL: Homepage, https://github.com/SamirHossain099/cpdrift
Project-URL: Archive, https://doi.org/10.5281/zenodo.22548754
Project-URL: Issues, https://github.com/SamirHossain099/cpdrift/issues
Author: Samir Hossain
License: MIT
License-File: LICENSE
Keywords: Cell Painting,JUMP,batch correction,image-based profiling,inductive evaluation,morphological profiling,representation drift,reproducibility
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Classifier: Topic :: Scientific/Engineering :: Image Processing
Requires-Python: >=3.10
Requires-Dist: numpy>=1.22
Requires-Dist: pandas>=1.5
Provides-Extra: analysis
Requires-Dist: pyarrow>=10; extra == 'analysis'
Requires-Dist: requests>=2.28; extra == 'analysis'
Requires-Dist: scikit-learn>=1.1; extra == 'analysis'
Requires-Dist: scipy>=1.9; extra == 'analysis'
Provides-Extra: dev
Requires-Dist: pytest>=7; extra == 'dev'
Requires-Dist: ruff>=0.4; extra == 'dev'
Provides-Extra: harmony
Requires-Dist: harmonypy>=0.0.10; extra == 'harmony'
Provides-Extra: plot
Requires-Dist: matplotlib>=3.6; extra == 'plot'
Provides-Extra: sphering
Requires-Dist: pycytominer>=1.7; extra == 'sphering'
Description-Content-Type: text/markdown

# cpdrift

Inductive evaluation and **representation drift** for batch correction in image-based profiling.

Batch-correction methods are usually benchmarked *transductively*: fitted on every batch, including
the ones they are scored on. Production screens are not like that. Plates accumulate over months,
and a correction has to be applied to batches that did not exist when it was fitted. `cpdrift`
provides the pieces needed to evaluate a correction the way it will actually be deployed.

```bash
pip install cpdrift
```

numpy and pandas only. MIT licensed.

[![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.22548754.svg)](https://doi.org/10.5281/zenodo.22548754)

## Citing

Archived at Zenodo under the concept DOI [10.5281/zenodo.22548754](https://doi.org/10.5281/zenodo.22548754), which
always resolves to the latest release. Cite that one rather than a version DOI, so the
citation does not go stale.

## Why it exists

Two methods sitting side by side in the same benchmark can have categorically different deployment
semantics, and the usual evaluation protocol does not distinguish them:

- **Fittable** methods expose a `fit` / `transform` split. Fit once, apply forward, and profiles you
  already analysed never move.
- **Refit-only** methods offer a single whole-matrix entry point. The only way to add a batch is to
  refit everything, which silently rewrites earlier results.

`harmonypy` is refit-only in every released version (0.0.9 through 2.0.0): no `transform`,
`predict`, `apply` or `fit_transform` has ever existed in the published API. `pycytominer.Spherize`
is a `BaseEstimator` with both. That asymmetry is what this package measures.

## Representation drift

The naive measure -- displacement between two corrected embeddings -- is wrong, because several
methods are defined only up to a rotation of the feature space. A pure basis change would register
as large drift while changing no downstream conclusion. So `cpdrift` gauge-fixes first:

1. Fit an orthogonal Procrustes rotation on **control** wells, whose biology is unchanged by
   construction.
2. Measure displacement on **treated** wells only.
3. Divide by the assay's own within-batch replicate distance.

`D~ = 1` means "already-analysed profiles moved as far as two replicates of the same perturbation
differ from each other". Uncorrected input gives exactly zero, and that is a unit test.

```python
import numpy as np
from cpdrift import Sphering, normalised_drift, replicate_distance

# fit on the batches you have, apply forward to the ones that arrive later
op = Sphering().fit(Z_anchor)
before = op.transform(Z_anchor)
...
after = op.transform(Z_anchor)          # no refit, so nothing moved

d = normalised_drift(before, after, is_control,
                     replicate_distance(before, pert_ids, batch=batch))
assert d < 1e-12
```

Without the gauge fix, a rotation alone scores **1.117** on this scale -- roughly twice the
replicate distance. Naive measurement manufactures the result it is looking for.

## Evaluating both axes

A method can climb a batch-mixing score by destroying biological signal, so a mixing number quoted
alone cannot detect that. Report both:

```python
from cpdrift import batch_mixing_score, mean_average_precision

mix = batch_mixing_score(Z, batch)              # batch homogeneity
mAP, _ = mean_average_precision(Z, pert_ids, batch=batch)   # biological retrieval
```

`batch_mixing_score` is a deliberately simple k-NN score. For a published implementation use
`scib`'s iLISI; the two rank methods consistently (per-source Spearman +0.77 to +1.00 in our
testing) but only the ordering is comparable, since they differ in neighbourhood size and distance.

## API

| name | what it does |
|---|---|
| `FITTABLE`, `REFIT_ONLY`, `REGISTRY` | the deployment-semantics taxonomy |
| `NoCorrection`, `MeanCenter`, `ComBat`, `Sphering` | fittable corrections, `fit` / `transform` |
| `Harmony` | refit-only; raises if you call `transform` |
| `NullDMSO` | the signal-free null model, for auditing whether a metric rewards degenerate solutions |
| `procrustes_rotation`, `representation_drift`, `normalised_drift`, `replicate_distance` | the drift metric |
| `mean_average_precision`, `average_precision`, `batch_mixing_score` | evaluation metrics |

`Sphering` reproduces `pycytominer.Spherize` to <=1e-6, including its choice to regularise the
singular values rather than the eigenvalues, but avoids its `np.linalg.svd(..., full_matrices=True)`
call -- which materialises an (n x n) matrix it never reads, 53.5 GB at 81,740 wells.

That equivalence is verified against **pycytominer 1.7 and later**, which is what the `sphering`
extra pins. Older releases give a different result, so the comparison test skips rather than fails
on them.

`Harmony` needs `harmonypy`, imported lazily and only when called:

```bash
pip install "cpdrift[harmony]"
```

## Note on threads

`cpdrift` deliberately does **not** set BLAS thread limits. Capping threads is an application
decision, and a library that mutates process-wide environment variables at import time is a hazard.
Set `OMP_NUM_THREADS` and friends yourself, before importing numpy.

## Licence

MIT.
