Metadata-Version: 2.4
Name: amr-clonalshare
Version: 1.0.0
Summary: Observational summaries of antimicrobial phenotypes by recorded lineage, with explicit phenotype definitions, input diagnostics, collection bounds and model-specific uncertainty.
Author-email: Maciej Kochanowski <maciej.kochanowski@piwet.pulawy.pl>
License-Expression: MIT
Project-URL: Homepage, https://github.com/maciejkochanowski/amr-clonalshare
Project-URL: Documentation, https://github.com/maciejkochanowski/amr-clonalshare/tree/main/docs
Project-URL: Repository, https://github.com/maciejkochanowski/amr-clonalshare
Project-URL: Issues, https://github.com/maciejkochanowski/amr-clonalshare/issues
Project-URL: Changelog, https://github.com/maciejkochanowski/amr-clonalshare/blob/main/CHANGELOG.md
Keywords: bioinformatics,antimicrobial-resistance,veterinary-surveillance,variance-component,population-structure,typing-resolution,intraclass-correlation,prevalence-decomposition,kitagawa-decomposition,interval-censored-mic,anytime-valid-inference,streptococcus-suis,serovar,ncbi-pathogen-detection,reproducible-research
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Classifier: Topic :: Scientific/Engineering :: Medical Science Apps.
Classifier: Intended Audience :: Healthcare Industry
Classifier: Operating System :: OS Independent
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
License-File: Licence.txt
Requires-Dist: numpy>=2.0
Requires-Dist: pandas>=2.2
Requires-Dist: scipy>=1.13
Requires-Dist: pyyaml>=6.0
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: pytest-cov>=4.0; extra == "dev"
Requires-Dist: hypothesis>=6.0; extra == "dev"
Requires-Dist: scikit-learn>=1.5; extra == "dev"
Requires-Dist: ruff>=0.6; extra == "dev"
Requires-Dist: mypy>=1.11; extra == "dev"
Requires-Dist: types-PyYAML>=6.0; extra == "dev"
Requires-Dist: build>=1.0; extra == "dev"
Provides-Extra: docs
Requires-Dist: mkdocs-material==9.5.39; extra == "docs"
Requires-Dist: mkdocstrings[python]==1.0.6; extra == "docs"
Dynamic: license-file

# amr-clonalshare

## Version 1.0.0

This release adds a matched-record comparison of two lineage definitions and a calibrated 95% interval for the lineage share of MIC variation. For exact MIC readings the interval inverts Wald's generalised F pivot. For dilution intervals it inverts a likelihood-ratio test calibrated by a null-wise parametric bootstrap. In prespecified simulations both kept nominal coverage in all 72 main Gaussian designs; one of 16 additional Gaussian designs fell just below (92.6%, Wilson 90.0–94.6%). The earlier moment estimate and its approximate F interval stay in the output as a labelled comparison. The simulation drivers and results are in `benchmarks/mic_inference/` and `benchmarks/results_mic_release/`.

Describe how recorded antimicrobial phenotypes vary across recorded lineages.
The package reads CSV tables of isolate identifiers, lineage labels and susceptibility calls or minimum inhibitory concentrations (MICs). It does not read genomes. Laboratory analysts can use the command line with a YAML configuration; bioinformaticians can use the same configuration through Python.

**Software version: 1.0.0.** Cite the version-specific Zenodo record; the [concept DOI](https://doi.org/10.5281/zenodo.22306353) groups all versions.

## Install and run

From PyPI or from this source directory:

```bash
python -m pip install amr-clonalshare   # or: python -m pip install .
amr-clonalshare --config examples/workflows/calls.yaml --check-input
amr-clonalshare --config examples/workflows/calls.yaml --results-dir out/calls
```

Python 3.11 or later and NumPy, pandas, SciPy and PyYAML are required. A container image with the pinned stack of the shipped records is published with every release as `ghcr.io/maciejkochanowski/amr-clonalshare:1.0.0` (see the [reproduction guide](docs/manual/08-container.md)).

Start with the [four executable recipes](examples/workflows/README.md): calls, MIC, two collections, and the optional population model. Their small fictional data teach the file format; they are not biological evidence. The [input manual](docs/manual/02-input.md) explains how to substitute your own tables.

## Choose a question

| User question | Route | Interpretation |
|---|---|---|
| How strongly do the recorded labels describe the recorded calls? | Calls recipe | Observed-scale lineage-membership share, support and uncertainty |
| What do recorded MIC intervals show at this label resolution? | MIC recipe | Dilution-scale variance component with a calibrated 95% interval |
| How do two recorded collections differ? | Contrast recipe | Descriptive lineage-composition and within-lineage rate terms |
| What does a Gaussian lineage population model imply? | Population recipe | Separate latent-liability ICC and explicit model assumptions |

Clinical S/I/R, WT/NWT and binary data have separate `phenotype_kind` settings. New configurations should state the positive outcome, source and any applicable AST standard/version. Clinical S/I/R defaults to R only; I is not silently treated as resistant. Legacy configurations remain readable for explicit historical replay.

## Read the result bundle

A completed run saves `clonal_share_result.json`, a CSV summary, Markdown and offline HTML reports, input-QC records and a completion manifest. JSON is the canonical analytical record; the other formats render its values. Inspect the manifest before treating a directory as a complete run. Existing output is protected; explicit `--overwrite` preserves a backup. Use a new output directory for a comparison run.

Read the phenotype definition, each method's retained cohort, method status and reasons before comparing estimates. Small or incomplete data can support a description even when a particular interval is withheld. A missing method is not a zero estimate.

Missingness checks describe observed associations. A nonsignificant check or equal typing fractions cannot establish representativeness. When labels are missing, a supported decomposition describes the typed subset; collection generalization remains unsupported. Exact finite-collection bounds show what missing binary outcomes could change within the recorded frame; these are not confidence intervals and do not extend to unrecorded isolates.

## Python using the same configuration

```python
from amr_clonalshare import load_config, run

config = load_config("examples/workflows/calls.yaml")
result = run(config, results_dir="out/calls_api", seed=20260913)
```

See the [API guide](docs/api.md) for result reading and the [results manual](docs/manual/05-results.md) for schema 2.0. The reader also accepts historical schema 1.0 records without manufacturing current fields.

## Evidence and limits

The empirical examples include a 677-isolate *Streptococcus suis* collection and a 7,049-isolate poultry-meat *Salmonella* input cell. Their original provenance, release-specific results, exclusions and negative findings remain in the examples and separately identified evidence archive. The MIC interval calibration and the S. suis reanalysis for this release are in `benchmarks/results_mic_release`. [REPRODUCIBILITY.md](REPRODUCIBILITY.md) explains the distinction.

Lineage association does not establish transmission, a resistance mechanism, intervention benefit or clinical utility. Changing the lineage definition changes the question. Population-model intervals depend on their stated assumptions, and passing a diagnostic does not certify those assumptions. The optional general Gaussian-probit route is selected explicitly with `interval_method: general`; it does not change the default. Read its [validation and resource guidance](docs/GENERAL_POPULATION_ICC.md).

## Documentation and licence

[User manual](docs/index.md) · [Methods](docs/methodology.md) · [API](docs/api.md) · [Changes](CHANGELOG.md) · [Code metadata](docs/CODE_METADATA.md)

Code is distributed under the MIT licence; see `LICENSE` and `Licence.txt`. Example-data provenance and licences are stated alongside each collection.

## Compare two lineage definitions on matched records

```bash
amr-clonalshare-compare --input examples/matched_lineages/input.csv \
  --id-column isolate_id --outcome positive --lineage-a lineage_a \
  --lineage-b lineage_b --output out/lineage_comparison
```

The four arms distinguish record restriction from label changes. The two common-frame arms have the same IDs and outcomes; unsupported arms remain diagnostics. Arm-wise bootstrap intervals are not intervals for paired differences. The example is fictional and teaches the format.

## Interval for one MIC table

```bash
amr-clonalshare-mic readings.csv --method bootstrap --seed 1 \
  --panel-edges "[-3,-2,-1,0,1,2,3]" --workers 8 --output rho.json
```

The input has columns `lo`, `hi` (log2 concentration bounds; equal for an exact reading) and `lineage`. When the records were read on more than one panel, name the column that says which with `--panel-column` and pass `--panel-edges` as a JSON object, one array of cut points per panel name; `--covariate-column` names a column whose levels (a laboratory, a country) enter the model as fixed effects. Use `--method exact` when every reading is exact. The main pipeline computes the same interval for every agent unless `censored.calibrated_interval: false`; `censored.workers` sets the number of processes and does not change the result. One interval for about 700 isolates takes several minutes on 16 cores.
