Metadata-Version: 2.4
Name: tadamz
Version: 1.0.0a16
Summary: package for targeted LC-MS(/MS) data analysis
Author: Patrick
Author-email: pkiefer@ethz.ch
License: MIT License
Classifier: Intended Audience :: Science/Research
Classifier: Intended Audience :: End Users/Desktop
Classifier: License :: OSI Approved :: MIT License
Classifier: Natural Language :: English
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Operating System :: OS Independent
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: emzed>=3.0.2
Requires-Dist: pyyaml
Requires-Dist: uncertainties
Requires-Dist: pybaselines==1.1.0
Requires-Dist: numpy>=2
Requires-Dist: seaborn
Requires-Dist: sklearn-migrator>=0.22.1
Dynamic: license-file

# tadamz

[![PyPI version](https://img.shields.io/badge/python-3.11%20%7C%203.12%20%7C%203.13-blue)](https://www.python.org/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)

**tadamz** is a specialized Python package for targeted LC-MS(/MS) (Liquid Chromatography-Mass Spectrometry) data analysis, built on top of the [emzed3](https://emzed.ethz.ch/) mass spectrometry framework. 

Developed at **ETH Zurich (Institute of Microbiology)**, `tadamz` provides a robust, modular, and highly configurable pipeline to extract, classify, normalize, calibrate, and quantify targeted compounds from MS1, MS/MS, or SRM chromatograms.

---

## 🚀 Key Features

*   **📊 Peak Extraction & Integration:** High-performance chromatography integration, baseline subtraction (via `pybaselines`), and extraction of chromatogram peaks based on targeted precursor $m/z$ and retention times.
*   **🤖 Peak Classification:** Integrates a Random Forest classifier (`scikit-learn`) to score peak quality (`peak_quality_score`), filtering out noise or false positives based on peak metrics (symmetry, FWHM, peak-to-noise ratio, etc.).
*   **🔄 RT Adaptation & Co-elution Alignment:** Corrects and adapts target retention times dynamically across batches and alignments using reference co-eluting compound peaks.
*   **⚖️ Calibration & Absolute Quantification:** Fits weighted (e.g., $1/x$, $1/x^2$, $1/s^2$) linear or quadratic calibration models to calibrant tables to achieve accurate absolute concentration calculations.
*   **📉 Normalization Options:** Includes sample-wise normalization, standard-free normalization, and natural abundance/isotopologue overlay correction for stable isotope labeling experiments.
*   **📝 Automated Report Generation:** Automatically generates comprehensive PDF reports featuring global abundance heatmaps and compound-wise acquisition-order trends (for retention times, FWHM, and abundance).
*   **💻 Interactive emzed GUI Integration:** Built-in tools for interactive table exploration, manual peak scoring, and data visualization via `emzed-spyder`.

---

## 📦 Installation

`tadamz` requires Python **>= 3.11** and the core `emzed` framework.

You can install `tadamz` directly via `pip` from the repository:

```bash
pip install git+https://gitlab.com/emzed3_extensions/targeted.git
```

This will automatically install dependencies, including:
*   `emzed >= 3.0.2`
*   `pyyaml`, `uncertainties`, `pybaselines == 1.1.0`
*   `numpy >= 2`, `seaborn`, `sklearn-migrator >= 0.22.1`

---

## 🛠️ Getting Started

### 1. Configuration (YAML)

`tadamz` workflows are driven by configuration dictionaries, typically stored as YAML files:

```yaml
extract_peaks:
  integration_algorithm: linear
  ms_data_type: MS_Chromatogram
  mz_tol_abs: 0.3
  peak_search_window_size: 60
  subtract_baseline: true

classify_peaks:
  scoring_model: random_forest_classification
  scoring_model_params:
    classifier_name: srm_peak_classifier

normalize_peaks:
  sample_wise: true
  correct_int_std_for_nat_abundance: false

processing_steps:
  - extract_peaks
  - classify_peaks

postprocessings:
  - postprocessing1
  - postprocessing2

postprocessing1:
  - classify_peaks
  - coeluting_peaks
  - normalize_peaks
```

### 2. Workflow Execution

Load targets, samples, and config, and execute the pipeline:

```python
from tadamz import run_workflow, postprocess_result_table, load_config
from tadamz.in_out import load_targets_table, load_samples_from_folder

# Load inputs
config = load_config("path/to/config.yaml")
targets_table = load_targets_table("path/to/targets_table.xlsx")
samples = load_samples_from_folder("path/to/raw_data_folder")

# Run main workflow
result = run_workflow(targets_table, samples, config)

# Post-process (e.g., RT coelution alignment & peak normalization)
result = postprocess_result_table(result, config, postprocess_id=0)
```

---

## 🤖 Training a custom Peak Classifier

Train custom peak classifiers for your targeted assays using manually scored peaks:

```python
from tadamz import generate_peak_classifier

generate_peak_classifier(
    classifier_name="my_custom_srm_classifier",
    path_to_folder="path/to/save/classifier",
    path_to_table="path/to/measured_peaks.table",
    ms_data_type="MS_Chromatogram",
    inspect=True, # Opens interactive GUI for manual scoring
)
```

---

## 📊 Generating Reports

`tadamz` can create comprehensive PDF reports of your targeted analysis:

```python
from tadamz.generate_report import generate_report

# Generate global, compound-wise, and calibration report figures. Calibration
# plots are derived from the calibration_model column in result_table.
global_plots, compound_plots, calibration_plots = generate_report(
    result_table,
    pdf_folder="path/to/output_folder",
    # Optional; omit when no calibration was performed. Calibration plots use
    # column names such as calibrate.value_col from the full configuration.
    config=config,
)

# Calibration plots are grouped by compound, then plot name.
curve = calibration_plots["compound_name"]["calibration_curves"]
```

---

## 📂 Project Structure

```text
tadamz/
├── src/
│   └── tadamz/                  # Main package source code
│       ├── calibration/         # Regression models & absolute calibration curve fitting
│       ├── data/                # Default configuration files & sample datasets
│       ├── scoring/             # Peak quality metrics & Random Forest classifiers
│       ├── in_out.py            # IO helpers for tables, YAML, JSON, and PKL files
│       ├── workflow.py          # Main workflow definition and runner functions
│       └── generate_report.py   # PDF report generator using matplotlib and seaborn
├── tests/                       # Unit tests & regression tests (run via pytest)
├── setup.cfg                    # Metadata, classifiers, and project dependencies
└── pyproject.toml               # Build system configuration
```

---

## 📄 License & Credits

*   **Author:** Patrick Kiefer ([pkiefer@ethz.ch](mailto:pkiefer@ethz.ch))
*   **Organization:** Institute of Microbiology, ETH Zurich
*   **License:** MIT License
