Metadata-Version: 2.4
Name: heterodecomp
Version: 0.1.1
Summary: Shared and individual pattern decomposition for heterogeneous data
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: numpy>=1.23
Requires-Dist: openpyxl>=3.1
Requires-Dist: pandas>=1.5
Requires-Dist: pillow>=9
Requires-Dist: scipy>=1.9
Requires-Dist: tensorly>=0.8
Requires-Dist: torch>=2.0
Provides-Extra: eeg
Requires-Dist: mne>=1.6; extra == "eeg"
Provides-Extra: test
Requires-Dist: build>=1; extra == "test"
Requires-Dist: twine>=5; extra == "test"

# HeteroDecomp

**HeteroDecomp** is a Python toolkit for decomposing repeated heterogeneous observations into:

- **shared patterns** present across observations;
- **individual patterns** specific to each observation; and
- **residual/noise components** that are not explained by the fitted decomposition.

It provides one consistent loading, preprocessing, fitting, evaluation, and export interface for time-series data, general matrices, functional-connectivity (FC) matrices, and common image formats.

> **Development status:** Alpha (`0.1.1`). The API and algorithm adapters may evolve in future releases.

## Highlights

- A single public API for nine shared/individual decomposition algorithms.
- Automatic loading and auditing of CSV, TSV, TXT, Excel, NPY/NPZ, and common image files.
- Conservative header/index detection for tabular data, missing-value handling, constant-column checks, and data-quality reports.
- Support for unequal sequence lengths through explicit or recommended truncation.
- Automatic scaling appropriate to the detected data type, with explicit scaling options when needed.
- Metrics for reconstruction quality, residuals, shared/individual energy allocation, and component overlap.
- Ground-truth recovery metrics for simulated datasets.
- Reproducible results through an explicit random seed.
- Structured output folders containing component files, per-item metrics, summary metrics, data reports, and model metadata.
- Optional EEG/P300 helpers based on MNE-Python.

## Installation

### Install from PyPI

```bash
python -m pip install heterodecomp
```

### Install the optional EEG support

MNE-Python is optional. Install the extra when you need to read EEG FIF files or use the P300 preparation workflow:

```bash
python -m pip install "heterodecomp[eeg]"
```

### Install from a local checkout

```bash
python -m pip install .
```

For editable development installation:

```bash
python -m pip install -e ".[test]"
```

HeteroDecomp requires **Python 3.10 or newer**.

## Quick start

The main workflow consists of loading a collection of 2-D arrays and fitting one decomposition algorithm. Inputs use the convention **observations/time points in rows and features/channels/variables in columns**.

```python
from heterodecomp import decompose

result = decompose(
    "path/to/data",
    kind="timeseries",
    algorithm="perpca",
    shared_rank=5,
    individual_rank=2,
    params={"epochs": 200, "lr": 0.01},
    seed=0,
    output_dir="heterodecomp_results/perpca",
)

print(result.summary_metrics)
print(result.output_dir)
```

`decompose()` accepts a directory, a single file, a sequence of file paths, a sequence of NumPy arrays, or a previously loaded `LoadedData` object. Set `save=False` to fit and inspect a result without writing component files.

For more control over loading and data auditing:

```python
from heterodecomp import decompose, load_data

loaded = load_data(
    "path/to/data",
    kind="auto",
    pattern="*.csv",
    missing="mean",
    constant="warn",
)

result = decompose(
    loaded,
    algorithm="perpca",
    shared_rank=5,
    individual_rank=2,
    scaling="zscore",
    save=False,
)

print(loaded.report.summary())
print(result.subject_metrics.head())
```

## Supported algorithms

Use `available_algorithms()` to retrieve the canonical algorithm names accepted by `decompose()`.

| Name | Description | Typical use |
| --- | --- | --- |
| `perpca` | Personalized PCA with explicit shared and individual subspaces. | General-purpose default when interpretable subspaces are important. |
| `hmf` | Hierarchical matrix factorization with global and local factors. | High-quality reconstruction with flexible factors. |
| `jive` | Joint and Individual Variation Explained. | Classical joint/individual variation analysis. |
| `robust_jive` | Robust JIVE variant. | Data containing sparse outliers or corrupted entries. |
| `rajive` | Robust angle/resampling-based joint variation adapter. | More robust estimation of shared structure. |
| `slide` | Shared and individual latent decomposition. | Settings where structure selection is part of the analysis. |
| `pertucker` | Personalized Tucker decomposition. | Data with meaningful multilinear row/column structure. |
| `percdl` | Personalized one-dimensional convolutional dictionary learning. | Time-series signals and localized repeated motifs. |
| `robust_pca` | Low-rank plus sparse-error baseline adapted to shared/individual outputs. | Robust low-rank modeling with sparse errors. |

Algorithm-specific keyword arguments are passed through the `params` dictionary. See the scripts in [`examples/`](examples/) for complete, runnable configurations.

## Input data

### Tabular data

Supported extensions include `.csv`, `.tsv`, `.txt`, `.xlsx`, `.xls`, `.npy`, and `.npz`. Each input item must resolve to a non-empty 2-D numeric array, and all items must have the same number of columns/features.

```python
from heterodecomp import load_data

loaded = load_data(
    "data/",
    kind="timeseries",  # "auto", "timeseries", "matrix", or "image"
    pattern="*.csv",
    header="auto",
    index="auto",
    missing="mean",     # "auto", "error", "mean", "zero", or "drop"
)
```

The loader records detected headers and indices, missing and imputed values, dropped rows, constant columns, shape consistency, and (for square matrices) symmetry information in `loaded.report`.

### Images

Common raster images (`.png`, `.jpg`, `.jpeg`, `.bmp`, `.tif`, `.tiff`, and `.webp`) can be analyzed with `kind="image"`. Images in one analysis must have identical dimensions. Images are converted to grayscale by default.

### Functional-connectivity matrices

For square symmetric matrices, use `kind="matrix"` and usually `scaling="none"`. HeteroDecomp preserves the detected symmetric structure in exported shared and individual components.

### Unequal sequence lengths

If time-series items have different row counts, provide `target_length` explicitly, or leave it unset to use the shortest available length during fitting. The data report also includes a recommended truncation length when lengths differ.

## Scaling

The `scaling` argument accepts:

- `"auto"` (default): `zscore` for time series, `none` for matrices, and `range` for images;
- `"none"`: no scaling;
- `"zscore"`: center and scale each feature within each item;
- `"global"`: scale all items using one global standard deviation;
- `"range"`: use image-style range scaling.

The fitted components are restored to the original data scale before they are returned and saved.

## Outputs

When `save=True` (the default), results are written to `output_dir`, or to a generated directory based on the input source. The output contains:

```text
<output_dir>/
├── shared/                 # shared component for each input item
├── individual/             # individual component for each input item
├── reconstruction/         # reconstructed input for each input item
├── residual/               # input minus reconstruction
├── subject_metrics.csv     # per-item quality and energy metrics
├── summary_metrics.json    # aggregate metrics
├── data_report.json        # input audit report
└── model_metadata.json     # algorithm, ranks, seed, and model metadata
```

Tabular components are saved as CSV by default; set `output_format="xlsx"` for Excel output. Image analyses additionally save `exact_components.npz` so the numerical signed arrays remain available even when display images are clipped or converted to 8-bit formats.

Typical metrics include reconstruction explained variance, MSE/RMSE, Pearson correlation, shared and individual energy fractions, and shared/individual component overlap. Metric names and available fields are exposed through `result.subject_metrics` and `result.summary_metrics`.

## Ground-truth evaluation

For simulated data with known shared and individual components, call `evaluate_ground_truth()` after fitting:

```python
from heterodecomp import decompose, evaluate_ground_truth

result = decompose(
    "data/full_signal",
    kind="timeseries",
    algorithm="perpca",
    shared_rank=3,
    individual_rank=3,
    target_length=150,
    scaling="zscore",
)

metrics = evaluate_ground_truth(
    result,
    "data/shared_signal",
    "data/individual_signal",
)

print(metrics[["shared_cosine", "individual_cosine"]].mean())
```

The evaluation returns per-item MSE, RMSE, Pearson correlation, and cosine similarity for both shared and individual components. When the analysis has an output directory, it also writes `ground_truth_metrics.csv` and `ground_truth_summary.json`.

## EEG and P300 workflow

The optional P300 helpers prepare target/non-target ERP averages from MNE FIF files and validate correspondence between the fitted PerPCA shared basis and a P300 reference:

```python
from heterodecomp import prepare_p300, validate_p300

prepared = prepare_p300(
    "path/to/fif_epochs",
    output_path="prepared_p300.npz",
    target_event="Target",
    nontarget_event="NonTarget",
)

validation = validate_p300(
    prepared,
    shared_rank=5,
    individual_rank=2,
    p300_window=(0.3, 0.6),
    permutations=1000,
)

print(validation.pearson_r, validation.permutation_p)
```

Reading FIF files requires the optional dependency:

```bash
python -m pip install "heterodecomp[eeg]"
```

## Examples

The [`examples/`](examples/) directory contains algorithm-specific scripts for BOLD time series, simulated ground-truth evaluation, FC matrices, images, and real BOLD data without ground truth.

After installing the package, update the placeholder data paths at the top of an example and run it from the project root:

```bash
python examples/01_perpca.py
```

The example scripts use `max_items=10` by default for a quick trial where applicable. Set it to `None` to process all available items. For real data without known components, set both `true_shared_dir` and `true_individual_dir` to `None` so only no-ground-truth metrics are reported.

## Reproducibility and progress reporting

Set `seed` in `decompose()` and algorithm parameters explicitly when you need reproducible runs. Iterative adapters print human-readable progress and quality metrics by default where supported; pass `{"progress": False}` in `params` to silence algorithm progress output.

## Development and tests

Install the test dependencies and run the test suite with:

```bash
python -m pip install -e ".[test]"
python -m unittest discover -s tests -v
```

The fixture-based integration tests are skipped automatically unless the expected external test data directory is available.

## License

This release does not declare a license in its package metadata. Add the project’s intended license before publishing if redistribution under a specific license is required.
