Metadata-Version: 2.5
Name: bertram-amla-nba
Version: 1.0.0
Summary: Helper utilities for EDA and preprocessing of binary classification tasks on tabular data.
Project-URL: Homepage, https://github.com/bertramhage/36120-26SP-group3-26666281-package
Project-URL: Repository, https://github.com/bertramhage/36120-26SP-group3-26666281-package
Project-URL: Issues, https://github.com/bertramhage/36120-26SP-group3-26666281-package/issues
Author-email: Bertram Hage <bertram.hage@gmail.com>
License-Expression: MIT
License-File: LICENSE
Keywords: eda,machine-learning,pandas,preprocessing,tabular
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.12
Requires-Dist: matplotlib>=3.9.2
Requires-Dist: pandas>=2.2.2
Requires-Dist: scikit-learn>=1.6.1
Description-Content-Type: text/markdown

# bertram-amla-nba

A helper package for binary classification ML-workflows on tabular data. It aids the
EDA, preprocessing and data transformation operations used binary classification ML tasks
so they can be reused across notebooks instead of being copy-pasted.

The package is split into two modules:

- `eda` — target/feature exploration plots and a per-feature summary table
- `transformation` — preprocessing steps that learn on train and can be reused on test

## Installation

```bash
pip install bertram-amla-nba
```

or with uv:

```bash
uv add bertram-amla-nba
```

Requires Python 3.12 or newer.

## Quick start

```python
import pandas as pd
from bertram_amla_nba import eda, transformation
```

### EDA

```python
eda.explore_target_distribution(train, target_column="label")
eda.explore_feature(train, "pts", target_column="label")
eda.explore_feature(train, "position", target_column="label", categorical=True)

summary = eda.feature_summary(train, target_column="label")
```

### Transformation

Each transformer returns `(dataframe, preprocessor)`. You can call it without a preprocessor on
the training set to learn the parameters, then pass that preprocessor back when applying on the test set.

```python
from bertram_amla_nba.transformation import (
    feet_inches_to_inches,
    remove_outliers,
    impute_values,
    encode_categorical,
)

# "6-2" -> 74.0
train["height_in"] = feet_inches_to_inches(train["height"])
test["height_in"] = feet_inches_to_inches(test["height"])

# Fit on train
train, outlier_p = remove_outliers(train, target_column="label")
train, impute_p = impute_values(train, target_column="label")
train, encoder = encode_categorical(train, columns=["position"])

# Replay on test
test, _ = remove_outliers(test, preprocessor=outlier_p)
test, _ = impute_values(test, preprocessor=impute_p)
test, _ = encode_categorical(test, columns=["position"], preprocessor=encoder)
```

- `remove_outliers` replaces values outside the IQR bounds with the column median
- `impute_values` fills missing values with the median (numeric) or mode (categorical)
- `encode_categorical` one-hot encodes the given columns and passes the rest through

## Development

```bash
uv sync
uv run pytest
```

## License

MIT
