Metadata-Version: 2.3
Name: frameworthy
Version: 0.1.6
Summary: Lightweight statistical validation library for data changes.
Author: joypauls
Author-email: joypauls <joypaulsen3@gmail.com>
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Requires-Dist: narwhals>=2.0
Requires-Dist: numpy>=2.2
Requires-Dist: scipy>=1.16
Requires-Python: >=3.11
Description-Content-Type: text/markdown

> ⚠️ WIP: All 0.1.x releases are experimental. 0.2.0 will be the first stable, production-ready release.

> Releases >0.1.5 are ready for testing, but the API is still subject to change.

# frameworthy

![PyPI Version](https://img.shields.io/pypi/v/frameworthy) 
[![PyPI pyversions](https://img.shields.io/pypi/pyversions/frameworthy.svg?x=1)](https://pypi.org/project/frameworthy/)
[![Tests](https://github.com/joypauls/frameworthy/actions/workflows/tests.yml/badge.svg)](https://github.com/joypauls/frameworthy/actions/workflows/tests.yml)
[![codecov](https://codecov.io/gh/joypauls/frameworthy/branch/main/graph/badge.svg?token=npu0JtY8hc)](https://codecov.io/gh/joypauls/frameworthy)

Frameworthy is a lightweight, dataframe-first library for statistically validating changes in data and metrics. Built for engineers and scientists who need reliable checks with statistical rigor, but are not looking to adopt a heavy platform. Works in data validation pipelines, in tests with assertions, or in exploratory settings with easily inspectable results.

<div align="center"><img src="docs/public/banner.png" width="600"></div>

## Getting Started

### Installation

Available on [PyPI](https://pypi.org/project/frameworthy/), use `pip` or your preferred package manager.

```bash
pip install frameworthy
# or
uv add frameworthy
```


### Examples

See `scripts/examples.py` for a few quick examples.

```python
import frameworthy as fw
import polars as pl # or pandas

# load your data if necessary
before_df = pl.read_csv("before.csv")
after_df = pl.read_csv("after.csv")

# run a check
result = (
    fw.check(after_df, before_df)
    .mean("column_name")
    .equivalent(within=0.1)
)

# inspect the results
print(result)
# or raise an exception on failure
result.assert_passed()
```

No data of your own yet? `check()` also accepts two numpy arrays directly, and `frameworthy` ships a couple of sample-data generators so you can try it with no extra steps:

```python
import frameworthy as fw

data = fw.sample_normal(shift=0.0)
result = fw.check(data.after, data.before).mean().equivalent(within=0.2)
print(result)
```


### Usage

Bring your data: two Pandas/Polars dataframes (before and after / pre and post). Also supports passing in numpy arrays directly.

A standard `frameworthy` check looks like this:

```python
fw.check(after_df, before_df).mean("column_name").equivalent(within=0.1)
```

1. Start a **check** with `fw.check(after_df, before_df)`
2. Specify the **metric** to check, with `.mean("column_name")`
3. Make a **claim** about the change, using `.equivalent()`, `.change_greater_than()`, or `.change_less_than()`
4. Examine the **results** or assert that the check passed with `.assert_passed()` for testing


## Methodology

These are the most important notes to be aware of for correct usage. For further details, see the more extensive [docs](https://joypauls.github.io/frameworthy/methodology).

### Confidence Intervals

The parameter `alpha` is a one-sided significance level everywhere; `.equivalent()`'s displayed CI is (1-2α), not (1-α). This is to maintain consistency with the TOST (Two One-Sided Tests) equivalence testing method.

Examples:
- `.equivalent()`
    - `alpha=0.05` → 90% CI 
    - `alpha=0.025` → 95% CI
- `.change_greater_than()` and `.change_less_than()`
    - `alpha=0.1` → 90% CI 
    - `alpha=0.05` → 95% CI

### Distribution Checks

`.distribution().equivalent()` doesn't use the BCa bootstrap that powers `.median()`/`.custom()`. The plug-in Wasserstein distance estimator is biased and, right where it matters most (two samples that are actually equivalent, so the true distance is 0 or close to it), the ordinary bootstrap is known to be unreliable at that boundary. Instead, it uses a subsampling/m-out-of-n bootstrap, which stays valid in that case.

Like `.mean()`/`.rate()`/`.median()`/`.custom()`, `.distribution()` supports both independent and paired samples (`paired_by=` across two dataframes, or `paired_column=` within a single dataframe). Pairing doesn't change the observed Wasserstein distance itself (it's a function of the two marginal distributions), but it does narrow the confidence interval when `before`/`after` are correlated, since the subsampling bootstrap then resamples matched pairs together instead of independently.


## Development

To run the tests:
```bash
make test
```

To run the examples:
```bash
uv run scripts/examples.py
```


## Feedback/Questions

Questions or feedback are welcome! Please open an [issue](https://github.com/joypauls/frameworthy/issues).


## License

MIT, but attribution is appreciated.
