Metadata-Version: 2.4
Name: sparsepr
Version: 0.1.0
Summary: Training-free sparse-attention acceleration for video diffusion and world models
Author: Pardis Taghavi
License-Expression: Apache-2.0
Project-URL: Documentation, https://github.com/PardisTaghavi/sparsepr-python#readme
Project-URL: Issues, https://github.com/PardisTaghavi/sparsepr-python/issues
Project-URL: Repository, https://github.com/PardisTaghavi/sparsepr-python
Keywords: attention,cuda,diffusion,sparse-attention,video-generation
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
License-File: THIRD_PARTY_NOTICES.md
Provides-Extra: hunyuan
Requires-Dist: accelerate<2,>=1.0; extra == "hunyuan"
Requires-Dist: diffusers<0.36,>=0.35; extra == "hunyuan"
Requires-Dist: flashinfer-python==0.6.2; extra == "hunyuan"
Requires-Dist: torch<2.8,>=2.7; extra == "hunyuan"
Requires-Dist: transformers<4.52,>=4.51; extra == "hunyuan"
Requires-Dist: triton<3.4,>=3.3; extra == "hunyuan"
Provides-Extra: cosmos3
Requires-Dist: accelerate<2,>=1.0; extra == "cosmos3"
Requires-Dist: diffusers<0.40,>=0.39; extra == "cosmos3"
Requires-Dist: flashinfer-python==0.6.2; extra == "cosmos3"
Requires-Dist: torch<2.8,>=2.7.1; extra == "cosmos3"
Requires-Dist: transformers<4.58,>=4.57; extra == "cosmos3"
Requires-Dist: triton<3.4,>=3.3; extra == "cosmos3"
Provides-Extra: dev
Requires-Dist: build>=1.2; extra == "dev"
Requires-Dist: pytest>=8; extra == "dev"
Requires-Dist: ruff>=0.8; extra == "dev"
Requires-Dist: twine>=6; extra == "dev"
Dynamic: license-file

# SparsePR

SparsePR is a training-free sparse-attention accelerator for video diffusion
and world models. It attaches to an existing pipeline instance, preserves that
pipeline's inference interface, and can restore the original attention
processors exactly.

> **Status:** alpha library extraction. Cosmos3 and HunyuanVideo are the first
> public adapters. Wan2.2 and Cosmos-Predict2.5 are intentionally not exposed
> until their integrations are instance-local and reversible.

## Design contract

- `import sparsepr` does not import PyTorch, Triton, Diffusers, or model code.
- Acceleration state belongs to one model instance.
- Every successful attachment is reversible.
- Unsupported models produce an adapter-by-adapter diagnostic.
- SparsePR does not replace the upstream pipeline's generation interface.

## Installation

The lightweight API and adapter registry install with:

```bash
python -m pip install sparsepr
```

Install one model environment at a time because the validated upstream
Diffusers versions differ:

```bash
python -m pip install "sparsepr[hunyuan]"
# or, in a separate environment
python -m pip install "sparsepr[cosmos3]"
```

GPU inference requires Linux and a CUDA-enabled PyTorch installation. H100 or
A100-80GB is recommended for supported full-model inference. Install PyTorch
from the index matching the host CUDA runtime before installing a SparsePR
extra.

## Usage

SparsePR returns a handle that delegates calls and attribute reads to the
original pipeline:

```python
import sparsepr

handle = sparsepr.accelerate(
    pipe,
    sparsepr.Config.balanced(),
    num_inference_steps=35,
    guidance_scale=6.0,
)

result = handle(prompt=prompt, image=image, num_inference_steps=35)
print(handle.stats().to_dict())
pipe = handle.detach()
```

The original pipeline may also remain the call target:

```python
handle = sparsepr.accelerate(pipe, num_inference_steps=35)
result = pipe(prompt=prompt, image=image, num_inference_steps=35)
sparsepr.detach(pipe)
```

Attachments can be scoped:

```python
with sparsepr.accelerate(pipe, num_inference_steps=35) as accelerated_pipe:
    result = accelerated_pipe(prompt=prompt, image=image)
# original attention processors have been restored
```

HunyuanVideo currently needs the prompt and output geometry when attaching,
because its sparse layout includes the encoded prompt length:

```python
handle = sparsepr.accelerate(
    pipe,
    prompt=prompt,
    height=720,
    width=1280,
    num_frames=129,
    num_inference_steps=50,
)
```

## Public API

- `sparsepr.accelerate(model, config=None, adapter=None, **context)`
- `sparsepr.detach(model_or_handle)`
- `sparsepr.is_supported(model, adapter=None)`
- `sparsepr.is_accelerated(model_or_handle)`
- `sparsepr.stats(model_or_handle)`
- `sparsepr.reset_generation_state(model_or_handle)`
- `sparsepr.list_adapters()`
- `sparsepr.register_adapter(adapter)`

See [`docs/ADAPTERS.md`](docs/ADAPTERS.md) for the adapter contract and
[`docs/RELEASING.md`](docs/RELEASING.md) for the PyPI checklist.

## Development

```bash
python -m pip install -e ".[dev]"
pytest -q
ruff check src tests
python -m build
twine check dist/*
```

## License

Apache-2.0. See `THIRD_PARTY_NOTICES.md` for incorporated components.
