Metadata-Version: 2.4
Name: papyrus_structure_pipeline
Version: 0.1.0
Summary: Papyrus Structure Pipeline
Author-email: "Olivier J. M. Béquignon" <olivier.bequignon.maintainer@gmail.com>
Maintainer-email: "Olivier J. M. Béquignon" <olivier.bequignon.maintainer@gmail.com>
License-Expression: MIT
Project-URL: Homepage, https://github.com/OlivierBeq/Papyrus_structure_pipeline
Project-URL: Repository, https://github.com/OlivierBeq/Papyrus_structure_pipeline
Project-URL: Issues, https://github.com/OlivierBeq/Papyrus_structure_pipeline/issues
Keywords: cheminformatics,molecule,standardization
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Science/Research
Classifier: Topic :: Scientific/Engineering :: Chemistry
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Typing :: Typed
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: rdkit
Requires-Dist: chembl-structure-pipeline
Provides-Extra: docs
Requires-Dist: sphinx; extra == "docs"
Requires-Dist: sphinx-rtd-theme; extra == "docs"
Requires-Dist: sphinx-autodoc-typehints; extra == "docs"
Provides-Extra: testing
Requires-Dist: pytest; extra == "testing"
Requires-Dist: pytest-cov; extra == "testing"
Dynamic: license-file

# 📜 Papyrus Structure Pipeline

<!-- Badges -->
<div align="center">

[![PyPI version](https://img.shields.io/pypi/v/papyrus-structure-pipeline.svg)](https://pypi.org/project/papyrus-structure-pipeline/)
[![Supported Python versions](https://img.shields.io/pypi/pyversions/papyrus-structure-pipeline.svg)](https://pypi.org/project/papyrus-structure-pipeline/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![Tests](https://github.com/OlivierBeq/Papyrus_structure_pipeline/actions/workflows/tests.yml/badge.svg)](https://github.com/OlivierBeq/Papyrus_structure_pipeline/actions/workflows/tests.yml)
[![Ruff](https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ruff/main/assets/badge/v2.json)](https://github.com/astral-sh/ruff)
<br>
</div>

A small, opinionated molecule standardization pipeline built on top of the [ChEMBL Structure Pipeline](https://github.com/chembl/ChEMBL_Structure_Pipeline) and [RDKit](https://www.rdkit.org/). It turns arbitrary input structures into a consistent parent form suitable for bioactivity datasets: salts and metals stripped, charges neutralized, tautomers canonicalized, and mixtures/inorganics/out-of-range molecular weights filtered out. First used to curate the [Papyrus](https://doi.org/10.5281/zenodo.7377161) bioactivity dataset.

## ✨ Features

- 🧪 **ChEMBL-based standardization** — wraps the ChEMBL Structure Pipeline's parent-structure extraction and normalization, run twice (before and after tautomer canonicalization) for a consistent final form.
- 🧂 **Configurable salt & metal stripping** — removes a user-extensible set of salts/metals beyond what the ChEMBL pipeline already handles.
- ⚛️ **Charge neutralization** — uncharges ionizable groups and zwitterions.
- 🔀 **Bounded tautomer canonicalization** — deterministic canonical tautomer picking with a capped search, so runtime stays predictable even for polyphenol-like molecules with many tautomerizable sites.
- 🧬 **Organic/inorganic filtering** — configurable allowed-atom list and C-C bond requirement.
- 🧩 **Mixture & size filtering** — drop multi-fragment mixtures and molecules outside a molecular-weight window.
- 🔍 **Rich failure reporting** — `return_type=True` reports *why* a molecule was rejected (mixture, inorganic, too small/large, standardization error) instead of just returning `None`.
- 🏷️ **Typed API** — ships `py.typed`, fully type-hinted, mypy-clean.

### Citing

If you use this package, please cite the Papyrus dataset it was first used in:

[![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.7377161.svg)](https://doi.org/10.5281/zenodo.7377161)

## 📦 Installation

```bash
pip install papyrus-structure-pipeline
```

Or from source:

```bash
git clone https://github.com/OlivierBeq/Papyrus_structure_pipeline.git
pip install ./Papyrus_structure_pipeline
```

## 🛠️ Requirements

- Python 3.11+
- [RDKit](https://www.rdkit.org/docs/Install.html)
- [ChEMBL Structure Pipeline](https://github.com/chembl/ChEMBL_Structure_Pipeline) (installed automatically)

## 💡 Usage

### Standardize a compound

Comparison to the [ChEMBL Structure Pipeline](https://github.com/chembl/ChEMBL_Structure_Pipeline):

```python
from rdkit import Chem
from chembl_structure_pipeline import standardizer as ChEMBL_standardizer
from papyrus_structure_pipeline import standardizer as Papyrus_standardizer

# CHEMBL1560279
smiles = "CCN(CC)C(=O)[n+]1ccc(OC)cc1.c1ccc([B-](c2ccccc2)(c2ccccc2)c2ccccc2)cc1"

mol = Chem.MolFromSmiles(smiles)
out1 = ChEMBL_standardizer.standardize_mol(mol)
out2 = Papyrus_standardizer.standardize(mol)

print(Chem.MolToSmiles(out1))
# CCN(CC)C(=O)[n+]1ccc(OC)cc1.c1ccc([B-](c2ccccc2)(c2ccccc2)c2ccccc2)cc1

print(Chem.MolToSmiles(out2))
# CCN(CC)C(=O)[n+]1ccc(OC)cc1
```

Get details on the standardization to identify why it fails for some molecules:

```python
smiles_list = [
    # erlotinib
    "n1cnc(c2cc(c(cc12)OCCOC)OCCOC)Nc1cc(ccc1)C#C",
    # midecamycin
    "CCC(=O)O[C@@H]1CC(=O)O[C@@H](C/C=C/C=C/[C@@H]([C@@H](C[C@@H]([C@@H]([C@H]1OC)O[C@H]2[C@@H]([C@H]([C@@H]([C@H](O2)C)O[C@H]3C[C@@]([C@H]([C@@H](O3)C)OC(=O)CC)(C)O)N(C)C)O)CC=O)C)O)C",
    # selenofolate
    "C1=CC(=CC=C1C(=O)NC(CCC(=O)OCC[Se]C#N)C(=O)O)NCC2=CN=C3C(=N2)C(=O)NC(=N3)N",
    # cisplatin
    "N.N.Cl[Pt]Cl",
]

for smiles in smiles_list:
    mol = Chem.MolFromSmiles(smiles)
    print(Papyrus_standardizer.standardize(mol, return_type=True))

# (<rdkit.Chem.rdchem.Mol object at 0x000000946F99B580>, <StandardizationResult.CORRECT_MOLECULE: 1>)
# (None, <StandardizationResult.NON_SMALL_MOLECULE: 2>)
# (None, <StandardizationResult.INORGANIC_MOLECULE: 3>)
# (None, <StandardizationResult.MIXTURE_MOLECULE: 4>)
```

Allow other atoms to be considered organic:

```python
smiles = "CCN(CC)C(=O)C1=CC=C(S1)C2=C3C=CC(=[N+](C)C)C=C3[Se]C4=C2C=CC(=C4)N(C)C.F[P-](F)(F)(F)(F)F"
mol = Chem.MolFromSmiles(smiles)

print(Papyrus_standardizer.standardize(mol, return_type=True))
# (None, <StandardizationResult.INORGANIC_MOLECULE: 3>)

Papyrus_standardizer.ORGANIC_ATOMS.append('Se')

print(Papyrus_standardizer.standardize(mol, return_type=True))
# (<rdkit.Chem.rdchem.Mol object at 0x0000009F24D15F90>, <StandardizationResult.CORRECT_MOLECULE: 1>)

Papyrus_standardizer.ORGANIC_ATOMS = Papyrus_standardizer.ORGANIC_ATOMS[:-1]

print(Papyrus_standardizer.standardize(mol, return_type=True))
# (None, <StandardizationResult.INORGANIC_MOLECULE: 3>)
```

Add custom substructures to be removed as salts:

```python
# lomitapide
smiles = "C1CN(CCC1NC(=O)C2=CC=CC=C2C3=CC=C(C=C3)C(F)(F)F)CCCCC4(C5=CC=CC=C5C6=CC=CC=C64)C(=O)NCC(F)(F)F.c1ccccc1"
mol = Chem.MolFromSmiles(smiles)

print(Papyrus_standardizer.standardize(mol, return_type=True))
# (None, <StandardizationResult.MIXTURE_MOLECULE: 4>)

Papyrus_standardizer.SALTS.append('c1ccccc1')

print(Papyrus_standardizer.standardize(mol, return_type=True))
# (<rdkit.Chem.rdchem.Mol object at 0x0000009F24D15F90>, <StandardizationResult.CORRECT_MOLECULE: 1>)

Papyrus_standardizer.SALTS = Papyrus_standardizer.SALTS[:-1]

print(Papyrus_standardizer.standardize(mol, return_type=True))
# (None, <StandardizationResult.MIXTURE_MOLECULE: 4>)
```

## 📄 License

This project is licensed under the MIT License - see the [LICENSE](https://github.com/OlivierBeq/Papyrus_structure_pipeline/blob/master/LICENSE) file for details.

## 📚 API Documentation

```python
def standardize(mol,
                remove_additional_salts=True, remove_additional_metals=True,
                filter_mixtures=True, filter_inorganic=True, filter_non_small_molecule=True,
                canonicalize_tautomer=True, small_molecule_min_mw=200, small_molecule_max_mw=800,
                tautomer_allow_stereo_removal=True, tautomer_max_tautomers=50, return_type=False,
                raise_error=True
                ) -> Chem.Mol | None:
```

Standardizes a molecule: ChEMBL parent extraction, salt/metal stripping, mixture/inorganic/size
filtering, charge neutralization, tautomer canonicalization, then ChEMBL parent extraction again.

#### Parameters

- ***mol  : Chem.Mol***
  RDKit molecule object to standardize.
- ***remove_additional_salts  : bool***
  Removes a custom set of fragments if present in the molecule object.
- ***remove_additional_metals  : bool***
  Removes metal fragments if present in the molecule object. Ignored if `remove_additional_salts`
  is set to `False`.
- ***filter_mixtures  : bool***
  Return `None` if the molecule is a mixture.
- ***filter_inorganic  : bool***
  Return `None` if the molecule is inorganic.
- ***filter_non_small_molecule  : bool***
  Return `None` if the molecule is not a small molecule.
- ***canonicalize_tautomer  : bool***
  Canonicalize the tautomeric state of the molecule.
- ***small_molecule_min_mw  : float***
  Molecular weight under which a molecule is considered too small.
- ***small_molecule_max_mw  : float***
  Molecular weight above which a molecule is considered too big.
- ***tautomer_allow_stereo_removal  : bool***
  Allow the tautomer search algorithm to remove stereocenters.
- ***tautomer_max_tautomers  : int***
  Maximum number of tautomers to consider by the tautomer search algorithm. Bounds worst-case
  runtime for molecules with many tautomerizable sites.
- ***return_type  : bool***
  Add a `StandardizationResult` to the return value.
- ***raise_error  : bool***
  Raise an exception upon failure, otherwise return `None` (or `(None, StandardizationResult.STANDARDIZATION_ERROR)`
  if `return_type` is also set).

________________

```python
def is_organic(mol, return_type=False) -> bool:
```

Returns whether the RDKit molecule is organic: no ChEMBL exclusion flag, at least one C-C bond,
and made up only of atoms listed in `ORGANIC_ATOMS`.

#### Parameters

- ***mol  : Chem.Mol***
  RDKit molecule object to check the organic nature of.
- ***return_type  : bool***
  Add an `InorganicSubtype` to the return value.

________________

```python
def is_small_molecule(mol, min_molwt=200, max_molwt=800) -> bool:
```

Returns whether the RDKit molecule has a molecular weight within `min_molwt` and `max_molwt`.

#### Parameters

- ***mol  : Chem.Mol***
  RDKit molecule object to check the molecular weight of.
- ***min_molwt  : float***
  Molecular weight under which a molecule is considered too small.
- ***max_molwt  : float***
  Molecular weight above which a molecule is considered too big.

________________

```python
def is_mixture(mol) -> bool:
```

Returns whether the RDKit molecule is composed of multiple fragments.

#### Parameters

- ***mol  : Chem.Mol***
  RDKit molecule object to check the fragment count of.

________________

### Result types

- ***StandardizationResult*** — outcome of `standardize()`: `CORRECT_MOLECULE`, `NON_SMALL_MOLECULE`,
  `INORGANIC_MOLECULE`, `MIXTURE_MOLECULE`, `STANDARDIZATION_ERROR`.
- ***InorganicSubtype*** — reason `is_organic()` returned `False`: `NO_CC_BOND`, `NOT_CHONPSFClIBrB`,
  `EXCLUSION_FLAG_SET`.
- ***SaltStrippingResult*** — outcome of the internal salt-stripping step: `STRIPPED_MOLECULE`, `EMPTY_MOLECULE`.

### Module-level configuration

- ***SALTS  : list[str]*** — extra salt SMILES stripped beyond the ChEMBL pipeline's own list; user-extensible.
- ***METALS  : list[str]*** — metal SMILES stripped when `remove_additional_metals=True`.
- ***ORGANIC_ATOMS  : list[str]*** — atom symbols allowed for a molecule to be considered organic.
