Metadata-Version: 2.4
Name: ultrams
Version: 0.1.3
Summary: Pretrained UltraMS models for mass spectrometry
License-Expression: Apache-2.0
Project-URL: Homepage, https://github.com/Dsadd4/UltraMS
Project-URL: Source, https://github.com/Dsadd4/UltraMS
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy<3,>=1.24
Requires-Dist: torch>=2.1
Requires-Dist: transformers<5,>=4.33
Requires-Dist: huggingface-hub<1,>=0.23
Dynamic: license-file

# UltraMS

Foundation models for MS/MS spectra. Obtain spectrum-level embeddings or fine-tune the encoder with PyTorch.

## Install

Python 3.10 or newer:

```bash
python -m pip install ultrams
```

For a particular CPU, CUDA, ROCm, or Apple Silicon setup, select the appropriate [PyTorch installation](https://pytorch.org/get-started/locally/) first. To install the current GitHub source instead:

```bash
python -m pip install git+https://github.com/Dsadd4/UltraMS.git
```

## Get a spectrum embedding

```python
from ultrams import UltraMS

model = UltraMS.from_pretrained("unsupervised")
embedding = model.encode(
    mz=[100.1, 121.1, 150.0], intensity=[20, 100, 35], precursor_mz=301.2
).embedding  # NumPy vector
```

## Fine-tune with PyTorch

`UltraMS` is a `torch.nn.Module`. The two rows below show the required data format; replace them with your labelled spectra.

```python
import torch
from torch.utils.data import DataLoader
from ultrams import UltraMS

dataset = [
    {"mz": [100.1, 121.1, 150.0], "intensity": [20, 100, 35], "precursor_mz": 301.2, "target": 0.5},
    {"mz": [102.1, 135.2, 167.3], "intensity": [40, 100, 25], "precursor_mz": 315.3, "target": 0.7},
]

device = "cuda" if torch.cuda.is_available() else "cpu"
model = UltraMS.from_pretrained("unsupervised", device=device).train()
head = torch.nn.Linear(model.embedding_dim, 1).to(device)
loader = DataLoader(dataset, batch_size=2, collate_fn=model.batch_converter())
optimizer = torch.optim.AdamW([*model.parameters(), *head.parameters()], lr=1e-5)

for batch in loader:
    batch = {name: value.to(device) for name, value in batch.items()}
    prediction = head(model(batch["peaks"], batch["attention_mask"], batch["precursor_mz"]))
    loss = torch.nn.functional.mse_loss(prediction.squeeze(-1), batch["target"])
    optimizer.zero_grad()
    loss.backward()
    optimizer.step()
    print(f"loss: {loss.item():.4f}")
```

## Fine-tune in one call

For a shorter route, provide the same spectra with a `label` field.

```python
import torch
from ultrams import UltraMS

labelled_spectra = [
    {"mz": [100.1, 121.1, 150.0], "intensity": [20, 100, 35], "precursor_mz": 301.2, "label": 0.5},
    {"mz": [102.1, 135.2, 167.3], "intensity": [40, 100, 25], "precursor_mz": 315.3, "label": 0.7},
]

device = "cuda" if torch.cuda.is_available() else "cpu"
model = UltraMS.from_pretrained("unsupervised", device=device)
predictor = model.finetune(labelled_spectra, task="regression", epochs=1)
prediction = predictor.predict([100.1, 121.1, 150.0], [20, 100, 35], precursor_mz=301.2)
print(prediction)
```

Use `task="classification"` for class labels. Fine-tuning saves the model, training configuration, and loss history in `ultrams_finetune/`.

## Choose a pretrained model

| Model | Learned representation | Use | Weights |
| --- | --- | --- | --- |
| **Unsupervised** (`"unsupervised"`) | Encoder spectrum-level embedding, $h_{\mathrm{CLS}}$ | General MS/MS representation and fine-tuning | [Hugging Face](https://huggingface.co/dsadd4/UltraMS-Unsupervised) |
| **MoNA contrastive** (`"mona"`) | Embedding projection of $h_{\mathrm{CLS}}$ | Spectrum similarity learned on MoNA | [Hugging Face](https://huggingface.co/dsadd4/UltraMS-MoNA-Contrastive) |
| **Search** (`"search"`) | Embedding projection of $h_{\mathrm{CLS}}$ | Spectrum-to-spectrum similarity used to build UltraAtlas | [Hugging Face](https://huggingface.co/dsadd4/UltraMS-Search) |

`encode(...).embedding` returns the representation in the table. To use a downloaded `model.pt` directly, call `UltraMS.from_checkpoint(path)`.

## Pretraining

The [UltraMSdata pretraining code](https://github.com/Dsadd4/UltraMS/blob/main/training/README.md) has a separate installation and requires prepared UltraMSdata.
