Metadata-Version: 2.4
Name: anticapt
Version: 1.0
Summary: Fine-Tuned Nucleotide Language Models for Predicting and Designing Anticancer Aptamers
Home-page: https://github.com/raghavagps/anticapt
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE.txt
Requires-Dist: pandas
Requires-Dist: numpy
Requires-Dist: scikit-learn==1.7.2
Requires-Dist: lightgbm
Requires-Dist: joblib
Requires-Dist: torch
Requires-Dist: transformers
Requires-Dist: einops
Dynamic: description
Dynamic: description-content-type
Dynamic: home-page
Dynamic: license-file
Dynamic: requires-dist
Dynamic: requires-python
Dynamic: summary

# AntiCapt: A method for predicting, designing and scanning anticancer single-stranded (ss) DNA aptamers

## Introduction

AntiCapt is developed for predicting, designing, and scanning anticancer
(AC) single-stranded (ssDNA) aptamers. Two models are incorporated -- both use the SAME
feature pipeline (fine-tuned HyenaDNA embeddings), with different
classifiers trained on different datasets:

- **Model 1 (main / default):** Anticancer aptamer v/s random oligonucleotides -- fine-tuned HyenaDNA embeddings + LightGBM
- **Model 2 (alternate):** Anticancer aptamer v/s general aptamers -- fine-tuned HyenaDNA embeddings + LightGBM
  (separate training data, scaler, and classifier)

**Modules/Jobs:** This program implements three modules (job types):

1. **Predict** -- predicting anticancer-aptamer potential of input DNA sequences.
2. **Design** -- generating all single-point mutants (every position x
   every alternative base) of input sequences and computing the
   anticancer-aptamer potential (score) of each mutant. Useful for
   identifying which single substitutions most improve predicted potency.
3. **Scan** -- creating all overlapping windows of a given length from
   longer input sequences and computing the anticancer-aptamer potential
   of each window. Useful for localizing the active region of a longer
   sequence.

## Installation

### Option 1: Install via PyPI (recommended)

```bash
pip install anticapt
```

This installs the `anticapt` command directly, with both models
(fine-tuned HyenaDNA checkpoints, LightGBM classifiers, and scalers)
bundled in -- nothing further to set up.

### Option 2: Install from source

```bash
git clone https://github.com/raghavagps/anticapt.git
cd anticapt
pip install -e .
```

Both models (Dataset1 and Dataset2 -- fine-tuned HyenaDNA checkpoints,
LightGBM classifiers, and scalers) are already included under
`anticapt/models/` and ready to use immediately -- no additional setup
needed. The layout is:

```
anticapt/models/
├── dataset1/
│   ├── finetuned_hyenadna/    <- fine-tuned HyenaDNA checkpoint
│   │                             (config.json, weights, tokenizer files)
│   ├── LGBM_model.joblib
│   └── LGBM_scaler.joblib
└── dataset2/
    ├── finetuned_hyenadna/    <- fine-tuned HyenaDNA checkpoint
    ├── LGBM_model.joblib
    └── LGBM_scaler.joblib
```

If you later want to update either model, see `models/README.txt` for
the layout requirements. If both datasets should share one fine-tuned
embedding model, copy/symlink it into both `finetuned_hyenadna/` folders.

`MAX_LENGTH` in `anticapt/hyenadna_features.py` is set to **128**,
matching the original fine-tuning pipeline -- change it there if yours
differs.

## Minimum Usage

```bash
anticapt -i example/aptamer.fasta
```

This predicts the anticancer-aptamer potential of sequences in FASTA
format, using Model 1 by default, threshold 0.5. Output saved to
`anticapt_output.csv`.

## Full Usage

```
anticapt [-h] -i INPUT [-o OUTPUT] [-j {1,2,3}] [-m {1,2}]
         [-t THRESHOLD] [-w WINLENG] [-d {1,2}]

optional arguments:
  -h, --help            show this help message and exit
  -i, --input           Input: ssDNA aptamer sequence(s) in FASTA format,
                         or one sequence per line (simple format)
  -o, --output          Output: file for saving results, default anticapt_output.csv
  -j, --job {1,2,3}     Job Type: 1:predict, 2:design, 3:scan, default 1
  -m, --model {1,2}     Model: 1:Anticancer aptamer v/s random
                         oligonucleotides (default), 2:Anticancer aptamer
                         v/s general aptamers; both use fine-tuned
                         HyenaDNA embeddings + LGBM
  -t, --threshold       Threshold: value between 0 and 1, default 0.5
  -w, --winleng         Window Length: scan mode only, default 30
  -d, --display {1,2}   Display: 1:AC-aptamers only, 2:all sequences,
                         default 2
```

## Examples

Predict, using Dataset2's model, only showing positives:
```bash
anticapt -i example/aptamer.fasta -m 2 -d 1 -o predict_results.csv
```

Design mutants of your sequences and rank by predicted score:
```bash
anticapt -i example/aptamer.fasta -j 2 -o design_results.csv
```

Scan a long sequence with a 25 nt window:
```bash
anticapt -i long_sequence.fasta -j 3 -w 25 -o scan_results.csv
```

## Input File

Two formats accepted:
1. **FASTA format** (standard) -- `>header` line followed by sequence.
2. **Simple format** -- one sequence per line, no headers; IDs are
   auto-generated as `Seq_1`, `Seq_2`, ...

Non-ACGT characters trigger a warning but are not stripped automatically
(check your sequences if you see this warning -- it usually means stray
whitespace, ambiguity codes, or an accidental protein/RNA sequence).

## Output File

Results are saved in CSV format. Column sets differ slightly by job type:

- **Predict:** `Sequence_ID, Sequence, Length, Score, Prediction`
- **Design:** `Sequence_ID, Position, Original_Base, Mutant_Base, Mutant_Sequence, Score, Prediction`
- **Scan:** `Sequence_ID, Start_Position, End_Position, Window_Sequence, Score, Prediction`

`Score` is the predicted probability (0 to 1) of the anticancer-aptamer
class; `Prediction` is the thresholded call at `--threshold`.

## Package Files

```
setup.py                              : Package metadata / build config
README.md                             : This file provides information about this package
LICENSE.txt                           : GNU GPL v3.0 license text
CITATION.cff                          : Citation information
src/anticapt/
    __init__.py                        : Package entry point, exposes core functions
    python_scripts/
        anticapt.py                    : Main CLI program (predict/design/scan)
        seq_io.py                      : FASTA/simple-format reading, output writing
        design_scan.py                 : Point-mutant and sliding-window generation
        hyenadna_features.py           : Fine-tuned HyenaDNA embedding extraction (both models)
    models/
        README.txt                     : Model directory layout
        dataset1/
            finetuned_hyenadna/          : Model 1's fine-tuned HyenaDNA checkpoint
            LGBM_model.joblib            : Model 1 classifier
            LGBM_scaler.joblib           : Model 1 scaler
        dataset2/
            finetuned_hyenadna/          : Model 2's fine-tuned HyenaDNA checkpoint
            LGBM_model.joblib            : Model 2 classifier
            LGBM_scaler.joblib           : Model 2 scaler
    example/
        aptamer.fasta                   : Test example file containing ssDNA aptamer sequences in FASTA format
```

## License

This project is licensed under the GNU General Public License v3.0 --
see the `LICENSE` file for details.

## Address for Contact

In case of any query please contact:
```
Prof. G. P. S. Raghava, Head Department of Computational Biology,
Indraprastha Institute of Information Technology (IIIT),
Okhla Phase III, New Delhi 110020; Phone: +91-11-26907444;
Email: raghava@iiitd.ac.in  Web: http://webs.iiitd.edu.in/raghava/
```
