Metadata-Version: 2.5
Name: genotopep
Version: 0.2.0.1
Summary: Reproducible peptide candidate generation from assemblies and predicted proteins
Project-URL: Homepage, https://github.com/ljunwon1114/GenoToPep
Project-URL: Repository, https://github.com/ljunwon1114/GenoToPep
Project-URL: Issues, https://github.com/ljunwon1114/GenoToPep/issues
Author: Jun Won Lee
License-Expression: GPL-3.0-or-later
License-File: LICENSE
Keywords: bioinformatics,genomics,metagenomics,peptides
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: GNU General Public License v3 or later (GPLv3+)
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Requires-Python: >=3.10
Requires-Dist: pyrodigal<4,>=3.6
Requires-Dist: pyyaml<7,>=6
Provides-Extra: dev
Requires-Dist: build>=1.2; extra == 'dev'
Requires-Dist: pytest<10,>=8; extra == 'dev'
Requires-Dist: ruff<1,>=0.9; extra == 'dev'
Description-Content-Type: text/markdown

# GenoToPep

[![PyPI](https://img.shields.io/pypi/v/genotopep.svg)](https://pypi.org/project/genotopep/)
[![Bioconda](https://img.shields.io/conda/vn/bioconda/genotopep.svg)](https://anaconda.org/bioconda/genotopep)
[![License](https://img.shields.io/pypi/l/genotopep.svg)](LICENSE)
[![CI](https://github.com/ljunwon1114/GenoToPep/actions/workflows/ci.yml/badge.svg?branch=main)](https://github.com/ljunwon1114/GenoToPep/actions/workflows/ci.yml)

GenoToPep is a reproducible command-line package for generating peptide candidate universes from metagenome assemblies or predicted-protein FASTA files.

## Installation

Python 3.10 or newer is required.

### PyPI

```bash
python -m pip install genotopep
```

### Bioconda

```bash
conda install -c conda-forge -c bioconda genotopep
```

For version pinning and available builds, see the [Bioconda package page](https://anaconda.org/bioconda/genotopep).

New releases are published to PyPI first; Bioconda follows after its recipe is updated.

## Quickstart

Check the candidate-universe size before generation.

```bash
genotopep estimate windows \
  --assembly assembly.fna.gz \
  --min-length 10 --max-length 50 --stride 1
```

Generate exhaustive windows from an assembly.

```bash
genotopep generate windows \
  --assembly assembly.fna.gz \
  --min-length 10 --max-length 50 --stride 1 \
  --output results/assembly-windows
```

> Windows strategy candidate counts grow rapidly with protein length and the requested length range. For large inputs, run `genotopep estimate` first.

Main output files:

```text
*.windows.peptides.faa.gz
*.windows.peptide_occurrences.tsv.gz
*.manifest.json
*.checksums.sha256
```

## Input routes

- Nucleotide assembly: `--assembly sample.fna[.gz]`
- Predicted proteins: `--proteins sample.faa[.gz]`
- Predicted proteins with genomic provenance: `--proteins sample.faa --gff sample.gff3[.gz]`

`--assembly` and `--proteins` are mutually exclusive. GFF3 is optional and is only valid with FAA input.

### FNA ORF calling

Assembly input is translated with Pyrodigal metagenome mode [1,2]. The default short-ORF threshold is:

```text
min_protein_aa = 10
min_gene_nt = 10 * 3 + 3 = 33 nt
closed = false
max_overlap = 0
```

Pyrodigal still applies its coding and start-site models: `33 nt` is an eligible minimum, not a guarantee that every open reading frame of that length is emitted. Contig-edge partial calls are retained and marked in the ORF provenance output. Terminal stop symbols are not written to protein FASTA.

## Generation strategies

### Intact

Emits the complete supplied or Pyrodigal-called protein when its length is within the inclusive range.

```bash
genotopep generate intact \
  --proteins proteins.faa.gz \
  --min-length 10 --max-length 100 \
  --output results/proteins-intact
```

### Windows

Emits contiguous subsequences for every inclusive length, using stride 1 by default.

```bash
genotopep generate windows \
  --proteins proteins.faa.gz \
  --min-length 10 --max-length 50 --stride 1 \
  --output results/proteins-windows
```

### Cleavage

Builds a boundary union from selected deterministic rules, creates elementary fragments, and emits every length-valid joining of adjacent fragments.

```bash
genotopep generate cleavage \
  --proteins proteins.faa.gz \
  --cleavage-mode trypsin cnbr \
  --min-length 10 --max-length 100 \
  --output results/proteins-cleavage
```

Supported lowercase modes follow simplified deterministic profiles derived from published specificity summaries and ExPASy PeptideCutter conventions [3-5]:

- `trypsin`: cuts after K/R unless followed by P; context-dependent exceptions and counter-exceptions are not modeled [3,5]
- `chymotrypsin`: cuts after F/Y unless followed by P and after W unless followed by M/P [3,5]
- `cnbr`: boundary after M; product chemistry, methionine oxidation, and condition-dependent incomplete cleavage are not modeled [5,6]
- `all`: union of trypsin, chymotrypsin, and CNBr boundaries

Omitting `--cleavage-mode` resolves to `all`. Multiple individual modes are accepted. `all` cannot be combined with an individual mode. The cited sources support residue-specific boundary conventions; the union of selected boundaries and enumeration of every contiguous adjacent-fragment joining are GenoToPep-defined candidate-generation semantics.

## Pre-generation estimates

`estimate` reports exact occurrence counts for workload planning without writing peptide outputs or performing sequence deduplication. An occurrence is a location-derived candidate, including duplicate sequences from different proteins or coordinates. Therefore `unique peptides <= retained occurrences`; the retained occurrence count is an exact upper bound on the final exact-sequence unique peptide count.

The exclusion count is occurrence-level. Because the exclusion TSV aggregates by source protein, strategy, and reason, one exclusion row can represent many excluded occurrences. `counts.exclusions` is the excluded-candidate count, whereas `counts.exclusion_groups` is the number of aggregate rows written.

```bash
genotopep estimate \
  --assembly assembly.fna.gz \
  --cleavage-mode all \
  --min-length 10 --max-length 100 --stride 1
```

Omitting strategy selection defaults to `all`; explicitly specifying `all` has the same effect. Both expand to `intact`, `windows`, and `cleavage`. For assembly input, Pyrodigal ORF calling obtains the protein sequences required for exact counting, but ORF and peptide output files are not written.

Estimate JSON schema 1 includes strategy-specific counts, arithmetic totals, and non-blocking scale warnings. A strategy with at least 10,000,000 candidate events (`occurrences + exclusions`) receives a `large_candidate_universe` warning. The threshold does not cap or sample generation, require `--force`, or predict runtime, memory, compressed size, or temporary-database size.

## Outputs

For input `sample.proteins.faa.gz` and strategy `windows`, the main outputs are:

```text
sample.proteins.windows.peptides.faa.gz
sample.proteins.windows.peptide_occurrences.tsv.gz
sample.proteins.windows.exclusions.tsv.gz
sample.proteins.command.txt
sample.proteins.run-config.yaml
sample.proteins.run.log
sample.proteins.manifest.json
sample.proteins.checksums.sha256
```

FNA runs additionally write:

```text
sample.orfs.faa.gz
sample.orfs.fna.gz
sample.orfs.gff3.gz
sample.orfs.tsv.gz
```

Column definitions, types, coordinate scope, and conditional empty-field semantics are documented in the [output schema](docs/OUTPUT_SCHEMA.md).

### Relationships among output files

- `*.peptides.faa.gz` contains one record per exact peptide sequence; its FASTA header is the unique `peptide_id` (`>sha256:...`).
- `*.peptide_occurrences.tsv.gz` joins to the peptide FASTA on `peptide_id`. Many occurrence rows can reference one FASTA record, so this is a many-to-one relationship from occurrences to peptides.
- For FNA input, `*.orfs.tsv.gz` joins `id` to occurrence `source_gene_id`. FAA input does not produce an ORF TSV; `source_gene_id` then identifies the supplied protein.
- `*.exclusions.tsv.gz` has no `peptide_id`. For FNA input, its `source_gene_id` can join to ORF `id`; exclusions and retained occurrences are mutually exclusive candidate events.
- The ORF TSV `protein_sequence` and `nucleotide_sequence` columns intentionally duplicate the sequences in `orfs.faa.gz` and `orfs.fna.gz`, respectively, so each ORF row is self-contained.

The occurrence TSV has one row per retained candidate occurrence. The exclusion TSV aggregates excluded candidates by source protein, strategy, and reason: `counts.exclusions` equals the sum of `excluded_occurrences`, while `counts.exclusion_groups` equals its data-row count. These TSV row counts must not be added as row counts. Use `counts.occurrences + counts.exclusions` for the complete candidate-event total.

## Reproducibility

FAA sequences are converted to uppercase and trailing terminal `*` symbols are removed before filtering, generation, hashing, and output. FNA sequences and Pyrodigal translations are also converted to uppercase. FASTA identifiers must be unique, and headers containing tab, control, or DEL characters are rejected.

Peptide identity is `sha256:<digest>` of the exact normalized amino-acid sequence. Exact sequence, rather than the digest, is the internal deduplication key, and unique peptide FASTA records are ordered by binary lexicographic sequence order.

`checksums.sha256` covers the complete run inventory. Manifest schema 2 records role-specific SHA-256 values over the exact compressed bytes of the peptide FASTA, occurrence TSV, and exclusion TSV, plus a combined `payload_sha256`. Deterministic payload gzip uses compression level 6, an empty embedded filename, and `mtime=0`, as recorded in `run-config.yaml`.

Compressed-byte hashes can depend on the Python/zlib implementation. For cross-environment scientific comparison, compare decompressed payload content in addition to recorded compressed-byte hashes. Self-authored hashes detect accidental corruption and verify internal consistency; they are not authentication against deliberate replacement. Adversarial tamper resistance requires an external trust anchor.

Outputs are written through an incomplete sibling directory and renamed only after successful completion. Concurrent generation attempts for the same destination use an exclusive sibling lock. After a hard process kill, remove stale lock or incomplete artifacts only after confirming that no matching generation process is running.

`--resume` accepts only an identical, complete run after verification of schema and code identity, input/GFF content hashes, parameters, normalization, compression, payload identity, inventory, and output checksums. Recorded paths are provenance only, so identical content can resume from a different path. Use a stable `--sample-id` when renamed input or GFF files must keep the same output prefix. Resume does not continue a partially written stage.

Schema migrations and other release-specific compatibility notes are maintained in [CHANGELOG.md](CHANGELOG.md).

## GFF3 mapping scope

The current mapper resolves a protein FASTA identifier against GFF3 `ID`, `protein_id`, `locus_tag`, or `Parent`. Missing and ambiguous mappings are errors. Single-CDS bacterial proteins are supported; multipart spliced CDS models are not yet supported.

FAA input without GFF3 provides protein-relative coordinates and records that genomic coordinates are unavailable. FAA with successful GFF3 mapping and FNA-called ORFs provide genomic coordinates.

## References

1. Larralde M. Pyrodigal: Python bindings and interface to Prodigal, an efficient method for gene prediction in prokaryotes. *Journal of Open Source Software*. 2022;7(72):4296. https://doi.org/10.21105/joss.04296
2. Hyatt D, Chen G-L, LoCascio PF, Land ML, Larimer FW, Hauser LJ. Prodigal: prokaryotic gene recognition and translation initiation site identification. *BMC Bioinformatics*. 2010;11:119. https://doi.org/10.1186/1471-2105-11-119
3. Keil B. *Specificity of Proteolysis*. Springer Berlin Heidelberg; 1992. https://doi.org/10.1007/978-3-642-48380-6
4. Gasteiger E, Hoogland C, Gattiker A, Duvaud S, Wilkins MR, Appel RD, Bairoch A. Protein Identification and Analysis Tools on the ExPASy Server. In: *The Proteomics Protocols Handbook*. Humana Press; 2005:571-607. https://doi.org/10.1385/1-59259-890-0:571
5. SIB Swiss Institute of Bioinformatics. PeptideCutter: cleavage specificities of selected enzymes and chemicals. https://web.expasy.org/peptide_cutter/peptidecutter_enzymes.html
6. Gross E, Witkop B. Selective cleavage of the methionyl peptide bonds in ribonuclease with cyanogen bromide. *Journal of the American Chemical Society*. 1961;83(6):1510-1511. https://doi.org/10.1021/ja01467a052

## License

Copyright (C) 2026 Jun Won Lee.

GenoToPep is licensed under GPL-3.0-or-later. Pyrodigal is a GPL-3.0-or-later dependency. License compatibility should be reviewed before public redistribution; this statement is not legal advice.
