Metadata-Version: 2.2
Name: bwamem
Version: 0.0.56
Summary: Bindings to bwa aligner
Author-email: Chang Ye <yech1990@gmail.com>
Maintainer-email: Chang Ye <yech1990@gmail.com>
License: MPL-2.0
Project-URL: Homepage, https://github.com/y9c/bwamem
Project-URL: Repository, https://github.com/y9c/bwamem
Project-URL: Documentation, https://y9c.github.io/bwamem/
Project-URL: Bug Tracker, https://github.com/y9c/bwamem/issues
Keywords: bioinformatics,alignment,bwa,genomics
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: POSIX :: Linux
Classifier: Operating System :: MacOS
Classifier: Programming Language :: C
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.8
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Requires-Python: >=3.8
Description-Content-Type: text/markdown
License-File: LICENSE.md
Requires-Dist: cffi>=1.0.0
Requires-Dist: setuptools>=75.3.2
Requires-Dist: wheel>=0.45.1
Provides-Extra: dev
Requires-Dist: pytest>=6.0; extra == "dev"
Requires-Dist: pytest-cov; extra == "dev"
Requires-Dist: ruff; extra == "dev"
Requires-Dist: twine; extra == "dev"
Provides-Extra: docs
Requires-Dist: sphinx; extra == "docs"
Requires-Dist: sphinx-rtd-theme; extra == "docs"

bwamem
======

Python bindings for the BWA-MEM aligner.

Installation
------------

```bash
pip install bwamem
```

Usage
-----

### Build Index

```python
from bwamem import BwaIndexer

indexer = BwaIndexer()
index_path = indexer.build_index('reference.fa')
```

### Single-End Alignment

```python
from bwamem import BwaAligner

aligner = BwaAligner('path/to/index')
alignments = aligner.align('ACGATCGCGATCGA')

for aln in alignments:
    print(f'{aln.ctg}:{aln.r_st} strand={aln.strand} mapq={aln.mapq}')
```

### Paired-End Alignment

```python
read1 = 'ACGATCGCGATCGA'
read2 = 'TTCGATCGATCGAT'

paired_alignments = aligner.align(read1, read2)

for pe_aln in paired_alignments:
    print(f'Insert size: {pe_aln.insert_size}, Proper pair: {pe_aln.is_proper_pair}')
```

### Retrieve Sequences from Index

```python
# Get full sequence
seq = aligner.seq('chr1')

# Get subsequence
subseq = aligner.seq('chr1', start=100, end=200)
```

### Monitor Index Building Progress

```python
from bwamem import BwaIndexer

# Progress messages are captured by default (no console spam)
indexer = BwaIndexer(capture_progress=True)
index_path = indexer.build_index('genome.fasta')

# Check progress info
progress = indexer.get_progress()
print(f"Status: {progress['status']}")
print(f"Iterations: {progress['iterations']}")
print(f"Characters processed: {progress['characters_processed']}")

# Get progress percentage (if available)
if indexer.progress_percent:
    print(f"Progress: {indexer.progress_percent:.1f}%")

# Access all captured messages
for msg in progress['messages']:
    print(msg)
```

### Control Index Building Verbosity

```python
from bwamem import BwaIndexer

# Verbosity levels:
# 0 = silent (no output)
# 1 = quiet (only warnings/errors) - default
# 2 = normal (standard BWA messages)
# 3+ = debug (verbose output)

# Silent mode
indexer = BwaIndexer(verbose=0)
indexer.build_index('genome.fasta')

# Normal mode with progress messages shown in console
indexer = BwaIndexer(verbose=2, capture_progress=False)
indexer.build_index('genome.fasta')

# Debug mode with captured progress
indexer = BwaIndexer(verbose=3, capture_progress=True)
indexer.build_index('genome.fasta')

# Custom algorithm and block size
indexer = BwaIndexer(algorithm='bwtsw', block_size=50000000)
indexer.build_index('genome.fasta')
```

### Read FASTA/FASTQ Files

```python
from bwamem import fastx_read, read_paired_fastx

# Single-end (supports both FASTA and FASTQ)
for read in fastx_read('sequences.fasta.gz'):
    print(f'{read.name}: {read.sequence}')

# Paired-end
for read1, read2 in read_paired_fastx('R1.fastq', 'R2.fastq'):
    print(f'{read1.name}, {read2.name}')
```

### Custom Options

```python
# Specify alignment parameters
aligner = BwaAligner('path/to/index', options='-x ont2d -A 1 -B 0')

# Set custom insert size for paired-end reads
aligner = BwaAligner('path/to/index', insert_model=(500, 50))
paired_alignments = aligner.align(read1, read2)
```

Command-line tool
-----------------

`bwamem map` runs BWA-MEM against one or more reference layers and emits a
SAM stream annotated with a `HI:Z:<layer>` tag used for hierarchical /
contamination-aware assignment:

```bash
# Single-end, hierarchical (contamination -> genes -> transcripts)
bwamem map -i contamination -i genes -i tx_transcript \
    -k 18,14,10 -n 0.05,0.15,0.30 -T 30,30,30 reads.fq

# Paired-end against a single reference
bwamem map -i ref -1 r1.fq.gz -2 r2.fq.gz -o out.sam
```

`HI:Z:<i>` records which layer claimed each read (`0`-based); reads that map
to no layer get `HI:Z:-1` (these are the candidates forwarded to the genome
stage). Each layer's parameters (`-k`, `-n`, `-T`) may be given as a single
value applied to all layers or a comma-separated list, one per layer.

### Parallelism (`-t / --threads`)

Alignment is parallelized across processes (fork, COW-shared index), not
threads, so the C alignment runs concurrently without holding the GIL:

```bash
bwamem map -t 8 -i ref -1 r1.fq.gz -2 r2.fq.gz
```

Single-threaded output is byte-identical to the parallel output. Because
input parsing is streamed lazily through a bounded buffer, memory stays
bounded regardless of input size.

`-t 1` (the default) keeps the historical single-process code path unchanged.

### SE primary-hit determinism

For single-end reads that have **perfectly tied, equal-scoring** alignments
(identical score / mapq / CIGAR / layer), BWA's `mem_align1` does not mark a
deterministic primary (the SE path — unlike the paired-end path — never calls
`mem_mark_primary_se`). As a result, the exact repeat-scaffold coordinate
reported for such reads can depend on process-internal state, and may differ
between runs or between the process path and fork workers.

This affects only the *reported repeat position* of a tiny fraction of
multi-mapper reads: the `HI:Z:<layer>` assignment, mapped/unmapped status,
score, mapq and CIGAR are all unaffected. Pipelines that consume the SAM by
`HI:Z` layer (e.g. the small_RNA genome-vs-not decision) are therefore
unaffected. This is pre-existing BWA behavior, not introduced by the
parallel path.

Alignment Attributes
--------------------

Each `Alignment` object contains the following attributes:

| Attribute | Description |
|-----------|-------------|
| `ctg` | Contig/reference name |
| `r_st` | Reference start position (0-based) |
| `r_en` | Reference end position (property) |
| `strand` | Strand: +1 for forward, -1 for reverse |
| `q_st`, `q_en` | Query start/end positions |
| `mapq` | Mapping quality score |
| `cigar` | CIGAR as list of `[length, op]` pairs |
| `cigar_str` | CIGAR string (property) |
| `NM` | Edit distance |
| `score` | Alignment score |
| `is_primary` | Primary alignment flag |

**Calculated properties** (computed on demand): `r_en`, `cigar_str`, `blen`, `mlen`

**CIGAR operations**: 0=M (match), 1=I (insertion), 2=D (deletion), 3=N (skip), 4=S (soft-clip), 5=H (hard-clip)

**PairedAlignment** contains: `read1`, `read2` (Alignment objects), `is_proper_pair` (bool), `insert_size` (int or None)

License
-------

- Python bindings: Mozilla Public License 2.0  
- BWA: GNU General Public License v3.0
