Metadata-Version: 2.3
Name: breakpoint2bedsv
Version: 1.2.1
Summary: Convert SV breakpoints from VCF/BCF to BED
License: GPL-3.0-or-later
Author: Geoffroy Véronique
Author-email: veronique.geoffroy@inserm.fr
Requires-Python: >=3.8
Classifier: License :: OSI Approved :: GNU General Public License v3 or later (GPLv3+)
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.8
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Requires-Dist: pysam (==0.22.1)
Requires-Dist: variant-extractor (==5.1.0)
Description-Content-Type: text/markdown


<div align="center">
  <h1 style="font-weight: bold; margin-bottom: 0.2em;">breakpoint2BedSV</h1>
  <h3 style="margin-top: 0;">Convert SV breakpoints from VCF/BCF to BED</h3>
</div>

- [Why extracting start/end SV breakpoints from a VCF is not trivial](#why-extracting-startend-sv-breakpoints-from-a-vcf-is-not-trivial)
- [Requirements](#requirements)
- [Quick Installation](#quick-installation)
  - [Install from PyPI](#install-from-pypi)
  - [Upgrade](#upgrade)
  - [Install from GitHub](#install-from-github)
  - [Python API](#python-api)
  - [Run the test suite](#run-the-test-suite)
- [Command line usage / Options](#command-line-usage--options)
- [Outputs](#outputs)
- [Variant filtering rules](#variant-filtering-rules)
  - [Behavior](#behavior)
- [How to cite?](#how-to-cite)
- [Exemple of use](#exemple-of-use)
- [License](#license)

## Why extracting start/end SV breakpoints from a VCF is not trivial

In an SV VCF, the first breakpoint is usually straightforward to retrieve from the `CHROM` and `POS` columns.
However, the second breakpoint is not encoded in a single standardized way and may appear in different fields depending on the SV type or the caller.

| SV representation                                               | First breakpoint | Second breakpoint                        |
| --------------------------------------------------------------- | ---------------- | ---------------------------------------- |
| Symbolic allele (`<DEL>`, `<DUP>`, `<INV>`, `<CNV>`)            | `CHROM:POS`      | usually `INFO/END`                       |
| Breakend notation (e.g. `]chr13:53040041]ATATATATACACACA`)      | `CHROM:POS`      | embedded in the `ALT` field              |
| Sequence notation (e.g. INS: `REF=A` and `ALT=ATGATTCGTTCTG...`)| `CHROM:POS`      | embedded in the `REF` field              |
| Sequence notation (e.g. DEL: `REF=TGGAATTAGCCTG...` and `ALT=T`)| `CHROM:POS`      | embedded in the `REF` field              |
| Caller-specific representations                                 | `CHROM:POS`      | may use alternative tags such as `SVEND` |

As a consequence, extracting both breakpoints from an SV VCF requires handling multiple representations.
`breakpoint2BedSV` addresses this issue by converting heterogeneous SV representations into a unified BED-like breakpoint format.

## Requirements
<i>cf</i> [`pyproject.toml`](pyproject.toml)

## Quick Installation

### Install from PyPI

The recommended way to install `breakpoint2BedSV` is with `pip`:

```bash
pip install breakpoint2bedsv
```

Then verify the installation:

```bash
breakpoint2bedsv --help
```

### Upgrade

To upgrade to the latest version:

```bash
pip install --upgrade breakpoint2bedsv
```

### Install from GitHub

To install the latest development version directly from GitHub:

```bash
git clone https://github.com/lgmgeo/breakpoint2BedSV.git
cd breakpoint2BedSV
poetry install
```

Then run:

```bash
poetry run breakpoint2bedsv --help
```

### Python API

`breakpoint2bedsv` can also be used as a Python library.

The public `convert()` function converts an SV VCF/BCF file into a BED file containing SV breakpoint intervals:

```python
from breakpoint2bedsv import convert

bed_file = convert(
    input_file="input.vcf.gz",
    output_file="breakpoints.bed",
)

print(bed_file)
```
The Python API can be used independently of the breakpoint2bedsv command-line interface, for example as part of another bioinformatics workflow or Python package.


### Run the test suite

To run all tests locally:

```bash
poetry run pytest -v
```

To list the collected tests without executing them:

```bash
poetry run pytest --collect-only
```

The test data and test scripts are located in the `tests/` directory.

All tests are also executed automatically through GitHub Actions on each push and pull request.


## Command line usage / Options

```bash
usage: breakpoint2bedsv [-h] [-V] [--log-file <File>] -i <File> [-d <Dir>] -o <File> [-T <Dir>] [-v]

Convert SV breakpoints from VCF/BCF to BED

optional arguments:
  -h, --help            show this help message and exit
  -V, --version         show program's version number and exit
  --log-file <File>     write log messages to the specified file

Input files:
  -i <File>, --input-file <File>
                        the SV VCF/BCF input file
                        VCF/VCF.gz/BCF files are supported
                        multi-allelic lines are not allowed
                        required

Output options:
  -o <File>, --output-file <File>
                        output BED file containing non redundant SV breakpoints
                        (VCF/BCF IDs are merged as a comma-separated list when
                        multiple variants share the same coordinates)
                        required

Behavior:
  -T <Dir>, --tmp-dir <Dir>
                        directory where temporary files will be created
                        if not provided, the system default temporary directory is used
  -v, --verbose         enable verbose output
```

## Outputs

Running the tool will generate a BED output file with SV start/end coordinates and the associated VCF ID.
Redundant genomic coordinates are merged into a single BED entry, with multiple VCF IDs reported as a comma-separated list.

## Variant filtering rules

`breakpoint2BedSV` only processes structural variants compatible with breakpoint-based BED representation.

During parsing, the following records are automatically ignored:
- FILTER = `MULTIALLELIC` (including MCNV-like multi-allelic CNV representations)
- ALT = `<BND>` (breakend complex rearrangements)
- ALT = `<CPX>` (complex structural variants)
- ALT = `<CTX>` (complex translocations)

### Behavior
- These variants are **skipped during parsing**
- They are **not written to the output BED file**
- The number of skipped records is reported as a warning in the standard output

## How to cite?

Please cite the following doi if you are using this tool in your research:<br>
[![DOI](./doc/zenodo.21134592.svg)](https://doi.org/10.5281/zenodo.21134592)

## Exemple of use
This API is used in [breakpoint2ExclFlagSV](https://github.com/lgmgeo/breakpoint2ExclFlagSV)

## License

breakpoint2bedsv is free software: you can redistribute it and/or modify
it under the terms of the GNU General Public License as published by
the Free Software Foundation, either version 3 of the License, or
(at your option) any later version.

breakpoint2bedsv is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
GNU General Public License for more details.

See the `LICENSE` file for the full license text.

