Metadata-Version: 2.4
Name: rs-xml2json
Version: 0.2.0
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.8
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Rust
Classifier: Topic :: Text Processing :: Markup :: XML
Classifier: Topic :: File Formats :: JSON
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Dist: pytest>=8 ; extra == 'dev'
Provides-Extra: dev
License-File: LICENCE.txt
Summary: Fast schema-aware XML-to-JSON converter written in Rust, with correct typing driven by XSD definitions.
Keywords: xml,json,xsd,schema,converter,rust,pyo3
Author-email: Ernesto Ruge <ernesto.ruge@binary-butterfly.de>
Requires-Python: >=3.8
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
Project-URL: Homepage, https://github.com/datex2-tools/rs-xml2json
Project-URL: Issues, https://github.com/datex2-tools/rs-xml2json/issues
Project-URL: Repository, https://github.com/datex2-tools/rs-xml2json

# rs-xml2json

A high-performance XML-to-JSON converter written in Rust with Python bindings. It uses XSD schema definitions to 
produce correctly typed JSON output — integers, floats, booleans, and arrays are represented as their proper JSON types 
rather than treating everything as strings.

The primary focus for now is DATEX II data, but the library should be generic enough to support any XML-based data 
format.


## Transparency: AI Usage

This library is mostly AI-generated and not fully reviewed, as I lack Rust skills. It's mostly tested against highly 
complex DATEX II schemas, and the results are convincing: it can handle the DATEX II schema variety, the outputs are
correct, and it's fast enough to handle the enormous DATEX II XMLs. In both aspects, correctness and performance, it's
better than any other library I found, even outside the Python ecosystem. Still, depending on your development 
preferences, you might not want to use this library.


## Features

- **Schema-aware conversion**: Uses XSD schemas to determine JSON types (string, integer, float, boolean) and array structures (via `maxOccurs`)
- **Correct array handling**: Elements with `maxOccurs="unbounded"` are always emitted as JSON arrays, even when only a single element is present
- **Selectable JSON mapping conventions**: pick the output shape that fits your downstream consumer
    - **BadgerFish** (default): attributes prefixed with `@`, mixed-content text under `$`
    - **Parker**: attributes dropped, simple elements collapse to their value, root element name stripped
    - **Abdera**: attributes nested under `attributes`, children nested under `children`
- **Streaming file I/O**: Converts large XML files with buffered reading/writing (8 MB buffers)
- **Pre-parsed schemas**: Parse an XSD once, reuse it across multiple conversions
- **XSD import/include support**: Automatically resolves and parses referenced schema files
- **`xsi:type` support**: Overrides element types dynamically based on `xsi:type` attributes
- **Python bindings via PyO3**: Use directly from Python as a native module


## Requirements

- Rust >= 1.85 (edition 2024)
- Python >= 3.8
- [maturin](https://github.com/PyO3/maturin) (for building the Python package)
- [uv](https://github.com/astral-sh/uv) (optional, for managing the Python environment)


## Installation

### Build and install the Python package

```bash
# Using maturin directly
maturin develop --release

# Or using uv
uv pip install -e .
```

This compiles the Rust code and installs the `rs_xml2json` Python module.


### Build a wheel using Docker

A Docker-based build produces a self-contained `.whl` file for your local architecture without needing Rust or maturin installed on the host. By default it matches your system Python version (e.g. 3.12, 3.13, 3.14).

Two build variants are available:

| Variant            | Dockerfile          | Wheel type  | Use case                                                   |
|--------------------|---------------------|-------------|------------------------------------------------------------|
| `debian` (default) | `Dockerfile.debian` | `manylinux` | Standard glibc-based distros (Ubuntu, Debian, Fedora, ...) |
| `alpine`           | `Dockerfile.alpine` | `musllinux` | Alpine-based / musl environments                           |

```bash
# Build a manylinux wheel matching your system Python version (output goes to dist/)
make build

# Build a musllinux wheel instead
make build VARIANT=alpine

# Build for a specific Python version
make build PYTHON_VERSION=3.13
make build PYTHON_VERSION=3.14

# Combine both options
make build VARIANT=alpine PYTHON_VERSION=3.14
```

The resulting wheel will be in the `dist/` directory. To build and install into a local `.venv` in one step:

```bash
make install
```

This creates the venv if it doesn't exist, builds the wheel, and installs it.


### Build as a Rust library only

```bash
cargo build --release
```


## Usage

### Python API

```python
from rs_xml2json import Convention, Schema, convert, convert_to_file, convert_bytes, \
    convert_with_schema, convert_to_file_with_schema, convert_bytes_with_schema
```


#### One-shot conversion (parses XSD each time)


```python
# Convert XML file to JSON string (default convention: BadgerFish)
json_string = convert("data.xml", "schema.xsd")

# Pick a different convention
json_string = convert("data.xml", "schema.xsd", Convention.PARKER)
json_string = convert("data.xml", "schema.xsd", Convention.ABDERA)

# Convert XML file to JSON file (streaming, low memory)
convert_to_file("data.xml", "schema.xsd", "output.json")

# Convert raw bytes
json_string = convert_bytes(xml_bytes, xsd_bytes)
```

#### Pre-parsed schema (recommended for multiple conversions)

```python
from rs_xml2json import Schema

# Parse the schema once
schema = Schema.from_file("schema.xsd")
# Or from bytes
schema = Schema.from_bytes(xsd_bytes)

# Reuse across conversions (convention argument is optional)
json_string = convert_with_schema("data.xml", schema)
convert_to_file_with_schema("data.xml", schema, "output.json", Convention.PARKER)
json_string = convert_bytes_with_schema(xml_bytes, schema)
```


### CLI

```bash
# Default (BadgerFish) — prints JSON to stdout
rs-xml2json data.xml schema.xsd

# Write to a file
rs-xml2json data.xml schema.xsd output.json

# Choose a convention
rs-xml2json data.xml schema.xsd --convention parker
rs-xml2json data.xml schema.xsd output.json --convention abdera
```


### Example

Given this XSD schema:

```xml
<xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema">
    <xs:element name="catalog" type="CatalogType"/>
    <xs:complexType name="CatalogType">
        <xs:sequence>
            <xs:element name="name" type="xs:string"/>
            <xs:element name="item" type="ItemType" minOccurs="0" maxOccurs="unbounded"/>
        </xs:sequence>
        <xs:attribute name="version" type="xs:integer"/>
    </xs:complexType>
    <xs:complexType name="ItemType">
        <xs:sequence>
            <xs:element name="title" type="xs:string"/>
            <xs:element name="price" type="xs:decimal"/>
            <xs:element name="quantity" type="xs:integer"/>
            <xs:element name="available" type="xs:boolean"/>
        </xs:sequence>
        <xs:attribute name="id" type="xs:integer"/>
    </xs:complexType>
</xs:schema>
```

And this XML:

```xml
<catalog version="2">
    <name>Test Catalog</name>
    <item id="1">
        <title>Widget</title>
        <price>9.99</price>
        <quantity>100</quantity>
        <available>true</available>
    </item>
</catalog>
```

The output JSON will be:

```json
{
  "catalog": {
    "@version": 2,
    "name": "Test Catalog",
    "item": [
      {
        "@id": 1,
        "title": "Widget",
        "price": 9.99,
        "quantity": 100,
        "available": true
      }
    ]
  }
}
```

Note that `item` is always an array (because the schema declares `maxOccurs="unbounded"`), and numeric/boolean values are properly typed.


## JSON mapping conventions

`rs-xml2json` supports three named conventions. The default is BadgerFish. Schema-driven type coercion (XSD types → JSON scalars) and array detection (`maxOccurs > 1` → always an array, even for a single item) apply to all three.

| XML construct                | BadgerFish (default)            | Parker                                         | Abdera                                           |
|------------------------------|---------------------------------|------------------------------------------------|--------------------------------------------------|
| Element with simple type     | Value (string, number, boolean) | Value                                          | Value                                            |
| Element with complex type    | Object                          | Object                                         | Object with optional `attributes` and `children` |
| Element with `maxOccurs > 1` | Array                           | Array                                          | Array                                            |
| Attribute                    | Key prefixed with `@`           | Dropped                                        | Nested under `attributes`                        |
| Mixed content text           | `$` key                         | Dropped (Parker doesn't handle mixed)          | Stored under `children` when no child elements   |
| Empty element                | `null`                          | `null`                                         | `null`                                           |
| Root element name            | Preserved as outer key          | Stripped (the root body is the top-level JSON) | Preserved as outer key                           |

#### Example output

For the XML shown above (`<catalog version="2"> ... </catalog>`):

**BadgerFish** (default):
```json
{ "catalog": { "@version": 2, "name": "Test Catalog", "item": [ { "@id": 1, "title": "Widget", "price": 9.99, "quantity": 100, "available": true } ] } }
```

**Parker** (no attributes, root name dropped):
```json
{ "name": "Test Catalog", "item": [ { "title": "Widget", "price": 9.99, "quantity": 100, "available": true } ] }
```

**Abdera** (attributes/children separated):
```json
{ "catalog": { "attributes": { "version": 2 }, "children": { "name": "Test Catalog", "item": [ { "attributes": { "id": 1 }, "children": { "title": "Widget", "price": 9.99, "quantity": 100, "available": true } } ] } } }
```


## Running tests

```bash
# Rust tests
cargo test

# Python CLI tests (requires the module to be installed first)
pip install -e ".[dev]"
maturin develop
pytest tests/test_cli.py -v
```

## Project structure

```
├── Cargo.toml              # Rust package manifest
├── pyproject.toml           # Python package manifest (maturin)
├── src/
│   ├── lib.rs              # Python bindings (PyO3 module)
│   ├── error.rs            # Error types
│   ├── schema/
│   │   ├── mod.rs
│   │   ├── model.rs        # Schema data model (ElementDef, TypeDef, etc.)
│   │   ├── parser.rs       # XSD parser
│   │   └── type_map.rs     # XSD built-in type → JSON type mapping
│   └── converter/
│       ├── mod.rs
│       ├── convention.rs   # Convention enum (BadgerFish/Parker/Abdera)
│       └── walker.rs       # XML → JSON conversion engine
├── python/
│   └── rs_xml2json/
│       ├── __init__.py     # Python package re-exports
│       └── cli.py          # `rs-xml2json` command-line entry point
├── tests/
│   ├── integration_test.rs # Rust integration tests
│   ├── test_cli.py         # Python CLI integration tests (pytest)
│   ├── sample.xsd          # Test schema
│   └── sample.xml          # Test data
└── data/
    └── schema.xsd          # Example schema
```

## License

See the project license file for details.

