Metadata-Version: 2.5
Name: openparser-sdk
Version: 1.0.3
Summary: Parse documents and extract structured data with Python
Project-URL: Homepage, https://openparser.dev
Project-URL: Documentation, https://docs.openparser.dev
Project-URL: Repository, https://github.com/eigenpal/openparser
Project-URL: Issues, https://github.com/eigenpal/openparser/issues
Project-URL: Changelog, https://github.com/eigenpal/openparser/blob/main/packages/sdk-python/CHANGELOG.md
Author-email: OpenParser <support@openparser.dev>
Maintainer-email: OpenParser <support@openparser.dev>
License: Apache-2.0
License-File: LICENSE
Keywords: ai,api,client,document-extraction,document-parsing,llm,ocr,openparser,sdk
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Information Technology
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: Implementation :: CPython
Classifier: Topic :: Internet :: WWW/HTTP
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Typing :: Typed
Requires-Python: >=3.10
Requires-Dist: attrs>=21.3.0
Requires-Dist: httpx<0.30.0,>=0.24.0
Requires-Dist: python-dateutil>=2.8.0
Requires-Dist: typing-extensions>=4.0.0
Description-Content-Type: text/markdown

# openparser-sdk

Parse and extract structured data from documents with the OpenParser API.

Install the PyPI distribution `openparser-sdk`; import the client as `openparser`.

[Documentation](https://docs.openparser.dev) · [OpenParser](https://openparser.dev)

[![license](https://img.shields.io/badge/license-Apache--2.0-3B5BDB?labelColor=555)](./LICENSE)

## Install

```bash
pip install openparser-sdk
```

Use Python 3.10 or newer. Create an API key in the OpenParser dashboard.

## Quick Start

```python
import os
from pathlib import Path
from openparser import OpenParserClient

client = OpenParserClient(api_key=os.environ["OPENPARSER_API_KEY"])

result = client.parse.sync(
    {"ocr_model": "paddleocr-vl-1.6", "output_format": "openparser@1"},
    file=Path("invoice.pdf"),
)

print(result.page_count)
```

Set `OPENPARSER_API_KEY` to create the client without constructor arguments.
Set `OPENPARSER_BASE_URL` to use a different API origin.

## Parse

The API holds synchronous requests for up to 300 seconds. It returns the result
when processing finishes or a durable job reference when the wait expires.

```python
# Sync: hold the connection until the result or timeout.
parsed = client.parse.sync(
    {"ocr_model": "paddleocr-vl-1.6"},
    file=Path("document.pdf"),
)

# Async: create a job and poll it later.
accepted = client.parse.async_(
    {"ocr_model": "paddleocr-vl-1.6"},
    file=Path("document.pdf"),
)
job = client.wait_for_job(accepted.id)

# Reuse a file-pool upload instead of inline bytes.
uploaded = client.files.upload(Path("document.pdf"))
parsed = client.parse.sync({"ocr_model": "paddleocr-vl-1.6", "file_id": uploaded.id})
```

The SDK adds an `Idempotency-Key` header to every parse and extract request. Pass
`idempotency_key=` when you need to control retries.

## Extract

```python
extracted = client.extract.sync(
    {
        "ocr_model": "paddleocr-vl-1.6",
        "llm_model": "openai/gpt-4.1-mini",
        "schema": {
            "type": "object",
            "properties": {"total": {"type": "number"}},
        },
    },
    file=Path("invoice.pdf"),
)

suggested = client.extract.suggest_schema(
    {
        "parse_job_id": "opj_...",
        "hint": "Invoice number, vendor, and total",
    }
)
```

### Grounding and lineage

Set `"grounding": "field"` to receive verified citations and a `lineage@1`
derivation DAG:

```python
extracted = client.extract.sync(
    {
        "ocr_model": "mistral-ocr-4",
        "llm_model": "openai/gpt-5.6-terra",
        "grounding": "field",
        "schema": {
            "type": "object",
            "properties": {"total": {"type": "number"}},
        },
    },
    file=Path("invoice.pdf"),
)

lineage = extracted.to_dict().get("lineage")
if lineage is not None:
    print(lineage["outputs"])
```

The graph connects each output value to its source evidence and the operations
that produced it. Evidence carries the closest recognition confidence supplied
by the selected OCR model. Applications can append normalization, calculation,
inference, and human-review activities without replacing the original machine
result.

`OpenParserClient` is synchronous. Methods named `async_` submit durable jobs
without waiting for processing to finish.

## Jobs

```python
jobs_page = client.jobs.list(status="succeeded", limit=25)
first_job_id = jobs_page.data[0].id
job = client.jobs.get("opj_...")
parse_result = client.jobs.result("opj_...", format="openparser@1")
source_bytes = client.jobs.source("opj_...")
```

`jobs.result()` returns parse representations. Extract output is available on
the job returned by `jobs.get()`.

## Files

```python
uploaded = client.files.upload(Path("contract.pdf"))
metadata = client.files.get(uploaded.id)
content = client.files.download(uploaded.id)
client.files.delete(uploaded.id)
```

Uploads accept `pathlib.Path`, a file handle, or `{"content": bytes, "filename": str, "mime_type": str?}`.

## Models

```python
ocr_models = client.models.list_ocr()
llm_models = client.models.list_llm(mode="search", q="claude")
first_ocr_model = ocr_models.data[0]
first_llm_model = llm_models.data[0]
```

## Pipelines

```python
pipeline = client.pipelines.create(
    {
        "name": "invoice-extract",
        "ocr_model": "paddleocr-vl-1.6",
        "llm_model": "anthropic/claude-sonnet-4",
        "schema": {
            "type": "object",
            "properties": {"vendor": {"type": "string"}},
            "required": ["vendor"],
        },
    }
)

listed = client.pipelines.list()
first_pipeline = listed.items[0]
current = client.pipelines.get(pipeline.id)
updated = client.pipelines.update(pipeline.id, {"name": "invoice-v2"})
client.pipelines.delete(pipeline.id)
```

## Errors

Every non-2xx response raises a typed subclass of `OpenParserError`:

| HTTP | Class                               |
| ---- | ----------------------------------- |
| 400  | `OpenParserValidationError`         |
| 401  | `OpenParserAuthError`               |
| 402  | `OpenParserPaymentRequiredError`    |
| 403  | `OpenParserForbiddenError`          |
| 404  | `OpenParserNotFoundError`           |
| 409  | `OpenParserConflictError`           |
| 413  | `OpenParserLimitExceededError`      |
| 415  | `OpenParserUnsupportedMediaError`   |
| 422  | `OpenParserUnprocessableError`      |
| 429  | `OpenParserRateLimitError`          |
| 503  | `OpenParserServiceUnavailableError` |
| 504  | `OpenParserGatewayTimeoutError`     |
| 5xx  | `OpenParserServerError`             |

Each error exposes the API response through `error.envelope`: `code`, `message`,
`request_id`, and `retryable`.

## Development

After changing the OpenAPI specification, regenerate the client and run its
checks:

```bash
packages/openparser-sdk-python/scripts/codegen.sh
packages/openparser-sdk-python/scripts/check-codegen.sh
uv run --project packages/openparser-sdk-python pytest
```

## License

[Apache-2.0](./LICENSE)
