Metadata-Version: 2.4
Name: sopdf
Version: 0.2.0
Summary: A high-performance PDF processing library with a permissive license.
Project-URL: Homepage, https://github.com/SoMarkAI/SoPDF
Project-URL: Repository, https://github.com/SoMarkAI/SoPDF
Project-URL: Bug Tracker, https://github.com/SoMarkAI/SoPDF/issues
License: Apache-2.0
License-File: LICENSE
Keywords: document,pdf,pikepdf,pypdfium2,render,text-extraction
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Multimedia :: Graphics
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Topic :: Text Processing
Requires-Python: >=3.10
Requires-Dist: opencv-python>=4.0
Requires-Dist: pikepdf>=8.0
Requires-Dist: pypdfium2>=4.0
Provides-Extra: dev
Requires-Dist: pytest-cov>=4.0; extra == 'dev'
Requires-Dist: pytest>=7.0; extra == 'dev'
Description-Content-Type: text/markdown

<div align="center">

# SoPDF

**The PDF processing library that belongs to everyone.**

[![PyPI version](https://img.shields.io/pypi/v/sopdf.svg)](https://pypi.org/project/sopdf/)
[![Python versions](https://img.shields.io/pypi/pyversions/sopdf.svg)](https://pypi.org/project/sopdf/)
[![License](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](LICENSE)

```
pip install sopdf
```

English | [中文](README_CN.md)

</div>

---

## Why SoPDF?

**1. 🚀 High Performance**

With parallel processing and other optimizations, SoPDF significantly outperforms alternatives: rendering up to **2.78x faster**, plain text extraction **2.7x faster**, and full-text search **3x faster** — while maintaining **99% word-level accuracy consistency** with PyMuPDF. See the [Performance Benchmark](#performance-benchmark) section, or run it yourself.

**2. ✨ Feature-Rich**

Built on [`pypdfium2`](https://github.com/pypdfium2-team/pypdfium2) (Google PDFium, for rendering and text) and [`pikepdf`](https://github.com/pikepdf/pikepdf) (libqpdf, for structure and writing). SoPDF covers the entire workflow from rendering and text extraction to structural editing.

**3. 🎯 Clean API**

Intuition as documentation. You would have designed it the same way.

**4. 🔓 Permissive License**

In PDF processing, feature-rich + open source often comes with a license unfriendly to the open-source ecosystem. But SoPDF delivers equivalent core capabilities under the **Apache 2.0 License** — no strings attached, no license audit, zero friction. Embed it, ship it, fork it. It's yours.

> If you find SoPDF helpful, please consider giving it a ⭐ Star — it really means a lot to us. Every star fuels our motivation to keep improving.

---

## Benchmarks

> Measured on Apple M-series (arm64, 10-core), Python 3.10, against a 50-page PDF fixture.
> Run the suite yourself: `python tests/benchmark/run_benchmarks.py`

### Rendering vs PyMuPDF

| Scenario | SoPDF | PyMuPDF | Speedup |
| --- | --- | --- | --- |
| Open document | 0.1 ms | 0.2 ms | **1.39× faster** |
| Render 1 page @ 72 DPI | 4.1 ms | 9.0 ms | **2.20× faster** |
| Render 1 page @ 150 DPI | 11.7 ms | 30.2 ms | **2.58× faster** |
| Render 1 page @ 300 DPI | 34.2 ms | 95.2 ms | **2.78× faster** |
| 50 pages sequential @ 150 DPI | 543 ms | 1468 ms | **2.70× faster** |
| 50 pages parallel @ 150 DPI | 418 ms | 446 ms | **1.07× faster** |

SoPDF wins at every DPI — and the margin widens at higher resolutions. In parallel mode, SoPDF achieves a **1.30× speedup** over its own sequential baseline (the gap narrows because sequential rendering is now much faster). PyMuPDF's thread-parallel path, on the other hand, actually *regresses* to 1510 ms (slower than sequential) because MuPDF serialises concurrent renders behind a global lock.

### Text Extraction vs PyMuPDF

| Scenario | SoPDF | PyMuPDF | Speedup |
| --- | --- | --- | --- |
| Plain text — 50 pages | 26.0 ms | 70.0 ms | **2.70× faster** |
| Text blocks — 50 pages | 63.6 ms | 70.4 ms | **1.11× faster** |
| Search 'benchmark' — 50 pages | 30.2 ms | 91.0 ms | **3.01× faster** |
| Region extract — 50 pages | 27.6 ms | 39.6 ms | **1.43× faster** |

Text search is the standout: **3× faster** than PyMuPDF. Plain-text extraction follows at **2.7×**. Correctness is verified — sopdf and PyMuPDF produce 99% word-level overlap on the same document, so the speed advantage carries no accuracy trade-off.

---

## Architecture

SoPDF runs two best-in-class C/C++ engines in tandem:

```
┌──────────────────────────────────────────┐
│               SoPDF Python API           │
├───────────────────┬──────────────────────┤
│   pypdfium2       │   pikepdf            │
│   (Google PDFium) │   (libqpdf)          │
│                   │                      │
│   • Rendering     │   • Structure reads  │
│   • Text extract  │   • All writes       │
│   • Search        │   • Save / compress  │
└───────────────────┴──────────────────────┘
```

A **dirty-flag + hot-reload** mechanism keeps the two engines in sync:
when you write via pikepdf (e.g. rotate a page), the next read operation
(e.g. render) automatically reserialises the document into pypdfium2 —
zero manual sync required.

Files are opened with **lazy loading / mmap** — a 500 MB PDF opens in
milliseconds and only the pages you actually access are loaded.

For image encoding, SoPDF uses **OpenCV** (`opencv-python`) rather than Pillow. OpenCV's zero-copy NumPy bridge with pypdfium2 delivers **1.6×–1.9× faster** encoding than the Pillow path — see [docs/en/RuntimeDependencyCompare.md](docs/en/RuntimeDependencyCompare.md) for the full analysis and benchmark breakdown.

---

## Quick Start

```bash
pip install sopdf
```

Requires Python 3.10+. All native dependencies (`pypdfium2`, `pikepdf`, `opencv-python`) ship pre-built wheels for macOS, Linux, and Windows — no compiler needed.

```python
import sopdf

# --- Open ---
# from a file path (near-instant thanks to lazy loading & mmap)
with sopdf.open("document.pdf") as doc:

    # --- Render ---
    img_bytes = doc[0].render(dpi=150)            # PNG bytes
    doc[0].render_to_file("page0.png", dpi=300)   # write to disk

    # parallel rendering across all pages
    images = sopdf.render_pages(doc.pages, dpi=150, parallel=True)

    # --- Extract text ---
    text = doc[0].get_text()
    blocks = doc[0].get_text_blocks()             # list[TextBlock] with bounding boxes

    # --- Search ---
    hits = doc[0].search("invoice", match_case=False)   # list[Rect]

    # --- Split & merge ---
    new_doc = doc.split(pages=[0, 1, 2], output="chapter1.pdf")
    doc.split_each(output_dir="pages/")
    sopdf.merge(["intro.pdf", "body.pdf"], output="book.pdf")

    # --- Save ---
    doc.append(new_doc)
    doc.save("out.pdf", compress=True, garbage=True)
    raw = doc.to_bytes()                          # no disk write

    # --- Rotate ---
    doc[0].rotation = 90

    # --- Metadata ---
    print(doc.metadata.title)                  # read
    doc.metadata.title = "Updated Title"       # write
    print(doc.metadata.creation_datetime)      # parsed Python datetime

    # --- Outline (table of contents) ---
    for item in doc.outline.items:
        print(f"[p{item.page + 1}] {item.title}")
    flat = doc.outline.to_list()               # PyMuPDF-compatible flat list

# --- Encrypted PDFs ---
with sopdf.open("protected.pdf", password="hunter2") as doc:
    doc.save("unlocked.pdf")                      # encryption stripped on save

# --- Open from bytes / stream ---
with open("document.pdf", "rb") as f:
    with sopdf.open(stream=f.read()) as doc:
        print(doc.page_count)

# --- Auto-repair corrupted PDFs ---
with sopdf.open("corrupted.pdf") as doc:
    doc.save("repaired.pdf")
```

---

## Features

| Capability | Examples |
|---|---|
| Open from path / bytes / stream | [01_open](examples/01_open) |
| Render pages to PNG / JPEG | [02_render](examples/02_render) |
| Batch & parallel rendering | [02_render](examples/02_render) |
| Extract plain text | [03_extract_text](examples/03_extract_text) |
| Extract text with bounding boxes | [03_extract_text](examples/03_extract_text) |
| Full-text search with hit rects | [04_search_text](examples/04_search_text) |
| Split pages into new document | [05_split](examples/05_split) |
| Merge multiple PDFs | [06_merge](examples/06_merge) |
| Save with compression | [07_save_compress](examples/07_save_compress) |
| Serialise to bytes (no disk write) | [07_save_compress](examples/07_save_compress) |
| Rotate pages | [08_rotate](examples/08_rotate) |
| Open & save encrypted PDFs | [09_decrypt](examples/09_decrypt) |
| Auto-repair corrupted PDFs | [10_repair](examples/10_repair) |
| Read & write document metadata | [11_metadata](examples/11_metadata) |
| Read document outline (TOC) | [12_outline](examples/12_outline) |

---

## License

Apache 2.0 — see [LICENSE](LICENSE).

SoPDF is free to use in personal projects, commercial products, and open-source libraries. No licensing fees, no attribution requirements beyond the standard Apache 2.0 notice.

## WeChat Group
<img src="./docs/assets/wechat.png" width="100">
