Metadata-Version: 2.4
Name: peafowl-dox
Version: 0.4.0
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: numpy<3.0.0,>=2.2.6
Requires-Dist: pillow<12.0.0,>=11.1.0
Requires-Dist: opencv-python<5.0.0,>=4.8.0
Requires-Dist: pymupdf<2.0.0,>=1.23.0
Requires-Dist: o365<3.0.0,>=2.0.9
Dynamic: description
Dynamic: description-content-type
Dynamic: requires-dist
Dynamic: requires-python

# Peafowl Dox

![Peafowl Dox Logo](https://i.postimg.cc/YqtjKKSq/peafowl-dox-logo.png)

A utility library for image and document processing. Essential tools for handling multipart uploads, PDF conversion, and preparing documents for OCR and ML pipelines.

## Table of Contents

*(links work on GitHub; PyPI's README renderer strips heading anchors, so they're inert there)*

- [Installation](#installation)
- [Quick Start](#quick-start)
- API Reference
        - [Image Upload Processing](#image-upload-processing)
        - [PDF Conversion](#pdf-conversion)
        - [Image Resizing](#image-resizing)
        - [Image Preprocessing](#image-preprocessing)
        - [Document Processor Class](#document-processor-class)
        - [Document Counting](#document-counting)
        - [OneDrive & SharePoint Integration](#onedrive--sharepoint-integration)
- [Error Handling](#error-handling)
- [Dependencies](#dependencies)
- [Changelog](#changelog)

---

## Installation

```bash
pip install peafowl-dox
```

**Note:** Package name uses hyphens for pip, but imports use underscores:

```python
import peafowl_dox  # underscore in import!
```

---

## Quick Start

```python
from fastapi import UploadFile
from peafowl_dox import multipart_to_array

@app.post("/upload/")
async def upload_image(file: UploadFile):
                image_array = multipart_to_array(file.file)
                print(f"Image shape: {image_array.shape}")
                return {"message": "Image processed successfully"}
```

---

## API Reference

As of `0.4.0`, every image/PDF feature lives on a class (`ImageTransformer`, `PdfConverter`). The original free functions are **still fully supported** — they're kept as backward-compatible aliases bound to those classes, so no existing code breaks. New code should prefer the class-based imports below:

| Still works (backward-compatible) | Recommended |
|---|---|
| `from peafowl_dox import multipart_to_array` | `from peafowl_dox import ImageTransformer`<br>`ImageTransformer().multipart_to_array(...)` |
| `from peafowl_dox import resize_image` | `ImageTransformer().resize_image(...)` |
| `from peafowl_dox import preprocess_image` | `ImageTransformer().preprocess_image(...)` |
| `from peafowl_dox import pdf_to_images` | `from peafowl_dox import PdfConverter`<br>`PdfConverter().pdf_to_images(...)` |
| `from peafowl_dox import DocumentProcessor` | unchanged — already class-based |
| `from peafowl_dox import OneDriveProvider` | unchanged — already class-based |

### Image Upload Processing

Convert multipart file uploads to numpy arrays.

```python
from peafowl_dox import ImageTransformer

transformer = ImageTransformer()
array = transformer.multipart_to_array(file.file)

# or, the free-function form:
from peafowl_dox import multipart_to_array
array = multipart_to_array(file.file)
```

**Returns:** `np.ndarray` with shape `(height, width, channels)`

---

### PDF Conversion

Convert PDF pages to image arrays.

```python
from peafowl_dox import PdfConverter

converter = PdfConverter()
images = converter.pdf_to_images("document.pdf", dpi=150)

# or: from peafowl_dox import pdf_to_images
```

---

### Image Resizing

Resize images with aspect ratio preservation.

```python
from peafowl_dox import ImageTransformer

transformer = ImageTransformer()
resized = transformer.resize_image(image, 1024)

# or: from peafowl_dox import resize_image
```

---

### Image Preprocessing

```python
from peafowl_dox import ImageTransformer

transformer = ImageTransformer()
ocr_ready = transformer.preprocess_image(image)

# or: from peafowl_dox import preprocess_image
```

---

### Document Processor Class

```python
from peafowl_dox import DocumentProcessor
processor = DocumentProcessor()
image = processor.process_image("image.jpg")
```

---

### Document Counting

Count how many documents appear on a single scanned page (e.g. multiple IDs/cards
scanned together), and optionally visualize the detected regions for debugging.

```python
from peafowl_dox import DocumentProcessor

processor = DocumentProcessor()
page = processor.process_image("scan_page1.jpg")

count = processor.count_documents_in_page(page)
if count > 2:
    print("Notify user to re-crop the scan.")

# Debug helper: draws a bounding box + area ratio over each detected document
debug_image = processor.draw_detected_documents(page)
```

**Note:** detection is based on local content (ink/print/photos), not a strict
document count — documents packed with very little whitespace between them may
still be merged into a single detection. Use `draw_detected_documents` to
visually check this on a given scan.

---

### OneDrive & SharePoint Integration

Manage files in Microsoft OneDrive and SharePoint using Microsoft Graph API.

```python
from peafowl_dox import OneDriveProvider

client = OneDriveProvider(
                client_id="AZURE_CLIENT_ID",
                client_secret="AZURE_CLIENT_SECRET",
                tenant_id="AZURE_TENANT_ID",
                target_resource_id="SHAREPOINT_SITE_ID_OR_USER_EMAIL",
                is_sharepoint=True
)
```

Supported features:

- Upload files (auto-create folders)
- Download files or folders (recursive)
- List contents
- Delete files
- Safe or recursive folder deletion

---

## Error Handling

```python
from peafowl_dox import (
                PeafowlDoxError,
                ImageProcessingError,
                PDFConversionError,
                OneDriveIntegrationError
)
```

---

## Dependencies

- Python >= 3.8
- numpy >= 2.2.6
- Pillow >= 11.1.0
- opencv-python >= 4.8.0
- PyMuPDF >= 1.23.0
- O365 >= 2.0.35

---

## Changelog

### [0.4.0] - 2026-07-17

- Added `DocumentProcessor.count_documents_in_page` and `DocumentProcessor.draw_detected_documents` to detect/count documents on a scanned page (contrast-enhanced adaptive thresholding + contour detection)
- Reorganized the package internals: `core/` split into `vision/` (image/PDF/document processing) and `storage/` (cloud integrations)
- `image_utils.py` and `pdf_converter.py` are now classes (`ImageTransformer`, `PdfConverter`) instead of free-function modules, for consistency with `DocumentProcessor`/`OneDriveProvider`. The old free-function API (`multipart_to_array`, `resize_image`, `preprocess_image`, `pdf_to_images`) keeps working unchanged — no breaking changes

### [0.3.0] - 2025-12-17

- Added OneDriveClient for Microsoft Graph API integration
- Added OneDriveIntegrationError to exception handling

### [0.2.0] - 2025-11-06

- Renamed `prepare_for_ocr` to `preprocess_image`
- Renamed `ImageProcessor` to `DocumentProcessor`

### [0.1.0] - 2025-11-06

- Initial release
