Metadata-Version: 2.5
Name: koji-ingest
Version: 0.6.0
Summary: Ingestion pipeline for the koji-db database: parsing, chunking, and multi-vector embeddings
Project-URL: Homepage, https://github.com/tkr-projects/tkr-koji
Project-URL: Repository, https://github.com/tkr-projects/tkr-koji
Project-URL: Source, https://github.com/tkr-projects/tkr-koji/tree/main/koji-ingest
Author: Koji Team
License-Expression: Apache-2.0
Keywords: chunking,docling,embeddings,ingestion,koji,rag,vector
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Typing :: Typed
Requires-Python: >=3.10
Requires-Dist: numpy>=1.24
Requires-Dist: structlog>=23.1
Provides-Extra: colnomic
Requires-Dist: colpali-engine<1,>=0.3.9; extra == 'colnomic'
Requires-Dist: pillow>=10; extra == 'colnomic'
Requires-Dist: torch>=2.2; extra == 'colnomic'
Requires-Dist: transformers<6,>=4.45; extra == 'colnomic'
Provides-Extra: cuda
Requires-Dist: accelerate>=0.30; extra == 'cuda'
Requires-Dist: bitsandbytes>=0.43; (platform_system == 'Linux' and platform_machine == 'x86_64') and extra == 'cuda'
Requires-Dist: faster-whisper>=1.0; extra == 'cuda'
Requires-Dist: sentence-transformers>=3.0; extra == 'cuda'
Provides-Extra: dev
Requires-Dist: pillow>=10; extra == 'dev'
Requires-Dist: pyarrow>=14.0; extra == 'dev'
Requires-Dist: pytest-asyncio>=0.21; extra == 'dev'
Requires-Dist: pytest-cov>=4.0; extra == 'dev'
Requires-Dist: pytest>=7.0; extra == 'dev'
Provides-Extra: document
Requires-Dist: docling-core>=2.3; extra == 'document'
Requires-Dist: docling<3,>=2.94; extra == 'document'
Provides-Extra: enrichment
Requires-Dist: mlx-embeddings>=0.1.0; extra == 'enrichment'
Requires-Dist: mlx-vlm>=0.4.3; extra == 'enrichment'
Requires-Dist: mlx>=0.20; extra == 'enrichment'
Provides-Extra: koji
Requires-Dist: pyarrow>=14.0; extra == 'koji'
Provides-Extra: modal
Requires-Dist: modal>=1.0; extra == 'modal'
Provides-Extra: rendering
Requires-Dist: pdf2image>=1.16; extra == 'rendering'
Description-Content-Type: text/markdown

# koji-ingest

Ingestion pipeline for Kōji — document parsing, chunking, multi-vector embeddings, and VLM enrichment.

koji-ingest transforms raw documents (PDF, DOCX, PPTX, audio, images, and more) into ColBERT-style multi-vector embeddings ready for storage and semantic search in Kōji.

## Installation

The distribution is published to PyPI as `koji-ingest`; the import module is `koji_ingest`.

```bash
# Core ingestion (parsing, chunking, embeddings)
pip install koji-ingest

# With document parsing support (Docling, PDF)
pip install "koji-ingest[document]"

# With LibreOffice-based rendering (DOCX/PPTX page images)
pip install "koji-ingest[rendering]"

# With MLX VLM enrichment engine (Apple Silicon)
pip install "koji-ingest[enrichment]"

# With Kōji storage integration (PyArrow)
pip install "koji-ingest[koji]"

# With the ColNomic multi-vector engine (colpali-engine + torch + transformers)
pip install "koji-ingest[colnomic]"

# Linux + NVIDIA only: enable 4bit/8bit quantization (bitsandbytes)
pip install "koji-ingest[colnomic,cuda]"
```

### Embedding-engine install matrix

| Platform | Install | Supported `quantization` |
| --- | --- | --- |
| Apple Silicon (MPS) | `pip install "koji-ingest[colnomic]"` | `fp16`, `bf16`, `fp32` |
| CPU-only | `pip install "koji-ingest[colnomic]"` | `fp16`, `bf16`, `fp32` |
| Linux + NVIDIA CUDA | `pip install "koji-ingest[colnomic,cuda]"` | `fp16`, `bf16`, `fp32`, `4bit`, `8bit` |

`bitsandbytes` is CUDA-only, so the `[cuda]` extra is a no-op on macOS — requesting `quantization="4bit"` on a non-CUDA device raises a clear error at config-validation time rather than at model load.

## Features

### Multi-Vector Embeddings

koji-ingest generates ColBERT-style multi-vector embeddings where each document is represented by a set of per-token vectors, enabling MaxSim late-interaction scoring:

```python
from koji_ingest import MultiVectorEmbedding
import numpy as np

data = np.random.randn(10, 128).astype(np.float32)
emb = MultiVectorEmbedding(num_tokens=10, dim=128, data=data)

blob = emb.to_blob()          # serialize for Lance storage
recovered = MultiVectorEmbedding.from_blob(blob)
```

### Document Parsing and Chunking

Parse and chunk documents in a variety of formats via Docling:

```python
from koji_ingest import parse, chunk, IngestConfig

config = IngestConfig()
parsed = parse("report.pdf", config=config)
chunks = chunk(parsed, config=config)
```

Supported formats include PDF, DOCX, PPTX, Markdown, HTML, images, and audio files.

### DenseTextEngine (Lightweight Embedding)

For text-only workloads without vision, use `DenseTextEngine` backed by `mlx-embeddings`:

```python
from koji_ingest import DenseTextEngine

engine = DenseTextEngine()
embedding = engine.embed("The quick brown fox")
```

### Bring your own vector

`koji-memory` and `koji-server` take a `Vec<f32>` from the caller at every
layer and never compute one ([`koji-memory/src/lib.rs`](../koji-memory/src/lib.rs)
`NewMemory::new`, `RecallQuery::new`; [`koji-server/src/routes/memory.rs`](../koji-server/src/routes/memory.rs)).
The embedding convention is therefore the caller's responsibility, on both
sides of the wire — nothing downstream will notice if your indexing path and
your retrieval path disagree.

The convention is published as data, not code:
[`src/koji_ingest/embedding/model_conventions.json`](src/koji_ingest/embedding/model_conventions.json)
is the source of truth. It gives each model its `dim`, `max_seq_length`,
`query_prefix`, `document_prefix`, `normalization`, and `separation_floor`,
and it ships inside the wheel.

**In Python**, call `embed_document` at ingest time and `embed_query` at search
time — and nothing else:

```python
from koji_ingest.embedding import embed_document, embed_query, close_engines

vectors = await embed_document(["chunk one", "chunk two"])  # index these
query_vector = await embed_query("what did the chunk say?")  # search with this
close_engines()  # releases the cached model
```

Both refuse a model the table does not cover, raising `ConventionError` before
any model is loaded. There is deliberately no `allow_unlisted` parameter on
this surface: a caller who wants a bare model constructs
`DenseTextEngine(allow_unlisted=True)` explicitly and owns that decision. No
helper that guesses a convention is exported from `koji_ingest.embedding`.

**In another language**, read the JSON and reproduce the same two call sites:
apply `document_prefix` to text you index and `query_prefix` to text you
search with, and never the same prefix to both. You do not need to install
`koji-ingest` or its MLX extra to read the table — vendor or fetch the file.

Then check your own configuration:

```python
from koji_ingest.embedding import measure_separation, load_convention

report = await measure_separation(...)    # your matched and unrelated pairs
floor = load_convention("BAAI/bge-small-en-v1.5").separation_floor
assert report.gap >= floor
```

The pass mark is the `separation_floor` from the table, not an absolute score.
That is the failure mode this prevents: stamping one query token onto both
sides (in the origin case, E5's `query: ` on a BGE model, at ingest and at
search) produces *higher* absolute match scores while separating matched from
unrelated pairs *worse*, so a "do relevant things score high" check passes the
broken setup. The regression test that pins this is
[`tests/test_embedding_separation.py`](tests/test_embedding_separation.py).

Note that `koji-memory` gains no embedding hook from any of this, and is not
meant to: callable hooks cannot cross the HTTP boundary, so the memory surface
stays stateless and vector-taking
([`koji-server/src/routes/memory.rs`](../koji-server/src/routes/memory.rs)).

### Gemma 4 E4B VLM Enrichment (Apple Silicon)

Augment parsed documents with VLM-generated descriptions, code analysis, formula interpretations, and document summaries using the Gemma 4 E4B model via `mlx-vlm`:

```python
from koji_ingest.enrichment import GemmaEnrichmentEngine

engine = GemmaEnrichmentEngine()
enriched = await engine.enrich(parsed_content)
# enriched.figures now contain model-generated captions
# enriched.summary contains a document-level summary
```

### GPU / Metal Memory Management

Release cached GPU or Metal memory between batch operations to prevent OOM errors:

```python
from koji_ingest.gpu import release_gpu_memory

for batch in large_corpus:
    process(batch)
    release_gpu_memory()   # returns allocations to OS after each batch
```

### LibreOffice Page Renderer

Convert DOCX and PPTX slides to page images for visual ingestion:

```python
from koji_ingest.parser import parse

# Renders pages via LibreOffice headless, returns per-page images in ParsedContent
result = parse("slides.pptx", config=config)
```

### Audio Chunking

Ingest long-form audio with automatic chunking:

```python
from koji_ingest import IngestConfig, AudioConfig

config = IngestConfig(audio=AudioConfig(chunk_duration_secs=30))
result = parse("interview.mp3", config=config)
```

### Full Ingestion Pipeline

```python
from koji_ingest import Ingester, IngestConfig

ingester = Ingester(config=IngestConfig())
result = await ingester.ingest("document.pdf")

print(result.chunks)       # List[TextChunk]
print(result.embeddings)   # List[MultiVectorEmbedding]
```

## Running on Modal

Remote execution moves the GPU stages onto a Modal container while the
database stays on your machine. Nothing below changes the default: without
these settings the pipeline is entirely local and `modal` is never imported.

### 1. Install the extra

```bash
pip install "koji-ingest[modal]"
```

### 2. Authenticate

```bash
modal setup
```

### 3. Create the HuggingFace token secret

Gemma is a gated checkpoint, so the container needs a token. Create a Modal
secret named `koji-hf-token` carrying `HF_TOKEN`:

```bash
modal secret create koji-hf-token HF_TOKEN=hf_xxxxxxxxxxxxxxxx
```

### 4. Deploy the app

```bash
modal deploy -m koji_ingest.modal_app
```

The deploy pins the published `koji-ingest` from PyPI so the container's
serialization format cannot drift from your client. To deploy your working
checkout instead — for developing against unreleased changes — set
`KOJI_MODAL_DEV=1`:

```bash
KOJI_MODAL_DEV=1 modal deploy -m koji_ingest.modal_app
```

### 5. Warm the weights volume, once

```bash
modal run -m koji_ingest.modal_app::download_weights
```

This fills the `koji-hf-cache` volume so containers start without
re-downloading model weights. Run it once per deployment target, not per
ingest.

### Entry point 1 — `ModalEmbeddingWorker` (embedding only)

Select the Modal embedding backend from config and let `Ingester.from_config`
build the coordinator for you. Only the embedding calls route to the
deployed class: parsing, transcription, and enrichment still run locally,
so chunk text (and page images when visual embedding is on) is all that
leaves the machine on this path:

```python
from koji_ingest import Ingester, IngestConfig

ingester = Ingester.from_config(
    IngestConfig(embedding_backend="modal")
)
result = await ingester.process("document.pdf")
```

The same worker serves query-time embedding, so a query vector is produced by
the identical model that produced the stored vectors:

```python
from koji_ingest import ModalEmbeddingWorker

worker = ModalEmbeddingWorker("koji-ingest", "ColNomicEmbedder")
query_vectors = await worker.encode_queries(
    ["what did the report conclude?"]
)
```

### Entry point 2 — `ModalIngester` (whole document)

`ModalIngester` sends the whole document up and runs parse → transcribe →
enrich → embed in one remote call, streaming stage events back:

```python
from koji_ingest import ModalIngester

ingester = ModalIngester("koji-ingest", "IngestPipeline")
result = await ingester.process("document.pdf")
```

### Which stages run where

| Stage | `ModalEmbeddingWorker` | `ModalIngester` | Default (local) |
| --- | --- | --- | --- |
| Parse | local | Modal | local |
| Transcribe | local | Modal | local |
| Enrich (VLM) | local | Modal | local |
| Embed | Modal | Modal | local |
| Relation building | **local** | **local** | local |
| `persist_result` | **local** | **local** | local |
| Kōji database | **local** | **local** | local |

### Weights volume retention

The `koji-hf-cache` volume holds re-downloadable HuggingFace weights and
nothing else, so it is safe to drop when a model changes or you want the
storage back:

```bash
modal volume delete koji-hf-cache
```

Re-run `modal run -m koji_ingest.modal_app::download_weights` afterwards to
refill it before the next ingest.

### What leaves your machine

Documents, figures, and audio leave the machine **only** when
`embedding_backend='modal'` or `ModalIngester` is selected explicitly. The
Kōji database, relation building, and `persist_result` never do — they run
locally on every path, remote or not.

## License

Apache-2.0
