Metadata-Version: 2.5
Name: pantogloss
Version: 0.9.0
Summary: TensorFlow/Keras many-to-English machine translation
Project-URL: Homepage, https://github.com/chrismattmann/pantogloss
Project-URL: Repository, https://github.com/chrismattmann/pantogloss
Project-URL: Model repository, https://huggingface.co/chrismattmann/pantogloss-500-en
Project-URL: Issues, https://github.com/chrismattmann/pantogloss/issues
Author: Chris A. Mattmann
License: Apache-2.0
License-File: LICENSE
License-File: NOTICE
Keywords: keras,machine translation,multilingual,tensorflow
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Operating System :: MacOS :: MacOS X
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Requires-Dist: huggingface-hub<2,>=0.26
Requires-Dist: nlcodec<0.6,>=0.5
Requires-Dist: numpy<3,>=1.26
Requires-Dist: sacremoses<0.3,>=0.2
Requires-Dist: tensorflow<2.19,>=2.18
Provides-Extra: conversion
Requires-Dist: ruamel-yaml>=0.17; extra == 'conversion'
Requires-Dist: torch<3,>=2.2; extra == 'conversion'
Provides-Extra: cuda
Requires-Dist: tensorflow[and-cuda]<2.19,>=2.18; extra == 'cuda'
Provides-Extra: evaluation
Requires-Dist: sacrebleu<3,>=2.4; extra == 'evaluation'
Provides-Extra: metal
Requires-Dist: tensorflow-metal<1.3,>=1.2; (sys_platform == 'darwin' and platform_machine == 'arm64') and extra == 'metal'
Requires-Dist: tensorflow<2.19,>=2.18; extra == 'metal'
Provides-Extra: test
Requires-Dist: pypdf<7,>=6.16; extra == 'test'
Requires-Dist: pytest-cov<8,>=5; extra == 'test'
Requires-Dist: pytest<10,>=8; extra == 'test'
Requires-Dist: sacrebleu<3,>=2.4; extra == 'test'
Provides-Extra: tika
Requires-Dist: pypdf<7,>=6.16; extra == 'tika'
Requires-Dist: tika<4,>=3.3; extra == 'tika'
Description-Content-Type: text/markdown

# Pantogloss

[![Tests](https://github.com/chrismattmann/pantogloss/actions/workflows/test.yml/badge.svg?branch=main)](https://github.com/chrismattmann/pantogloss/actions/workflows/test.yml)
[![License: Apache-2.0](https://img.shields.io/badge/License-Apache--2.0-blue.svg)](LICENSE)

Pantogloss is a TensorFlow/Keras many-to-English machine-translation library.
Its first model, `pantogloss-500-en`, was converted and numerically validated
from the model described in *Many-to-English Machine Translation
Tools, Data, and Pretrained Models* (ACL-IJCNLP 2021).

The Python package is distributed through PyPI, while the initial model is kept
in a separate public Hugging Face repository. Pantogloss 0.3.0 and later download
it anonymously by default; cached or explicit Hugging Face credentials remain
supported for private and gated model repositories.

The codebase and converted model are licensed under Apache-2.0. This repository
is private during initial development.

## Intended API

```python
from pantogloss import Translator

translator = Translator.from_pretrained("pantogloss-500-en")
print(translator.translate("Comment allez-vous ?"))
```

Greedy decoding remains the default because the full evaluation found it about
16 times faster than beam-4. The validated quality preset enables beam-4 with
length penalty 0.6 without changing the return type:

```python
print(translator.translate("Comment allez-vous ?", preset="quality"))
```

Use `preset="fast"` for the explicit greedy preset, or `preset="custom"` with
`beam_size` and `length_penalty` for advanced tuning. Explicit decoding values
also override either standard preset.

The CLI exposes the same choice:

```bash
pantogloss translate --preset fast "Comment allez-vous ?"
pantogloss translate --preset quality "Comment allez-vous ?"
pantogloss translate --preset custom --beam-size 8 --length-penalty 1.0 \
  "Comment allez-vous ?"
```

Pantogloss selects the first TensorFlow GPU automatically and enables memory
growth. Device choice can also be made explicit:

```python
translator = Translator.from_pretrained("pantogloss-500-en", device="gpu")
print(translator.device_info)
```

Using `device="gpu"` fails clearly if TensorFlow cannot see a GPU; use
`device="cpu"` to force CPU inference.

Greedy and beam translation use encode-once, graph-compiled TensorFlow decoding
loops with decoder self-attention and cross-attention key/value caches by
default. If a TensorFlow backend cannot compile a loop, Pantogloss falls back
to its eager cached decoder. Eager execution can also be selected explicitly:

```python
translator = Translator.from_pretrained("pantogloss-500-en", compiled_decode=False)
```

For exact comparison with the original full-prefix beam calculation, construct
the translator with `cached_beam=False`. This slower numerical reference may
choose different near-tied beam candidates than incremental cached attention.

Structured results are additive to the original string API:

```python
result = translator.translate_detailed("Bonjour le monde.")
print(result.text, result.source_tokens, result.target_tokens)
print(result.execution_device, result.elapsed_seconds)
```

For extracted documents, one shared model can segment, batch, and reconstruct
text while retaining blank lines and paragraph boundaries:

```python
from pantogloss import DocumentTranslator

documents = DocumentTranslator(translator)
result = documents.translate(
    extracted_text,
    source_language="fr",
    max_source_tokens=512,
    long_input="split",
)
print(result.text)
```

Long segments can use `split`, `truncate`, or `error` policy. Failures are
isolated to individual segments and recorded in `result.segments`; successful
neighbors remain ordered. Layout-only segments such as page numbers, dot
leaders, and separator rules are preserved verbatim instead of being sent to
the translation model.

Install `pantogloss[tika]` to use the `pantogloss-tika` extraction-to-English
command without a translation server or Docker container. It supports bounded
trials and page ranges, displays progress, writes output incrementally, and
resumes from a fingerprint-protected JSONL checkpoint. See
`examples/tika_to_english.py` for its compatibility wrapper.

The same workflow is available as a stable Python API:

```python
from pantogloss.tika import translate_document

result = translate_document(
    "report.pdf",
    output="report.en.txt",
    device="auto",
)
print(result.diagnostics)
print(result.runtime)
```

Apple Silicon GPU document runs automatically use short-lived TensorFlow
workers, committing five chunks before each worker exits. This bounds Metal's
retained unified-memory allocations without changing translations or checkpoint
compatibility. Tune the interval with `metal_worker_chunks=N` in Python or
`--metal-worker-chunks N` on the command line; use `None` in Python or `0` on
the command line to disable recycling. CPU and CUDA runs remain in-process.

New checkpoints record per-chunk diagnostics, translation timing, throughput,
peak process memory, and effective device metadata. Inspect or validate them
offline without importing TensorFlow or Tika:

```bash
pantogloss-tika inspect report.en.txt.jsonl
pantogloss-tika validate report.en.txt.jsonl --output report.en.txt
pantogloss-tika review report.en.txt.jsonl --output review.jsonl
pantogloss-tika review report.en.txt.jsonl --format csv --output review.csv
pantogloss-tika review report.en.txt.jsonl \
  --include-preserved formula --format csv --output formulas.csv
```

Diagnostics are conservative review signals—not translation-quality scores.
They flag empty output, retained multi-character source-script runs, fourfold
word repetition, and extreme character-length ratios; scientific symbols such
as `α`, `β`, and `π` are intentionally not treated as untranslated prose.
The model-free `review` command joins every finding back to its aligned source
and translation with chunk, segment, character-offset, token, truncation, and
error metadata. JSONL is the default for automated processing; CSV is convenient
for spreadsheet review.
Use `--include-preserved formula`, `layout_only`, or `all` to add informational
rows for deliberately untranslated spans without turning them into diagnostic
warnings. During classifier development, `tools/audit_formula_detection.py`
replays the explainable formula policy over an existing checkpoint and can emit
candidate-level JSONL with signal counts, natural-word count, and math density.

Formula-heavy spans are preserved verbatim by default because general-purpose
translation models can turn extracted equations into plausible but invented
prose. Checkpoints record these spans with `preservation_reason="formula"`, and
`inspect` reports preservation counts. Normal prose containing occasional
mathematical notation remains translatable. Whole-span preservation requires
zero natural-language words; mixed prose and notation stays in the translation
and diagnostic path rather than silently retaining source-language prose. Use
`preserve_formulas=False` with `DocumentTranslator.translate()` or
`translate_document()`, or pass
`--translate-formulas` to `pantogloss-tika`, to restore the prior behavior. The
Tika choice is protected by the checkpoint fingerprint.

Schema-1 checkpoints created by Pantogloss 0.5 remain inspectable and can be
resumed when their source checksum, model revision, and decoding options match.

The August 2026 full-document validation translated a 285,527-character Russian
dissertation on an Apple M3 Max in 58 durable chunks. Twelve recycled Metal
workers bounded peak process RSS at 6.10 GiB. All 6,517 aligned segments
completed without failures or truncation, 793 layout-only segments were
preserved verbatim, and the 286,361-character English result contained no
Cyrillic runs. The result was byte-identical across exact padding and bounded
power-of-two padding, including an interrupted/resumed run. A matching
three-page CUDA trial completed 139 segments with zero failures or truncation.

Install the accelerator backend for the machine:

```bash
# Linux with an NVIDIA GPU
pip install 'pantogloss[cuda]'

# Apple Silicon
pip install 'pantogloss[metal]'
```

Both use the same `device="auto"` or `device="gpu"` Python API. The CUDA extra
does not install or replace the host NVIDIA driver. The Metal extra uses Apple's
TensorFlow PluggableDevice and the TensorFlow 2.18 runtime combination validated
by the Bytewise project.

Source batches use bounded power-of-two padded widths by default (for example,
16, 32, and 64 through the active source-token limit). Padding remains masked
and does not change source token counts. This bounds accelerator allocation
shapes for long document runs; use `source_padding="exact"` with
`Translator.from_pretrained()` or `--source-padding exact` as a reference mode.
The policy and actual padded width are included in runtime and segment metadata
and the policy is protected by Tika checkpoint fingerprints.

Greedy decoding runs on the selected device. On Apple Silicon, beam decoding
uses a correctness-first CPU execution fallback because Panto-500 validation
found shape-sensitive corruption in Metal beam-expanded inference. CUDA beam
decoding remains on GPU. Pantogloss records the effective beam execution device
in evaluation manifests instead of silently claiming Metal placement.

The model is stored separately in the public Hugging Face repository
`chrismattmann/pantogloss-500-en`; it is never included in the Python wheel.

The default `token=None` uses a locally cached Hugging Face credential when one
exists but does not require one for public repositories. Use `token=False` to
force anonymous access or pass a token explicitly without storing it:

```python
import os

translator = Translator.from_pretrained(token=os.environ["HF_TOKEN"])
```

## Command line

The `pantogloss` command loads the model once and supports arguments, files, and
line-oriented Unix pipelines:

```bash
pantogloss info
pantogloss translate "Comment allez-vous ?"
printf 'Hola señor\nWie geht es Ihnen?\n' | pantogloss translate --device gpu
pantogloss translate --input source.txt --output english.txt --batch-size 16
pantogloss translate --beam-size 4 --length-penalty 0.6 "Hola señor"
pantogloss translate --max-source-tokens 512 --source-length-policy truncate "..."
pantogloss translate --input source.txt --preserve-empty-lines \
  --report timing.json --output english.txt
```

Use `--json` for ordered JSON Lines records containing the source index,
translation, error, inference time, and preservation status. Empty input lines
are omitted by default; `--preserve-empty-lines` copies them without model
inference. Batch failures are recursively isolated so successful neighbors are
retained and the command exits 1 after writing all results; `--fail-fast`
restores immediate termination. `--input-encoding`, `--output-encoding`, and
`--encoding-errors` control file decoding and encoding. Progress is automatic
on terminals and can be forced or disabled with `--progress` or `--no-progress`.

Use `--offline` to require an already cached model snapshot. Translation data
goes to stdout (or `--output`); model and device diagnostics are suppressed by
default so pipelines remain clean. Use `--verbose` for Pantogloss loading
progress, `--report` for a JSON timing/throughput summary, or
`--tensorflow-logs` for TensorFlow, CUDA, and Metal startup diagnostics.

## Development status

The complete 307-variable Keras model has been converted locally from all 308
learned PyTorch tensors (the target embedding and output projection are tied).
Greedy parity against the archived RTG implementation passes across a ten-language
batch: token IDs and translations match exactly, while final logits have a
maximum absolute error of 1.24e-5. With the original beam size 4 and length
penalty 0.6, all decoded four-best candidate sets match. One near-tied example
changes top rank because of framework floating-point ordering. Model version
0.1.0 is released in the public Hugging Face repository at an immutable commit.

The source model and generated artifacts stay under the ignored `artifacts/`
directory. To reproduce conversion after acquiring the source archive:

```bash
python tools/convert_rtg_checkpoint.py \
  artifacts/source/rtg500eng-tfm9L6L768d-bsz720k-stp200k-ens05 \
  artifacts/converted/pantogloss-500-en-candidate
```

Run the reference parity harness with:

```bash
CUDA_VISIBLE_DEVICES=-1 python tools/check_parity.py \
  artifacts/source/rtg500eng-tfm9L6L768d-bsz720k-stp200k-ens05 \
  artifacts/converted/pantogloss-500-en-candidate
```

To require and verify real GPU placement:

```bash
python tools/check_gpu.py artifacts/converted/pantogloss-500-en-candidate
```

## Apple Silicon validation

Pantogloss uses the same hardware-neutral GPU API for CUDA and Metal. On an
M-series Mac with Python 3.12 and Xcode command-line tools installed:

```bash
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[metal,test]'
hf auth login
python tools/check_platform.py --device cpu
python tools/check_platform.py --device gpu
python tools/benchmark_inference.py --device gpu --runs 5
```

The portable platform report identifies the selected backend as `cpu`, `cuda`,
or `metal`, verifies the first model variable's actual TensorFlow placement,
and runs a real translation. Metal placement and inference are validated on an
Apple M3 Max with TensorFlow 2.18.1. Before compiled decoding, a short batch-one
sentence had warmed medians of 0.545 seconds on CPU and 0.633 seconds on Metal.

## Decoding benchmark

Use the same input repeated into batches of 1, 8, 16, and 32:

```bash
for batch in 1 8 16 32; do
  python tools/benchmark_inference.py --device gpu --runs 5 \
    --batch-size "$batch"
done
```

The August 2026 TensorFlow 2.18.1 validation produced the following warmed
throughput. CPU and CUDA were measured on Linux; Metal was measured on an Apple
M3 Max with 128 GB unified memory.

| Batch size | CPU | CUDA (RTX 3080 Ti Laptop) | Metal (M3 Max) |
|---:|---:|---:|---:|
| 1 | 11.6/s | 14.2/s | 3.83/s |
| 8 | 62.1/s | 94.9/s | 29.15/s |
| 16 | 96.8/s | 160.6/s | 60.20/s |
| 32 | 143.0/s | 330.0/s | 115.28/s |

The M3 Max batch-one median was 0.253 seconds with compiled cached decoding,
down from the pre-compilation measurement of 0.633 seconds. Cold model load and
first-call graph compilation are reported separately from the warmed runs.

The benchmark JSON also reports total process peak RSS and, where supported by
the TensorFlow backend, allocator current memory, peak memory, and the peak
increment above its post-warmup baseline. At batch 32, CUDA's allocator rose by
22.5 MiB above the 2,114.5 MiB model baseline. Peak process RSS was approximately
8.4 GiB on CPU, 5.9 GiB with CUDA, and 4.0 GiB with Metal. TensorFlow Metal 1.2
reports zero for its allocator counters, so process RSS is the meaningful Metal
memory measurement.

Replay real checkpoint segments to test padding parity and long-run allocation
behavior without rerunning Tika:

```bash
python tools/check_padding_parity.py report.en.txt.jsonl \
  --device gpu --segments 256 --offline

python tools/stress_document_memory.py report.en.txt.jsonl \
  --device gpu --padding power_of_two --chunks 40 --offline \
  --output memory-stress.jsonl --max-growth-gib 3
```

The stress artifact is appended and flushed after every chunk, so partial runs
remain inspectable after interruption.

## Translation evaluation

Pantogloss includes a versioned evaluation runner and a checksum-pinned,
project-authored CC0 smoke corpus covering 12 languages and seven scripts. It
supports durable resumable translation artifacts, model-free rescoring,
adaptive batch recovery, deterministic paired-bootstrap confidence intervals,
and aligned regression comparisons. Reports include BLEU, chrF, per-language
diagnostics, failures, empty and unknown-token outputs, latency, and throughput.

Install the development extras and run the complete automated suite:

```bash
# Linux CUDA
python -m pip install -e '.[cuda,evaluation,test]'

# Apple Silicon Metal
python -m pip install -e '.[metal,evaluation,test]'

python -m pytest
```

The August 2026 smoke validation used Panto-500 revision
`250fc3b4122d79ac0734b28b368d2c1d68f72f7e` and TensorFlow 2.18.1:

| Platform and decoding | BLEU | chrF | Failures |
|---|---:|---:|---:|
| Linux CPU greedy | 64.72 | 73.29 | 0 |
| Apple M3 Max Metal greedy | 64.72 | 73.29 | 0 |
| Linux CPU beam-4 | 70.96 | 76.64 | 0 |
| Apple M3 Max CPU beam fallback | 70.96 | 76.64 | 0 |

The M3 beam fallback matched CPU exactly across all 12 translations: zero
metric delta, zero disagreements, and zero new failures. The automated suite
passed on both Kubuntu and Apple Silicon; hardware validation supplements the
routine tests because hosted CI does not provide these accelerators.

The targeted weak-language fixture adds two CC0 regression examples each for
Lao, Yoruba, Hausa, Igbo, Khmer, and Burmese, plus a public-safe per-language
diagnostic exporter. See `evaluation/README.md` for commands, checked-in reports, reproducibility
details, and the important limits on interpreting this deliberately small
regression fixture.

The full public-safe 8,250-example FLORES+ summary reports greedy BLEU 31.48,
chrF 57.33, and COMET 0.82975 versus beam-4 BLEU 32.42, chrF 58.03, and COMET
0.83528. Beam improved all three aggregate metrics with zero new failures, but
ran at only 6.14% of greedy throughput. The report contains no gated benchmark
sentences or translations.

### Pantogloss 0.9.0 cross-platform validation

Cached compiled beam completed all 8,250 FLORES+ examples on the RTX 3080 Ti
with zero failures at 10.47 sentences/second, an 8.57-times speedup over the
uncached numerical reference. Aggregate quality changed by only -0.088 BLEU
and -0.128 chrF. The corresponding Apple M3 Max gate passed 39 focused tests,
placed fast decoding on Metal, used the correctness-preserving CPU fallback for
beam decoding, and exactly matched all 12 checked-in weak-language CUDA
translations and metrics. The M3 three-sentence benchmarks sustained 10.80
sentences/second for `fast` and 10.26 for `quality`, with peak RSS near 3.90
GiB.

Release regression policy requires zero runtime failures, exact cross-platform
greedy agreement on the checked-in fixtures, no regression below their pinned
quality thresholds, aggregate cached-beam deltas within 0.25 BLEU and chrF of
the uncached reference, and at least a four-times cached-beam throughput
improvement. Pantogloss 0.9.0 passes every gate.

### Pantogloss 0.8.0 Apple M3 release validation

The public PyPI wheel was independently installed into a fresh Python 3.12
environment on an Apple M3 Max with TensorFlow 2.18.1. Pantogloss identified
the Metal backend and loaded Panto-500 revision
`250fc3b4122d79ac0734b28b368d2c1d68f72f7e` on `/GPU:0`. The `fast` preset
resolved to greedy decoding on Metal, while `quality` resolved to beam size 4,
length penalty 0.6, and the validated `/CPU:0` beam fallback.

A three-sentence French, Russian, and Spanish batch completed with zero
failures. Both presets returned the same English translations for these simple
inputs. Excluding model loading, fast completed in 3.72 seconds (0.806
lines/second) and quality in 10.22 seconds (0.294 lines/second), making quality
2.74 times slower for this small M3 batch. The Python API, structured execution
metadata, CLI JSONL, and timing-report paths all passed their assertions.

See the [Platform Validation wiki page](https://github.com/chrismattmann/pantogloss/wiki/Platform-Validation)
for reusable release-validation commands and guidance on interpreting
warm-up-sensitive timings. The [Pantogloss wiki](https://github.com/chrismattmann/pantogloss/wiki)
also covers installation, decoding presets, Tika document translation, and
quality evaluation.
