Metadata-Version: 2.5
Name: batchledger
Version: 0.1.0
Summary: Resume expensive Python batch jobs with a local SQLite ledger and zero dependencies.
Project-URL: Homepage, https://github.com/Malanaa/batchledger
Project-URL: Issues, https://github.com/Malanaa/batchledger/issues
Project-URL: Source, https://github.com/Malanaa/batchledger
Author: Syed Abdullah Imam
License-Expression: MIT
License-File: LICENSE
Keywords: batch,cache,checkpoint,etl,resume,sqlite
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.10
Provides-Extra: dev
Requires-Dist: build>=1.2; extra == 'dev'
Requires-Dist: pytest-cov>=5; extra == 'dev'
Requires-Dist: pytest>=8; extra == 'dev'
Requires-Dist: ruff>=0.11; extra == 'dev'
Requires-Dist: twine>=6; extra == 'dev'
Description-Content-Type: text/markdown

# batchledger

**Your batch script crashed at item 8,401. Start again at 8,401.**

Checkpoint expensive Python work in a local SQLite file. Restart the same script
and reuse completed results. No server, decorators, framework, or runtime dependencies.

Useful for API enrichment, document processing, data migrations, and long local
experiments where restarting from scratch costs time or money.

```bash
pip install batchledger
```

Requires Python 3.10+. This is an early **0.1** release with a deliberately small API.

## Start in four lines

```python
from batchledger import Ledger


def enrich(record):
    # Replace with your expensive API call or transformation.
    return {"id": record["id"], "score": len(record["text"])}


records = [{"id": "a", "text": "hello"}, {"id": "b", "text": "world"}]

with Ledger("enrichment.sqlite", namespace="enrich", version="1") as ledger:
    for result in ledger.map(enrich, records):
        print(result)
```

Run it again: `enrich` is not called for previously completed inputs. Results
arrive in input order, whether computed or restored. Every result is committed
**before** it is yielded. Inputs are consumed lazily, one item at a time.

## What makes it useful

- **Resume after a crash:** completed checkpoints survive process termination.
- **Inspect progress:** see completed, failed, and interrupted work with a CLI.
- **Prevent stale reuse:** input fingerprints and explicit task versions identify results.
- **Keep going after errors:** record failed items while processing the rest.
- **No accidental duplicate writers:** an OS lock rejects competing ledger instances.
- **Readable storage:** SQLite and JSON; loading results does not unpickle Python objects.

## Track outcomes and failures

```python
with Ledger("jobs.sqlite", namespace="divide", version="1") as ledger:
    for outcome in ledger.run(lambda x: 10 / x, [2, 0, 5], on_error="record"):
        if outcome.ok:
            print(outcome.value, outcome.cached, outcome.attempts)
        else:
            print(outcome.key, outcome.error)
```

The default `on_error="raise"` saves the failure and re-raises the original
exception. `"record"` yields an unsuccessful `Outcome` and continues. Failed
and interrupted items get one new attempt each time they are encountered on a
subsequent run. There is no hidden retry loop or retry limit; repeated runs can
repeat failed calls. `KeyboardInterrupt` and `SystemExit` always propagate.

`Outcome` exposes `key`, `value`, `cached`, cumulative `attempts`, `error`, and
the `ok` property. A successful `None` result has `ok=True`.

## Stable keys and deliberate invalidation

By default, the key is a SHA-256 fingerprint of canonical JSON input. Dictionary
key order does not matter. Identical inputs share a result, including duplicates
within one run. Changed inputs get new checkpoints.

Use an explicit key when your records have an ID:

```python
with Ledger("enrichment.sqlite", namespace="enrich-by-id", version="1") as ledger:
    results = list(ledger.map(enrich, records, key=lambda record: record["id"]))
```

Reusing that ID with changed input raises `InputChangedError` instead of silently
returning an outdated result. **Bump `version` when your function, prompt, model,
configuration, or external source changes.** Function code and external state
are not inspected automatically. Different namespaces and versions coexist in
one database. Keys must be non-empty strings.

Inputs and outputs must be plain JSON values: `None`, booleans, integers, finite
floats, strings, lists, and dictionaries with string keys. Tuples, bytes, sets,
datetimes, custom classes, NumPy values, and non-finite floats raise
`SerializationError`; convert them explicitly. Both fresh and cached outputs
are decoded from JSON for consistent behavior. Avoid mutating inputs in your function.

## Inspect or export without rerunning code

```bash
batchledger status enrichment.sqlite
# {"done": 2, "failed": 0, "running": 0, "total": 2}

batchledger export enrichment.sqlite > results.jsonl
batchledger export jobs.sqlite --status failed
batchledger status enrichment.sqlite --namespace enrich --version 1
python -m batchledger --help
```

These commands open the database read-only and work while the writer is active.
`running` means an attempt has started but has no committed success/failure yet;
after a crash that row remains `running` until it is encountered again.
Only encountered inputs appear in counts; the unconsumed input stream is not queued.

`inspect_records(path, namespace=None, version=None, status=None)` is the equivalent
Python iterator. Exported records contain namespace, version, key, input, result,
status, error, attempts, and a Unix `updated_at` timestamp. CLI exit codes are
`0` for successful inspection and `2` for usage/database errors; recorded job
failures do not make the inspection command fail.

## Guarantees and limits

- **At least once, not exactly once.** If an API call succeeds but the process
  dies before its checkpoint commits, that call runs again. Use provider-side
  idempotency keys for writes, payments, or other non-repeatable side effects.
  Checkpointing also cannot make the code consuming yielded results exactly-once.
- One sequential writer per database, on a **local filesystem**. Instances are
  not thread-safe. No distributed execution, async API, network-filesystem
  support, or worker pool is provided in this release.
- Always use `with Ledger(...)`. Consume or close an iterator before starting
  another run. The OS releases the writer lock on exit/crash. Leave the `.lock`
  file in place; deleting it while a writer is active can defeat locking.
- Use one canonical file path. Do not hard-link an active ledger, or modify its
  database/lock files externally. Symlink aliases are resolved before locking.
- Results, inputs, and exception messages are stored **unencrypted**. Choose an
  appropriate local location, and exclude ledger files from source control.
- Each JSON item/result must fit in memory. There is no eviction or expiration.
  Use a new database for a clean slate. Back up a live database with SQLite's
  backup API, rather than copying only its main file while WAL is active.

## Try an actual interruption

From a checkout after installation:

```bash
python examples/resume_demo.py  # intentional failure at item 3
python examples/resume_demo.py  # items 1–2 cached; items 3–5 computed
batchledger status demo.sqlite
```

## Related tools

[Joblib Memory](https://joblib.readthedocs.io/en/latest/memory.html) is a mature
choice for function-result caching, including scientific Python objects.
[persist-queue](https://github.com/peter-wangxu/persist-queue) provides persistent
queues. [Prefect](https://docs.prefect.io/) supports larger workflow orchestration.
Batchledger focuses on a narrower job: turning a synchronous loop over JSON
records into an inspectable, restartable batch script.

## Development

```bash
python -m venv .venv
source .venv/bin/activate  # Windows: .venv\Scripts\activate
python -m pip install -e '.[dev]'
pytest --cov=batchledger --cov-report=term-missing
ruff check .
python -m build
python -m twine check dist/*
```

Tests include real subprocess termination and restart, process-level lock
contention, failure recovery, lazy inputs, JSON round trips, and CLI exports.
CI runs on Linux, macOS, and Windows across supported Python versions.

MIT licensed. Created by [Syed Abdullah Imam](https://github.com/Malanaa).
