Metadata-Version: 2.4
Name: durabl
Version: 0.1.0
Summary: A durable execution runtime for LLM agent workflows.
Author: Aadi Malaviya
License-Expression: MIT
Project-URL: Homepage, https://github.com/Aad-im/durabl
Project-URL: Repository, https://github.com/Aad-im/durabl
Project-URL: Issues, https://github.com/Aad-im/durabl/issues
Keywords: durable-execution,workflow,event-sourcing,deterministic-replay,exactly-once,agents,llm,fault-tolerance
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Operating System :: POSIX
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Libraries :: Application Frameworks
Classifier: Topic :: System :: Distributed Computing
Classifier: Typing :: Typed
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pydantic>=2
Provides-Extra: dev
Requires-Dist: pytest; extra == "dev"
Requires-Dist: pytest-timeout; extra == "dev"
Requires-Dist: hypothesis; extra == "dev"
Dynamic: license-file

# durabl

**An LLM agent that dies at step 38 of 40 normally starts over from step 1 — and
sends the email it already sent.** durabl makes the restart resume instead.

Below is one agent, killed with `SIGKILL` partway through and started again.
Nothing was re-run, and the email went out once.

![The durabl dashboard after a killed run resumed](https://raw.githubusercontent.com/Aad-im/durabl/main/docs/img/dashboard-resumed.png)

The seven amber steps came back out of the log and cost nothing. The three teal
ones actually ran. Each line across the list is a process boundary: the first is
where the kill landed, the second is where a durable timer expired and a new
process picked the run up an hour of wall-clock later.

It works by writing every step to an append-only log, rebuilding the agent's
state by replaying that log, and committing each external side effect together
with its completion record in a **single database transaction** — so a crash can
never leave one without the other.

```
10,712 crash scenarios   ·   11,882 process deaths injected
0 duplicate side effects   ·   0 divergent final states
```

Those numbers are measured, not asserted: `python scripts/chaos_report.py`
regenerates them, and the test suite fails if this README disagrees with what it
produced.

It is Temporal's durable-execution model, scoped to agent workflows and small
enough for one person to read end to end: 4,536 lines of runtime across
17 modules, one SQLite file, no services, no network.

```bash
pip install git+https://github.com/Aad-im/durabl.git
```

No API key and no network are needed to try it — the example ships a stub model.

---

## The idea in one example

```python
from durabl import workflow, Runtime, ModelCall, ToolCall, Sleep

@workflow
def research_agent(ctx, topic):
    plan = yield ModelCall(f"Make a research plan for {topic}")
    results = []
    for query in plan["queries"]:
        results.append((yield ToolCall("search", {"q": query})))
    yield Sleep(hours=1)
    summary = yield ModelCall(f"Summarize: {results}")
    yield ToolCall("send_email",
                   {"to": ctx.input["email"], "body": summary["text"]},
                   idempotency_key=f"summary-{ctx.run_id}")
    return summary

rt = Runtime(db="agent.db", llm=my_llm, tools=my_tools)
run_id = rt.start(research_agent, {"topic": "sea otters", "email": "a@b.c"})

rt.resume_all()   # on process boot: picks up whatever was interrupted
rt.tick()         # fires due timers
```

The workflow is an ordinary generator. It yields descriptions of work; the
runtime decides whether to *do* that work or hand back what it did last time.
Kill the process anywhere in there, start it again, and `resume_all()` puts it
back exactly where it was — without re-running the searches and without
re-sending the email.

`Sleep(hours=1)` is a durable timer, not `time.sleep`. The process is free to
exit; the run is picked back up an hour later by whichever process is alive
then.

## How resume works

There is no serialised "agent state" anywhere. On resume durabl re-runs the
workflow function from the top, and for each step the function yields, looks at
the log:

```
    workflow yields ──▶  is this step in the log?
                            │
        ┌───────────────────┼────────────────────┐
        │                   │                    │
  recorded, done      log ends here       recorded, different
        │                   │                    │
   return the          execute it          NonDeterminismError
   recorded result     for real            (the workflow drifted)
   (zero I/O)
```

The third branch is the one that makes this trustworthy. A runtime that
matched positionally without checking would cheerfully hand a search result to
a step expecting an email receipt. durabl fingerprints every step and compares,
so a workflow that reached for `time.time()` or `random` is caught rather than
silently producing wrong state. `ctx.now()`, `ctx.random()` and `ctx.uuid()`
are the deterministic substitutes; their values are frozen into the log on
first execution.

## The transaction boundary

This is the load-bearing part. A tool call is three writes:

```
  ┌─ txn ─┐        ┌─ txn ─┐                    ┌──────── txn ────────┐
  │STEP_  │        │TOOL_  │   the actual       │ side_effects row    │
  │STARTED│  ───▶  │DISPAT-│ ──▶ tool call ──▶  │ +  STEP_COMPLETED   │
  │       │        │CHED   │                    │ (one transaction)   │
  └───────┘        └───────┘                    └─────────────────────┘
       ▲                ▲                     ▲            ▲
    crash here      crash here           crash here    crash here
    tool did        tool did             AMBIGUOUS     fully recorded,
    not run         not run              (see below)   replay serves it
```

If the recorded effect and the completion event were two transactions, a crash
between them would leave an effect the log has never heard of (so replay runs
the tool again) or a completion the deduplication table has never heard of (so
another run holding the same idempotency key does it again). Both are real
bugs; both ship in `tests/chaos/broken_runtime.py`, and the harness is required
to catch them.

`TOOL_DISPATCHED` exists to shrink the ambiguous window. Without it, *any*
crash inside a step would be indistinguishable from a crash around the call
itself, and every non-retryable tool would have to give up on every crash.

## The chaos harness

`tests/chaos/` injects a real process death — a `BaseException` the runtime
never catches, the same class `KeyboardInterrupt` belongs to — at every
instrumented point in the runtime, then throws the in-memory runtime away and
builds a fresh one against the same database file. That is the same path a
restarted process takes.

| | |
|---|---|
| triples (workflow x crash point x seed) | **10,712** |
| workflows | 50 (10 crash points, 26 seeds each) |
| process deaths injected | 11,882 |
| resumed to the clean run's exact final state | 10,330 |
| correctly refused to continue (see Limits) | 382 |
| **duplicate side effects** | **0** |
| **divergent final states** | **0** |
| violations of any kind | 0 |
| wall clock | 37s on 11 cores |

Crash points, and how many triples hit each:

| crash point | triples | |
|---|---:|---|
| `BEFORE_STEP_STARTED` | 1,300 | nothing recorded yet |
| `AFTER_STEP_STARTED` | 1,300 | intent recorded, tool not called |
| `AFTER_TOOL_DISPATCHED` | 1,300 | about to leave the process |
| `AFTER_TOOL_INVOKED` | 1,300 | tool returned, result not yet committed |
| `AFTER_MODEL_INVOKED` | 624 | model returned, result not yet committed |
| `AFTER_STEP_COMMITTED` | 1,300 | durable, mid-workflow |
| `BEFORE_CTX_RECORD` | 806 | a `ctx.now()` value not yet durable |
| `DURING_REPLAY` | 1,170 | died *while resuming* from an earlier death |
| `BEFORE_TIMER_EVENT` | 312 | timer due, completion not yet appended |
| `MID_COMPACTION` | 1,300 | snapshot written, old rows not yet dropped |

A further 2,288 (workflow, crash point) pairs are skipped
because that crash cannot occur in that workflow — a workflow with no timers
has no timer to die in the middle of. Skipping them is deliberate: a triple
that never crashed would prove nothing, and the harness asserts that every
triple it counts actually took at least one process death.

There is no crash point between "side effect committed" and "event appended",
because there is no instant between them. `tests/chaos/test_negative_control.py`
asserts that the correct runtime declares no probe there, and the broken
variants had to invent their own name to have somewhere to die.

Reproduce it:

```bash
python scripts/chaos_report.py     # prints the table above, writes docs/chaos_report.json
```

### The negative control

A harness that has never failed is not a harness. 5 deliberately wrong runtimes
ship alongside the real one. Each is run across the scenarios where its
particular bug can actually show itself, with a crash aimed at the window that
bug opens, and each is *required* to be caught every single time:

| broken variant | what it gets wrong | caught by | caught |
|---|---|---|---:|
| `split-effect-then-event` | records the side effect, commits, then appends STEP_COMPLETED separately | torn write: a side_effects row with no completion event | **69/69** |
| `split-event-then-effect` | appends STEP_COMPLETED, commits, then records the side effect separately | torn write: a completion event with no side_effects row | **69/69** |
| `ignores-tool-mode` | re-runs an AT_MOST_ONCE tool after an ambiguous crash | an AT_MOST_ONCE tool invoked more times than the clean run invoked it | **9/9** |
| `skips-dedup-check` | never consults side_effects before calling a tool | an AT_MOST_ONCE tool invoked twice; no crash required | **6/6** |
| `duplicates-everything` | calls every tool twice | every applicable triple | **33/33** |
| *(the real runtime)* | nothing | — | 0/81 |

186 of 186 triples flagged. The last row is the control on the control:
the real runtime, put through the same schedules, is flagged zero times out of
81 — a harness that shouted at everything would be no more useful than one
that shouted at nothing. Regenerate with
`python scripts/negative_control_report.py`.

The first four mutations are plausible bugs — the kind you would write if you
had not thought hard about crash windows. The fifth is not plausible at all; it
is there because if a runtime that duplicates *every* side effect can get
through, no number this harness prints means anything.

Note what the two `split-*` rows imply about the main grid: it **cannot** catch
them, and that is not a gap. The bug's window is the gap between two
transactions, and the real runtime has no such gap, so the grid's crash points
have nowhere to stand. The broken variants had to declare a probe of their own
to have somewhere to die, and the targeted schedules above are what aim at it.

The torn-write check runs on the **crashed** database, before anything
resumes. That matters: durabl's recovery path consults both the log and the
side-effects table, so it self-heals from a split transaction — a check that
only looked at the end state would pass a runtime that was quietly broken.

## Guarantees and Limits

Read this section as the contract. It is deliberately narrower than "durable
execution gives you exactly-once".

**Guaranteed, unconditionally:**

- **Workflow state is exactly-once.** Replaying a log reproduces a
  byte-identical final state, executing nothing and calling no tools. A
  workflow that drifts out of determinism raises `NonDeterminismError` instead
  of producing plausible-looking wrong state.
- **The log is gapless.** Sequence numbers are contiguous with no holes and no
  duplicates — 1..N, or watermark+1..N once a prefix has been compacted away —
  enforced by a `BEFORE INSERT` trigger, not by the Python process that is the
  thing that crashes.
- **No duplicate side effects, ever.** One idempotency key produces at most one
  row in `side_effects`. Recording refuses to overwrite rather than upserting,
  so the bug this project exists to prevent is loud rather than quiet.
- **A caller-supplied idempotency key holds across runs, crashes included.**
  Not just when the first run finished the effect: if it died mid-call, a
  second run holding the same key is in the same ambiguous window and gets the
  same answer rather than doing the effect again.
- **Timers fire exactly once, never early.** Against the deadline recorded when
  the timer was *armed* — a run picked up two days later does not restart its
  hour-long sleep. Completion and disarming are one transaction.
- **Compaction is transparent.** Replay after compaction produces the same
  final state as replay of the uncompacted log; the test keeps both and
  compares.

**Guaranteed per tool mode:**

- `IDEMPOTENT` — **exactly-once observable effect, at-least-once invocation.**
  After a crash inside the ambiguous window durabl calls the tool again. If
  your tool is genuinely idempotent (a read, a `PUT` to a known key, an API
  that dedupes on a client-supplied key) nobody can tell. If it is not, you
  declared the wrong mode.
- `AT_MOST_ONCE` — **no duplicates. Not exactly-once.** After a crash inside
  the ambiguous window durabl dead-letters the step and stops the run rather
  than guess. The tool was invoked zero or one times and durabl cannot tell you
  which. Saying "exactly-once" here would be a lie: an unknown duplicate has
  been traded for a known stall, and the stall is visible in `durabl dlq`.

**Not guaranteed:**

- **The ambiguous window cannot be closed.** Between handing control to a tool
  and that tool's result becoming durable, a dead process leaves no evidence of
  whether the request landed. This is the two-generals problem; no runtime
  solves it without help from the far end. durabl makes the window as small as
  a `TOOL_DISPATCHED` marker can make it, and tells you which side of it you
  are on.
- **A workflow that got *longer* is not detected.** It is indistinguishable
  from a workflow that crashed where the log ends. Shrinking, reordering, or
  changing a step's contents *is* detected.
- **Model calls are not deduplicated.** `ModelCall` has no mode and is
  re-executed after an ambiguous crash. A duplicate costs tokens and is
  observable to nobody outside the process; stalling a run over it would be the
  worse trade.
- **Single node.** SQLite, one writer, no coordination. Two processes driving
  the same run concurrently is not a supported configuration.
- **Not an agent framework.** durabl does not decide what the agent should do,
  choose tools, or manage prompts. It makes whatever you decided survive.

`docs/SEMANTICS.md` states all of this precisely, with the state table for what
the log holds after a crash.

## Try it

```bash
python examples/research_agent.py          # runs a stub agent end to end
# hit Ctrl-C somewhere in the middle
python examples/research_agent.py          # watch it resume
```

The demo prints which steps were replayed and which were executed, and appends
to `email_outbox.log` every time the email tool actually fires. Across any
number of kills and restarts, that file has exactly one line.

## Command line

```bash
python -m durabl runs                   # every run and its status
python -m durabl show <run_id>          # the event log as a readable timeline
python -m durabl show <run_id> --trace  # ...and which steps replay serves vs executes
python -m durabl resume <run_id>        # pick one run back up
python -m durabl dlq                    # dead-lettered steps and why
python -m durabl replay <run_id> --dry  # rebuild state, touch nothing
python -m durabl metrics <run_id>       # steps, retries, replays, wall clock vs step time
python -m durabl tick                   # fire due timers
python -m durabl compact [<run_id>]     # fold a resolved prefix into a snapshot
python -m durabl ui                     # watch runs in a browser, read-only
```

`--trace` is the one to look at. It prints, line by line, which steps came out
of the log and which actually ran — which is the whole mechanism, made visible.

## Watching it happen

```bash
python -m durabl ui --db agent.db       # then open the address it prints
```

A local page that polls the log and draws one card per step. It is worth having
for one reason: it reads the database, not the process, so it keeps rendering
while the agent you are watching is being killed.

Here is the demo above, at the moment the process was killed:

![The dashboard with a run frozen mid-flight](https://raw.githubusercontent.com/Aad-im/durabl/main/docs/img/dashboard-crashed.png)

Four searches are durable. The fifth is `DISPATCHED` — the process died after
control left durabl and before the result committed, which is the one window a
crash cannot be resolved from. Nothing else in the system can tell you that;
`TOOL_DISPATCHED` is written precisely so this state is nameable.

Run it again and the page draws a line at the seq where the new process picked
up — one line per boundary, so a run that was killed *and* slept on a durable
timer shows both, rather than only the most recent. That is the screenshot at
the top of this page.

To record the sequence without racing it by hand:

```bash
python scripts/demo.py                 # real SIGKILL, same take every time
python scripts/demo.py --kill-after 2  # die earlier, fewer durable steps
``` Run the agent in one window,
the dashboard in another, and kill the agent — the page holds the run frozen
mid-flight, showing which steps were already durable and which one was inside
the ambiguous window. Start the agent again and a labelled line appears across
the step list at the seq where the new process picked up: everything above it
was served from the log, everything below it actually ran.

Nothing in the log ever records "this step was replayed" — replay writes no
event, which is the point. The line is derived from `RUN_RESUMED`, so anyone
can reconstruct it from the log after the fact.

The dashboard opens the database read-only (`PRAGMA query_only`) and never
migrates it, so watching a run cannot perturb it. It binds loopback only and
loads no external assets.

## Tests

```bash
pytest -q     # 240 tests, 0 skipped, 0 xfailed, 83s
```

Zero skips and zero xfails are a requirement, not an outcome: nothing in this
suite is allowed to pass by being switched off. The only test doubles are the
LLM and the tools — the outside world. No part of durabl is ever mocked; the
broken runtimes in the negative control are real `Runtime` subclasses.

The eight correctness oracles from `SPEC.md` section 6 each have their own
file:

| oracle | file |
|---|---|
| 1. gapless sequence | `tests/test_oracle_01_gapless.py` |
| 2. replay determinism | `tests/test_oracle_02_replay_determinism.py` |
| 3. no duplicate side effects | `tests/test_oracle_03_no_duplicate_side_effects.py` |
| 4. crash equivalence | `tests/chaos/test_chaos.py` |
| 5. timer durability | `tests/test_oracle_05_timer_durability.py` |
| 6. retry accounting | `tests/test_oracle_06_retry_accounting.py` |
| 7. compaction transparency | `tests/test_oracle_07_compaction.py` |
| 8. non-determinism detection | `tests/test_oracle_08_nondeterminism.py` |

Oracle 8 was made to pass before any tool execution existed. It is what proves
replay is real rather than cosmetic: without it, "resume" is just a runtime
handing back recorded values in order and hoping they line up.

## What is in the box

| module | |
|---|---|
| `durabl/store.py` | SQLite event log; gapless trigger; the only way to write |
| `durabl/runtime.py` | the driver: replay, execution, the transaction boundary |
| `durabl/steps.py` | `ModelCall`, `ToolCall`, `Sleep`, and step fingerprints |
| `durabl/tools.py` | tool registration and the `IDEMPOTENT` / `AT_MOST_ONCE` contract |
| `durabl/context.py` | `ctx.now()`, `ctx.random()`, `ctx.uuid()` |
| `durabl/compaction.py` | snapshots; a snapshot is a trimmed log, not a state dump |
| `durabl/observability.py` | per-run metrics and the `show` timeline |
| `durabl/instrument.py` | the named probe points the chaos harness dies at |
| `durabl/canonical.py` | one JSON encoding, so "byte-identical" means something |
| `durabl/cli.py` | `python -m durabl` |
| `durabl/ui.py` | the read-only browser dashboard |

Further reading: `docs/SEMANTICS.md` (the contract),
`docs/ARCHITECTURE.md` (how it works), `docs/DECISIONS.md` (every place the
spec left room, what was chosen, and why).

## Requirements

Python 3.11+. Runtime dependencies: the standard library and `pydantic`.
Development: `pytest`, `pytest-timeout`, `hypothesis`. Nothing else, and no
network access at any point.

---

*Every number in this README is generated. `scripts/collect_numbers.py`
measures them, `scripts/build_readme.py` renders this file from
`docs/README.template.md`, and `tests/test_docs_numbers.py` fails the suite if
the two disagree.*
