Metadata-Version: 2.4
Name: harnix
Version: 0.9.0
Summary: A complete, model-agnostic LLM agent harness — loop, tools, context, memory, permissions, hooks, subagents, skills, MCP — with deterministic record/replay and trajectory evaluation built in.
Project-URL: Homepage, https://github.com/thejenilsoni/agent-harness
Project-URL: Repository, https://github.com/thejenilsoni/agent-harness
Project-URL: Issues, https://github.com/thejenilsoni/agent-harness/issues
Author: harnix contributors
License-Expression: MIT
License-File: LICENSE
Keywords: agent,agent-harness,agents,evaluation,llm,record-replay,reproducibility,scaffolding,tools
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Testing
Requires-Python: >=3.9
Provides-Extra: dev
Requires-Dist: build; extra == 'dev'
Requires-Dist: pytest-cov; extra == 'dev'
Requires-Dist: pytest>=7.0; extra == 'dev'
Requires-Dist: ruff; extra == 'dev'
Requires-Dist: twine; extra == 'dev'
Description-Content-Type: text/markdown

# harnix

**A complete, model-agnostic LLM agent harness — loop, tools, context, memory, permissions, hooks, sub-agents, skills, MCP — with deterministic record/replay and trajectory evaluation built in.**

[![CI](https://github.com/thejenilsoni/agent-harness/actions/workflows/ci.yml/badge.svg)](https://github.com/thejenilsoni/agent-harness/actions/workflows/ci.yml)
[![Python](https://img.shields.io/badge/python-3.9%2B-blue.svg)](https://www.python.org/)
[![License: MIT](https://img.shields.io/badge/license-MIT-green.svg)](LICENSE)
[![Dependencies: 0](https://img.shields.io/badge/runtime%20deps-0-brightgreen.svg)](pyproject.toml)

An **agent harness** is the software *around* a language model that turns its
completions into actions: the control loop, the tool layer, context management,
memory, safety/permissions, lifecycle hooks, sub-agents, and observability. A
recent survey formalizes it as a six-component tuple **H = (E, T, C, S, L, V)** and
argues that the harness — not the model — is the primary determinant of agent
reliability at scale.

`harnix` implements **all six** components in a small, zero-dependency
library — and adds the thing most harnesses lack: **reproducibility and
measurability by construction**. Record a run's model calls once and replay them
to get byte-identical trajectories, with no network and no cost. That makes agents
testable in CI like ordinary code, and makes agent evaluation honest and
comparable.

```python
from harnix import Agent, Workspace, PermissionPolicy, Memory, standard_toolset

ws = Workspace("./sandbox")
agent = Agent(
    model=my_model,                                  # any LLM (see below)
    tools=standard_toolset(ws, memory=Memory()),     # file + shell + memory tools
    permissions=PermissionPolicy.allow_all(),        # blocks dangerous commands
)
traj = agent.run("Create a README and run the tests.")
print(traj.final_output, traj.num_steps, traj.usage.total_tokens)
```

---

## The six harness components, and how we cover them

| | Component | What `harnix` provides |
| --- | --- | --- |
| **E** | Execution loop | `Agent` loop with budgets, retries, and a doom-loop guard (`agent.py`) |
| **T** | Tool registry | `@tool` (auto JSON-schema from type hints), built-in file/shell/memory tools, MCP tools |
| **C** | Context manager | `WindowedContext`: token-budget windowing + tool-output offloading + summary hook |
| **S** | State store | `Memory` (durable facts/notes, `MEMORY.md`-style digest) + `Session` save/resume |
| **L** | Lifecycle hooks | `Hook` pre/post-tool control points that **block or modify** calls; `PermissionPolicy` enforced as a hook |
| **V** | Evaluation interface | `Trajectory` + record/replay `Cassette` + scorers + process metrics |

Plus the differentiators that live across all of them: **deterministic
record/replay**, **trajectory-level metrics**, and **sub-agents** & **skills**.

## Installation

```bash
pip install harnix          # core, zero dependencies
pip install "harnix[dev]"   # + pytest, ruff, build tooling
```

## Safety is structural, not advisory

Permissions are a policy consulted by a hook *before* every tool call, so a
blocked call cannot be reasoned around by the model. Dangerous shell commands are
screened regardless of policy:

```python
from harnix import PermissionPolicy, PermissionRule, Decision

policy = PermissionPolicy(
    rules=[
        PermissionRule("read_file", Decision.ALLOW),
        PermissionRule("bash", Decision.ASK),     # ASK -> resolved by an approver
        PermissionRule("*", Decision.DENY),
    ],
    default=Decision.DENY,
)
# `rm -rf /`, fork bombs, `curl ... | sh`, disk wipes, etc. are DENIED outright.
```

Lifecycle hooks add deterministic control:

```python
from harnix import Hook
from harnix.hooks import HookDecision

class NoSecrets(Hook):
    def pre_tool_use(self, call):
        if "AWS_SECRET" in str(call.arguments):
            return HookDecision.block("refusing to handle secrets")
        return HookDecision.proceed()
```

## Plugging in a model

A model maps messages (+ tool schemas) to a response — that's the only contract.
Ready-made adapters ship for the common providers and speak each HTTP API
directly, so **no vendor SDK is required** — the core stays zero-dependency:

```python
from harnix.providers import OpenAIModel, AnthropicModel, OllamaModel, VLLMModel

model = OpenAIModel("gpt-4o-mini")                       # $OPENAI_API_KEY
model = AnthropicModel("claude-sonnet-4-20250514")       # $ANTHROPIC_API_KEY
model = OllamaModel("llama3.1")                           # local, no key
model = VLLMModel("meta-llama/Llama-3.1-8B-Instruct",    # your own server
                  base_url="http://gpu:8000/v1")

agent = Agent(model=model, tools=[...])
```

Any OpenAI-compatible gateway (Together, Groq, Fireworks, OpenRouter, LM Studio,
DeepSeek, …) works via `OpenAICompatibleModel`, and a `"provider:model"` string
resolves to a configured model:

```python
from harnix import load_model, OpenAICompatibleModel

model = load_model("anthropic:claude-sonnet-4-20250514", max_tokens=2048)
model = OpenAICompatibleModel("llama-3.1-70b",
                              base_url="https://api.groq.com/openai/v1",
                              api_key_env="GROQ_API_KEY")
```

For a custom or in-house client, wrap any callable with `FunctionModel`:

```python
from harnix import FunctionModel
from harnix.types import Message

def call_llm(messages, tools):
    return Message.assistant(content="...")   # translate your client's output

agent = Agent(model=FunctionModel(call_llm), tools=[...])
```

## The headline feature: deterministic record / replay

Wrap any model — a live provider above included — to capture its calls once, then
replay them forever with no network and no cost:

```python
from harnix import Cassette, RecordingModel, ReplayModel
from harnix.providers import OpenAIModel

model = OpenAIModel("gpt-4o-mini")

cassette = Cassette()
Agent(model=RecordingModel(model, cassette), tools=tools).run(task)
cassette.save("fixtures/task.json")          # record once

cassette = Cassette.load("fixtures/task.json")
agent = Agent(model=ReplayModel(cassette, name=model.name), tools=tools)
agent.run(task)                               # replay forever: deterministic, offline, free
```

```python
from harnix.assertions import assert_deterministic, assert_used_tool

def test_agent():
    traj = run_with_replay()
    assert_used_tool(traj, "run_tests")
    assert traj.succeeded

def test_reproducible():
    assert_deterministic(run_with_replay, times=3)
```

## The differential test-bench: *hold the model fixed, measure the harness*

Deterministic replay pins the model's decisions. Once the model is fixed, you can
change the **harness** — context policy, tool set, permissions, system prompt — and
measure the effect *offline and for free*. This is the thing full-stack harnesses
can't do: [Harness-Bench](https://arxiv.org/abs/2605.27922) establishes it with
5,194 live trajectories and real GPU spend; here it's a pure function of recorded
runs. Three primitives build on the cassette:

**1. Trajectory diff** — a semantic, step-aligned diff (not a text diff): where two
runs first diverge, which tool calls changed, and how the process metrics moved.

```python
from harnix import diff_trajectories

d = diff_trajectories(baseline_traj, candidate_traj)
print(d.render())          # first divergence at step 1; tool errors 0 -> 1; ...
d.identical                # False
```

**2. Golden-trajectory regression** — snapshot testing for agent behavior. Record a
run once, commit it as a fixture, and fail CI with a readable diff when behavior drifts:

```python
from harnix import assert_matches_golden

def test_agent_behavior_is_stable():
    traj = run_with_replay()
    assert_matches_golden(traj, "fixtures/agent.golden.json")
    # first run bootstraps the golden; HARNIX_UPDATE_GOLDEN=1 re-records
```

**3. Harness A/B ablation** — run two harness configs over a suite (model pinned by
the cassette) and get a verdict on which wins, and by how much, on accuracy *and* process cost:

```python
from harnix import compare_harnesses

report = compare_harnesses(baseline_factory, candidate_factory, suite)
print(report.summary())
#   verdict     : candidate wins: +12.5% accuracy (not significant, p=0.5)
#   accuracy    : 75.0% -> 87.5% (delta +12.5%)  |  avg tokens 240 -> 190 (-50)
#   significance: accuracy not significant (McNemar p=0.5, 95% CI [-0.1, +0.4])
#                 total_tokens: -50 significant (sign test p=0.008, 95% CI [-62, -38])
```

The comparison is **paired** (the same tasks run under both harnesses), so the
verdict is backed by the right paired tests, not a point estimate: **McNemar's
exact test** for the accuracy delta, the **sign test** for each process metric,
and a **seeded bootstrap CI** (reproducible run to run). That's the difference
between "candidate wins +12.5%" and knowing whether that lead is real or noise on
a small suite — often the accuracy delta isn't significant while the efficiency
gain clearly is. When several metrics are screened at once the p-values are
**corrected for multiple comparisons** (Holm by default; also Bonferroni or
Benjamini–Hochberg), across only the metrics that actually varied:

```python
report = compare_harnesses(a, b, suite, correction="holm")  # or "bonferroni" / "benjamini-hochberg" / "none"
```

### Fixtures survive schema changes

Cassettes and goldens are committed and outlive the code that wrote them, so both
carry a **schema version** and are **auto-migrated on load** — an old fixture
keeps working after the format evolves, and a fixture *newer* than your installed
version fails loudly instead of being misread. Upgrade committed fixtures in place
with the CLI:

```bash
harnix migrate old.cassette.json         # v1 -> v2, in place
harnix migrate run.json --dry-run        # report the change without writing
```

When a candidate harness changes the request text (so the exact cassette key
misses), enable **resilient replay** — it falls back to the nearest recorded
request and reports the count of approximate hits, so the ablation stays honest:

```python
model = ReplayModel(cassette, resilient=True)   # counterfactual replay
```

Run the whole workflow: [`examples/harness_ablation.py`](examples/harness_ablation.py).
See [`docs/STRATEGY.md`](docs/STRATEGY.md) for why this is the project's moat.

### Golden regression in CI: the pytest plugin

Installing `harnix` registers a pytest plugin (via the `pytest11` entry
point), so golden-trajectory regression testing needs no boilerplate — just the
`golden` fixture:

```python
def test_agent_behaviour_is_stable(golden):
    traj = run_agent_with_replay()      # deterministic, offline
    golden.check(traj)                  # -> goldens/test_agent_behaviour_is_stable.json
```

The first run bootstraps the fixture; later runs diff against it and fail with a
rendered `TrajectoryDiff` when behavior drifts. Flags:

```bash
pytest --update-golden          # re-record goldens after an intended change
pytest --golden-require-exists  # CI: a missing golden is a failure, not a bootstrap
```

Goldens live in a `goldens/` folder beside each test file by default; set
`harnix_golden_dir` in your pytest config for a central location.

## harnix studio: a local web UI

The test-bench is visual by nature — a diff, a divergence point, an A/B verdict
with confidence intervals. **studio** is a local, offline web UI that renders all
of it. Point it at a directory of trajectory / cassette / report JSON:

```bash
harnix serve ./runs        # opens http://127.0.0.1:8765 in your browser
python -m harnix.web ./runs --no-browser --port 8080   # equivalent
```

It has five views:

- **Artifacts** — every run in the directory, badged by kind.
- **Trajectory** — a step timeline: assistant text, tool calls (pretty-printed
  args), tool results (ok/error, latency), and the metrics header.
- **Diff** — pick two runs: the first divergence is highlighted, per-step changes
  are aligned, and metric deltas render as an inline-SVG chart.
- **Report** — an A/B ablation report with accuracy bars, a delta chart with 95%
  CI error bars, and a significance table (adjusted p-values, CIs, verdict).
- **Live run** — drive an agent against a provider and watch its steps stream in
  (Server-Sent Events); the recorded cassette + trajectory are saved into the
  served directory and appear in the sidebar. Defaults to an **offline scripted
  provider** so it works with zero setup — no API key, no network.

Same ethos as the rest of the library: **zero runtime dependencies** (stdlib
`http.server` backend, vanilla-JS frontend — no npm, no framework, no CDN, works
fully offline), and it ships inside the wheel.

> **Security:** studio is a **single-user local tool**. It binds to `127.0.0.1`
> and refuses non-loopback `Host` headers (a DNS-rebinding guard), so a web page
> you visit can't reach it. Don't bind it to a public interface: the live-run
> endpoint can spend API-key money and, with the `files` toolset, run a sandboxed
> shell.

## Sub-agents, skills, memory, MCP

```python
from harnix import SubAgent, SkillRegistry, MCPClient

# Delegate an isolated subtask to a child agent (fresh context, narrow tools).
delegate = SubAgent(model=model, tools=[search], name="researcher").as_tool()

# Progressive-disclosure skills: catalog in context, full body loaded on demand.
skills = SkillRegistry(); skills.load_dir("./skills")
agent = Agent(model=model, tools=[delegate, skills.as_tool()])

# Connect external MCP tool servers (dependency-free stdio client).
with MCPClient(["python", "my_mcp_server.py"]) as mcp:
    agent = Agent(model=model, tools=mcp.list_tools())
```

## Parallel tool execution

When a single step issues several independent (I/O-bound) tool calls — shell,
HTTP, file reads — run them concurrently. Results are always returned in the
original call order, so the recorded trajectory stays stable and replayable:

```python
agent = Agent(model=model, tools=tools, parallel_tools=True, max_workers=8)
```

(Default is sequential; enable it per agent. Your tools and hooks should be
thread-safe when you turn it on.)

## Interfaces: the harness is headless

`harnix` is an *engine*, not a UI. There is no required frontend — you
drive it however you like, because everything an interface needs is in the
`Trajectory` and the `Callback` stream:

- **Library / API** — `agent.run(task)` (the primary surface)
- **CLI** — `harnix inspect|diff|report|demo`
- **Web UI** — `harnix serve ./runs` (studio: a local, offline trace/diff/ablation explorer)
- **TUI** — `harnix tui` (a built-in `curses` terminal UI, stdlib-only)
- **REPL** — see [`examples/repl_frontend.py`](examples/repl_frontend.py); a ~30-line stdin loop *is* a frontend
- **Web backend** — call `agent.run()` from a request handler and stream `Callback` events
- **Chat bots** — wire `agent.run()` to Slack/Telegram/Discord, like OpenHarness's Ohmo

The same headless engine backs all of them; swap the adapter, keep the harness.

### The built-in TUI

```
┌──────────────────────────────────────────────────────────┐
│ you> What is 2 + 3?                                        │
│   · add({'a': 2, 'b': 3})                                  │
│       ↳ [ok] 5                                             │
│ assistant> The answer is 5.                                │
│ [completed] 2 steps · 8 tokens                             │
│                                                            │
│ ───────────────────────────────────────────────── ready  │
│ you> ▋                                                     │
└──────────────────────────────────────────────────────────┘
```

A scrolling transcript streams steps live as the agent runs (PgUp/PgDn to
scroll); type a task and press Enter; `/quit` to exit. Run it with the bundled
offline demo via `harnix tui`, or drive a real model from code:

```python
from harnix import TUIApp, Agent
TUIApp(lambda: Agent(model=my_model, tools=my_tools)).run()
```

The rendering and transcript state are pure functions (tested without a
terminal); only the `curses` driver touches a TTY.

## Evaluating across a task suite

```python
from harnix import Suite, Task, includes, run_suite, Pricing

suite = Suite(name="arith", tasks=[
    Task(id="add", prompt="What is 2 + 3?", scorer=includes("5")),
])
report = run_suite(lambda: build_agent(), suite, pricing=Pricing(0.003, 0.015))
print(report.summary())     # accuracy + process metrics (steps, tokens, cost, redundant calls)
```

Don't want to hand-enter rates? `pricing_for("gpt-4o-mini")` returns an
approximate `Pricing` from a built-in table, so `cost_usd` populates automatically
in `harnix inspect` and the studio UI (inferred from a run's recorded
model). Prices are approximate and dated — override with your own `Pricing`, or
extend the table via `register_pricing(prefix, in_per_1m, out_per_1m)`.

## What's in the box

| Module | Purpose |
| --- | --- |
| `agent` | the loop, `Budget`, retries, doom-loop guard, **parallel tool execution**, `Callback` tracing |
| `tools` | `@tool` decorator + `ToolRegistry` |
| `builtins` | file / shell / memory tools, `standard_toolset` |
| `workspace` | sandboxed filesystem root |
| `permissions` | `PermissionPolicy`, rules, dangerous-command screening |
| `hooks` | pre/post-tool hooks (block/modify), `PermissionHook`, `TruncateOutputHook` |
| `context` | `WindowedContext` token budgeting + offloading |
| `memory` | `Memory` + `Session` persistence |
| `skills` | `SkillRegistry` progressive disclosure |
| `subagent` | `SubAgent` delegation tool |
| `mcp` | dependency-free MCP stdio client |
| `model` | `Model` + `Scripted`/`Function`/`Recording`/`Replay` adapters |
| `providers` | live adapters: OpenAI, Anthropic, Ollama, vLLM, any OpenAI-compatible endpoint (stdlib HTTP, no SDK) |
| `cassette` | content-addressed record/replay store (+ resilient nearest-request replay) |
| `diff` | semantic, step-aligned `TrajectoryDiff` between two runs |
| `regression` | golden-trajectory snapshot testing (`assert_matches_golden`) |
| `compare` | harness A/B ablation (`compare_harnesses`, `ComparisonReport`) |
| `pricing` | approximate per-model price presets so `cost_usd` populates automatically (`pricing_for`, `register_pricing`) |
| `stats` | paired significance: McNemar + sign test + bootstrap CIs + multiple-comparison correction |
| `migrations` | schema versioning + auto-migration for cassettes/goldens |
| `pytest_plugin` | `golden` fixture + `--update-golden` for CI regression testing |
| `eval` / `metrics` / `assertions` | tasks, scorers, reports, trajectory metrics, pytest helpers |
| `config` | layered `HarnessConfig` |
| `tui` | built-in `curses` terminal UI (a frontend over `agent.run()`) |
| `web` | **studio**: local, offline web UI — trajectory/diff/ablation explorer + live run (stdlib server, vanilla JS) |
| `cli` | `harnix version | inspect | diff | migrate | report | demo | tui | serve` |

Run the full offline demo:

```bash
python examples/full_harness.py     # workspace + tools + permissions + hooks + memory + replay
```

## How this relates to other harnesses

Full-stack harnesses like [OpenHands](https://github.com/All-Hands-AI/OpenHands)
and [HKUDS/OpenHarness](https://github.com/HKUDS/OpenHarness) are excellent at
**capability** (rich tools, memory, multi-agent), but ship no record/replay,
deterministic reproducibility, or trajectory-level evaluation.
[Harness-Bench](https://arxiv.org/abs/2605.27922) shows the harness alone drives
substantial, model-dependent swings in task completion — which is exactly why
being able to *reproduce and measure* a run matters. `harnix` aims to be a
complete harness where that reproducibility and measurability are first-class,
while staying small enough (zero dependencies) to read in an afternoon.

See [`docs/RESEARCH.md`](docs/RESEARCH.md) for the full survey, the formal
taxonomy, design rationale, and an honest self-critique; and
[`docs/research-paper.md`](docs/research-paper.md) for the accompanying paper,
*"Deterministic Replay and Trajectory-Level Evaluation for LLM Agent Harnesses."*

## Development

```bash
pip install -e ".[dev]"
pytest                                   # 118 tests
ruff check src tests
python experiments/run_experiments.py    # reproduce the paper's numbers
```

## License

MIT — see [LICENSE](LICENSE).
