Metadata-Version: 2.4
Name: truereward-harwell
Version: 0.3.1
Summary: The recorded-state reader for truereward: run a recorded case on the Harwell engine, compare what it left with the approved baseline, answer with a case result.
Author: truereward contributors
License-Expression: Apache-2.0
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: truereward==0.3.1
Requires-Dist: harwell>=0.29.4
Requires-Dist: jsonschema>=4.18
Dynamic: license-file

# truereward-harwell

The recorded-state reader for `truereward`, over the Harwell engine.

It grades a change to a codebase that has no unit test suite. A case records
what a job leaves behind when the code is right, and claims the parts of it
the case is about: data sets, spool files and tables. Grading runs the same job
against the changed codebase and reads those claims against the approved
baseline.

- **A case is a starting state, a run and claims**, and nothing unclaimed is
  compared. A case passes when every claim of it holds, fails when one does
  not, and is ungradable under a declared code when one could not be evaluated.
  A claim that names an observable and reads no field of it asserts that
  observable whole, byte for byte; an observable no claim names is not read at
  all, so a correct change is never failed on something nobody was asking
  about.
- **A case is data.** A directory with `case.json`, a `start/` snapshot and an
  `expected/` snapshot, validated against the JSON Schema shipped in the
  package. Every field a case reads is resolved to bytes when the case is
  recorded, so grading never opens the candidate's source to decide what is
  compared.
- **`grade_case(case, codebase)`** runs the case twice in a contained run,
  reads the claims over the first end state and the baseline, and answers with
  a `truereward.CaseResult`. The two runs are compared with each other, over
  the same claimed observables, under the author's own mask, which answers the
  determinism check.
- **`record.record(...)`** records a case: refuse a case that names no claim,
  run the job twice, refuse a run that did not complete or that varied in
  something the case claims, keep what it left as the approved baseline,
  freeze every field map, and read every claim back over the baseline it just
  approved.
- **`selftest.run(...)`** asks whether a case set would catch a bug: the
  reference change must score 1.0, the unchanged codebase 0.0, and every
  seeded defect less than 1.0.

A task's test file:

```python
from functools import partial
from pathlib import Path

import pytest
from truereward_harwell import cases, grade_case

CASES = Path(__file__).parent / "cases"


@pytest.mark.parametrize("case", cases(CASES), ids=lambda c: c.id)
def test_case(case, grade):
    grade(case, partial(grade_case, codebase="/grade"))
```

Run it with `pytest -p truereward_harwell`, which brings the `truereward`
plugin with it.

Licensed under Apache-2.0.
