Metadata-Version: 2.4
Name: codeer-edu-guardrails
Version: 0.1.0a1
Summary: Educational response checks with five classifications and grounded diagnostics.
Project-URL: Homepage, https://www.codeer.ai
Author: Codeer.AI
Classifier: Development Status :: 3 - Alpha
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Education
Requires-Python: >=3.11
Requires-Dist: httpx<1,>=0.27
Requires-Dist: pydantic<3,>=2.8
Description-Content-Type: text/markdown

# Codeer Education Guardrails

Check a candidate tutoring response and receive five classifications with
diagnostics and verifiable input quotes. Python 3.11 or newer is required.

## Install

```sh
python -m pip install --pre codeer-edu-guardrails
```

Use a **Codeer API key** issued for this service. A Gemini provider key is not
accepted. You do not need to install the Codeer CLI or run a model locally.

```python
from getpass import getpass
from codeer_edu_guardrails import Guardrails

with Guardrails(api_key=getpass("Codeer API key: ")) as client:
    result = client.check(
        conversation_history=input("Question and tutoring history: "),
        candidate_response=input("Candidate tutor response: "),
        # student_context="Optional background supplied by the caller",
    )
    print(result.model_dump_json(indent=2))
```

Alternatively set `CODEER_API_KEY` securely in your environment and construct
`Guardrails()` without arguments. Never commit keys into source code.

## Results

`result.checks` contains exactly these original MRBench criteria:

| Criterion | Labels |
| --- | --- |
| `Mistake_Identification` | `Yes`, `To some extent`, `No` |
| `Mistake_Location` | `Yes`, `To some extent`, `No` |
| `Revealing_of_the_Answer` | `No`, `Yes (and the answer is correct)`, `Yes (but the answer is incorrect)` |
| `Providing_Guidance` | `Yes`, `To some extent`, `No` |
| `Actionability` | `Yes`, `To some extent`, `No` |

Each item has `label`, `confidence` (`high`, `medium`, `low`), `diagnosis`,
`evidence` (a list of `{source, quote}`), and `limitations`. The SDK checks that
each quote occurs verbatim in its specified input field. It does not verify the
semantic correctness of the diagnosis. Confidence is not a calibrated probability.

```python
guidance = result.checks["Providing_Guidance"]
print(guidance.label, guidance.diagnosis)
```

Request, chat and response IDs support troubleshooting. `scheme_version`
identifies the SDK's response contract. `model`, `agent_history_id` and
`cost_credits` are null when the service does not expose them. An enclosing JSON
Markdown fence may be removed; `normalized_json_fence` reports this formatting
normalization. Label strings and input text are never rewritten.
`service_completion_verified` indicates whether the service supplied an explicit
successful inference status. When status metadata is absent, the client validates
the synchronous response and its full result but does not invent that status.

## Errors and timeouts

```python
from codeer_edu_guardrails import GuardrailsError

try:
    with Guardrails() as client:
        result = client.check(conversation_history=history, candidate_response=candidate)
except GuardrailsError as error:
    print(type(error).__name__, error.request_id, error.chat_id)
```

Errors distinguish invalid input, authentication/permission, rate limits,
timeouts, service failure and invalid model output. A successful result always
contains all five checks. Missing evidence or invalid labels cause an error,
not fabricated classifications. There is no content-level abstention category.

`Guardrails(timeout=120)` sets the HTTP inactivity timeout in seconds, which is
not an end-to-end response-time guarantee. No requests are automatically retried;
a timeout can occur after the server has accepted work. Preserve the request/chat
IDs and inspect the outcome before retrying.

## Service and scope

Each check sends your supplied content to Codeer and creates a separate saved chat.
The backend uses a Gemini 3.8 Flash Agent. The generator retains responsibility
for deciding whether and how to revise its response. This SDK does not rewrite
answers or produce an overall pass/revise decision.

The alpha is for integration and local development. Classification and diagnostic
quality are not yet validated for release thresholds or multiple languages.
Content language is not restricted. Curriculum alignment, personalized help
dosage and complete answer correctness are not separately validated capabilities.
Do not infer them directly from these five labels.

`agent_id=` (or `CODEER_GUARDRAILS_AGENT_ID`) and `base_url=` allow an explicitly
configured Codeer deployment. Custom agents must implement the same five-check
contract. The external Chat API executes the published configuration; selecting
an Agent ID does not pin an immutable AgentHistory. Published service changes
must therefore be managed separately from SDK versions.

The criterion names follow [MRBench](https://github.com/kaushal0494/UnifyingAITutorEvaluation).
This package distributes no benchmark records, gold labels, customer data or keys.
