Metadata-Version: 2.4
Name: pnotp
Version: 0.4.0
Summary: Official Python client for the P¬P platform — projects, training, deployments, and inference from one API key.
Project-URL: Homepage, https://pnotp.ai
Author-email: PnotP <support@pnotp.ai>
Requires-Python: >=3.9
Requires-Dist: httpx<0.29,>=0.24
Provides-Extra: private
Requires-Dist: numpy>=1.23; extra == 'private'
Provides-Extra: server
Requires-Dist: fastapi<1.0,>=0.110; extra == 'server'
Requires-Dist: pillow>=9.0; extra == 'server'
Requires-Dist: pydantic<3.0,>=2.0; extra == 'server'
Requires-Dist: python-multipart>=0.0.6; extra == 'server'
Requires-Dist: uvicorn<0.29,>=0.23; extra == 'server'
Description-Content-Type: text/markdown

# pnotp

Official Python client for the **P¬P** platform. One account API key drives the
whole platform: create projects and models, train (with pause / resume / cancel),
deploy models, and call live inference endpoints.

## Install

```bash
pip install pnotp
```

## Authenticate

Mint an account API key in **Studio → Settings → API keys** (it starts with
`pnotp_sk_` and is shown once). Then:

```bash
export PNOTP_API_KEY="pnotp_sk_..."
```

```python
import pnotp
px = pnotp.Client()                       # reads PNOTP_API_KEY
# or: px = pnotp.Client(api_key="pnotp_sk_...")
```

Point it at a local dev server with `base_url=` or `PNOTP_BASE_URL`
(e.g. `http://localhost:8000/v1`). The default is the hosted backend.

## Train a model end to end

```python
import pnotp
px = pnotp.Client()

project = px.projects.create(name="Chest X-ray", task_type="binary_image_classification")

model = px.models.create(
    project["id"],
    name="cnn-v1",
    architecture=pnotp.architectures.image_classifier(num_classes=2),
)

# dataset = a .zip of class-named folders (image) or a .csv (tabular)
job = px.train(model.id, dataset="train.zip", epochs=5)
job.wait(on_progress=lambda j: print(j.status))

# pause / resume / cancel any time
# job.pause(); job.resume(); job.cancel()
```

## Deploy + call inference

```python
dep = px.deployments.upload(project["id"], name="prod", checkpoint="model.pt")
deployment_id = dep["deployment"]["id"]

result = px.predict(deployment_id, "xray.jpg")
print(result)                              # {"predictions": [...], "class_names": [...]}

px.deployments.stats(deployment_id, window="24h")
```

External consumers can call your endpoint with a per-deployment key (`pnp_...`):

```python
key = px.deployments.create_key(deployment_id)["api_key"]
px.deployments.predict(deployment_id, "xray.jpg", api_key=key)
```

## Data-compliant inference (zero retention)

Pass `compliant=True` at deploy time to make an endpoint **data-compliant**. The
platform then guarantees, on every call to that endpoint:

- **Zero retention** — the input you send is never persisted; the temp file is
  deleted before the response returns.
- **Redacted monitoring** — error text is kept out of the stats/monitoring log
  (only status code and latency are recorded).
- **An audit receipt** — each response carries a `compliance` block
  (`{"compliant": true, "retention": "none", "deployment_id", "timestamp"}`).

```python
dep = px.deployments.upload(project["id"], name="prod", checkpoint="model.pt", compliant=True)
# or a catalog model:
dep = px.deployments.deploy_open_source(project["id"], "mobilenet_v3_small", compliant=True)
# or flip it on an existing endpoint:
px.deployments.set_compliant(deployment_id, True)

result = px.predict(deployment_id, "xray.jpg")
print(result["compliance"])   # {'compliant': True, 'retention': 'none', ...}
```

It's a platform capability: `compliant=True` raises `APIError` (409) when the
platform hasn't enabled it. Probe first with
`px.deployments.compliance_available()`. `compliant=False` (the default) is
unchanged behaviour.

**Agents inherit it for free.** `pnotp.model` runs inference through the same
path, so an agent that calls a compliant endpoint is itself data-compliant —
compliance is a property of the *deployment*, not the agent or the call site. To
ship a data-compliant agent into a customer's process, deploy its models with
`compliant=True`; nothing else changes. See `docs/data-compliance.md`.

## Data-compliant LLM + agents

PnotP also hosts a **zero-retention LLM brain**. The inference layer stores
nothing — neither the prompt nor the completion — and each call returns a
`receipt` you can keep as the audit artifact:

```python
out = px.chat("In one sentence, what is a master's degree?")
print(out["completion"])
print(out["receipt"])   # {"mode": "compliant", "zero_retention": True,
                        #  "plaintext_stored": "none", "cost": {...}}
```

Use the same brain for an agent with `model="pnotp-compliant"` (build + run in
four lines):

```python
a   = px.agents.create("Advisor")
px.agents.publish(a["id"], model="pnotp-compliant", tools=["programs_db"])
run = px.agents.run(a["id"], "I love AI and maths — which programs fit?", wait=True)
print(run.final_answer())
```

A compliant agent run keeps no question, no reasoning, and no loop state — only
metered metadata + the final answer (resume of a compliant run is rejected). A
BYO-key brain still routes tokens to *that* vendor. **Tool-calling** works via the
SDK (`px.chat(..., tools=[...])`) or the **OpenAI-compatible** `POST /v1/chat/completions`
(drop-in for the Vercel AI SDK / LangChain / `openai` — base URL `<api>/v1`, model
`pnotp-compliant`). **Pricing** is **$1.11/1M output, $0.11/1M input** (≈1.8× under
OpenRouter), metered per token — **billed against your credits** (metering is live; an empty balance
returns `402`). Live rates (no key): `GET /v1/llm/pricing`. Full guide: `docs/data-compliance.md`.

**Vision (multimodal).** The served brain is a vision model (Qwen3-VL-8B), so an
`image` is understood VISUALLY — it sees the picture, not just text in it.
Documents are parsed and audio is speech-to-text. Image bytes go straight to our
own GPU and are never stored or logged.

```python
px.chat("What is in this image?", image="photo.jpg")     # real vision
px.chat("Summarize this contract", file="contract.pdf")  # document parse
px.chat("What does the customer want?", audio="voicenote.ogg")  # speech-to-text
# or the standard OpenAI image_url content-part form:
px.chat([{"role": "user", "content": [
    {"type": "text", "text": "Describe this"},
    {"type": "image_url", "image_url": {"url": "data:image/png;base64,...."}}]}])
```

## Agents

An **agent** is a BYO reasoning model (the "brain") driving a tool-calling loop
over tools you register, plus platform built-ins. **PnotP hosts no LLMs** — the
brain is your **Gemini** / **DeepSeek** key (any OpenAI-style `/chat/completions`
with tool-calling), or a text-gen model you **deployed on PnotP**.

```python
import pnotp

agent = pnotp.Agent("gemini:gemini-2.5-pro",          # reads GEMINI_API_KEY
                    system="You are a careful research assistant.")

@agent.tool
def search_pubmed(query: str, max_results: int = 5) -> list[dict]:
    "Search PubMed and return the top hits."          # docstring -> tool description
    return my_client.search(query, k=max_results)     # JSON Schema is derived from the signature

run = agent.run("Recent trials on pediatric pneumonia from chest X-rays?")
print(run.output, run.stop_reason)                    # final answer + "stop" | "max_steps"

# stream events for a live UI
for event in agent.stream("..."):
    ...   # TextDelta / ToolCallStarted / ToolCallFinished / StepBoundary / Done
```

The **`pnotp.model`** built-in lets the agent call any model you've deployed
(Tier 1; billed per platform inference). The brain can be your own deployment via
`pnotp.endpoints.deployed("my-llm")` or the `"pnotp:<deployment_id>"` shorthand.
Structured output: pass `response_format=MyDataclass` to parse the final answer.

The loop is bounded and observable: `max_steps`, per-tool `tool_timeout`,
`max_tool_retries` (a schema-invalid or timed-out call is re-asked without
spending the step budget), and `on_tool_error` (`"return"` feeds the error back to
the brain — the resilient default — or `"raise"`).

## Help & Toby

The SDK documents itself — offline, no API key, no network. It's the fastest way
for you (or a coding agent like Claude Code / Codex) to get unstuck:

```python
import pnotp

pnotp.help()                       # overview + the list of topics
pnotp.help("agents")               # a topic page
pnotp.help.search("pause a training job")   # search topics + the live API reference
pnotp.help(pnotp.Agent)            # any symbol's real signature + docstring
pnotp.help.errors                  # the exception catalog
```

The API reference is generated by introspection, so it always matches your
installed version. Every exception also carries a `.help` pointer.

**Toby** is the hosted assistant — natural-language answers, authenticated by your
account key (needs credits). It falls back to local doc search when there's no
key/credits/network, so it never raises:

```python
print(pnotp.help.toby("How do I deploy a model and then call it?"))

chat = pnotp.help.toby.session()   # multi-turn
chat.ask("How do I make an image classifier?")
chat.ask("And how do I pause its training?")
```

Inside a runtime agent, let the brain self-serve docs when it's unsure:

```python
agent = pnotp.Agent("gemini:gemini-2.5-pro", tools=[pnotp.help.as_tool()])
```

See `AGENTS.md` for a coding-agent-oriented cheat sheet.

## Surface

| Area        | Calls |
|-------------|-------|
| Projects    | `projects.list/create/get` |
| Models      | `models.create/list/get/add_version/versions` · `model.train(...)` |
| Training    | `train(...)` → `TrainingJob` · `job.wait/progress/pause/resume/cancel` · `jobs.list/get` |
| Deployments | `deployments.upload/list/get/predict/stats/pause/resume/delete` · `create_key/list_keys/revoke_key` |
| Compliant LLM | `chat(text_or_messages, ...)` → `{completion, usage, model, receipt}` (zero-retention; pricing MODELED) |
| Hosted agents | `agents.create/get/list/publish(model="pnotp-compliant", tools=, ...)/run(..., wait=)` → `run.final_answer/steps/progress` |
| Credits     | `credits.balance()` |
| API keys    | `api_keys.create/list/revoke` |
| Agents      | `Agent(model, tools=, system=, ...)` · `agent.run/stream` · `@agent.tool` / `@pnotp.tool` · `pnotp.endpoints.gemini/deepseek/openai_compatible/deployed` · `pnotp.model` |
| Help        | `help()` / `help("topic")` / `help(symbol)` · `help.search/topics/errors` · `help.toby("...")` / `help.toby.session()` · `help.as_tool()` |

Errors raise `pnotp.APIError` subclasses: `AuthError` (401/403),
`NotFoundError` (404), `InsufficientCreditsError` (402). Agent-loop failures raise
`pnotp.AgentError` subclasses: `MaxStepsExceeded`, `ToolError`, `ModelEndpointError`.

---

The legacy `pnotp_api` package (a standalone PathoVision wrapper) still lives in
this repo and is unrelated to the platform client above.
