Metadata-Version: 2.5
Name: viktron-sdk
Version: 0.8.0
Summary: Production-grade observability and governance SDK for AI agents
Project-URL: Homepage, https://viktron.ai
Project-URL: Documentation, https://viktron.ai/docs
Project-URL: Repository, https://github.com/vikasvardhanv/viktron-agentsALR
License: Apache-2.0
License-File: LICENSE
Keywords: agents,ai,governance,llm,observability,opentelemetry
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Software Development :: Libraries
Requires-Python: >=3.9
Requires-Dist: httpx>=0.24
Provides-Extra: ag2
Requires-Dist: ag2[tracing]; extra == 'ag2'
Requires-Dist: opentelemetry-exporter-otlp-proto-http>=1.24; extra == 'ag2'
Requires-Dist: opentelemetry-sdk>=1.24; extra == 'ag2'
Provides-Extra: all
Requires-Dist: anthropic>=0.20; extra == 'all'
Requires-Dist: langchain>=0.1; extra == 'all'
Requires-Dist: openai>=1.0; extra == 'all'
Requires-Dist: opentelemetry-exporter-otlp-proto-http>=1.24; extra == 'all'
Requires-Dist: opentelemetry-sdk>=1.24; extra == 'all'
Provides-Extra: anthropic
Requires-Dist: anthropic>=0.20; extra == 'anthropic'
Provides-Extra: langchain
Requires-Dist: langchain>=0.1; extra == 'langchain'
Provides-Extra: openai
Requires-Dist: openai>=1.0; extra == 'openai'
Provides-Extra: otel
Requires-Dist: opentelemetry-exporter-otlp-proto-http>=1.24; extra == 'otel'
Requires-Dist: opentelemetry-sdk>=1.24; extra == 'otel'
Provides-Extra: pydantic-ai
Requires-Dist: opentelemetry-exporter-otlp-proto-http>=1.24; extra == 'pydantic-ai'
Requires-Dist: opentelemetry-sdk>=1.24; extra == 'pydantic-ai'
Requires-Dist: pydantic-ai-slim>=1.0; extra == 'pydantic-ai'
Description-Content-Type: text/markdown

# Viktron SDK

[![PyPI](https://img.shields.io/pypi/v/viktron-sdk.svg)](https://pypi.org/project/viktron-sdk/)
[![License](https://img.shields.io/badge/license-Apache--2.0-blue.svg)](LICENSE)
[![Python](https://img.shields.io/pypi/pyversions/viktron-sdk.svg)](https://pypi.org/project/viktron-sdk/)

Framework-agnostic governance SDK for AI agents — policy enforcement, telemetry, and real-time observability for production AI systems.

See [BENCHMARK.md](BENCHMARK.md) for a reproducible runtime-overhead benchmark, [examples/](examples/) for a runnable quickstart with no API key required, and [CONTRIBUTING.md](CONTRIBUTING.md) to contribute.

## Install

```bash
pip install viktron-sdk
# With optional provider extras:
pip install "viktron-sdk[openai]"
pip install "viktron-sdk[anthropic]"
pip install "viktron-sdk[langchain]"
pip install "viktron-sdk[otel]"          # any OpenTelemetry-instrumented framework
pip install "viktron-sdk[pydantic-ai]"
pip install "viktron-sdk[ag2]"
pip install "viktron-sdk[all]"
```

## Quick Start

Two lines to your first trace:

```python
import viktron_sdk
viktron_sdk.init(api_key="vk_live_...")  # or set VIKTRON_API_KEY and call init()
```

### `@viktron_sdk.trace` — one-line decorator instrumentation, with a real run tree

```python
import viktron_sdk

viktron_sdk.init(
    api_key="vk_live_...",
    agent_id="my-research-agent",
    display_name="Research Agent v2",
    framework="crewai",
    llm_provider="anthropic",
    llm_model="claude-3-sonnet",
)

# Works on both sync and async functions. Nested calls form a run tree
# automatically — no manual trace_id/parent plumbing required.
@viktron_sdk.trace(as_type="tool", skip_args=["api_key"])
async def search_web(query: str, api_key: str):
    ...

@viktron_sdk.trace(as_type="agent")
async def run_agent(task: str):
    # search_web's span is recorded as a CHILD of run_agent's span —
    # same trace_id, parent_span_id set automatically — purely from
    # ambient context. Works through asyncio.gather() fan-out too.
    return await search_web(task)

# Flush before exit (atexit handler covers normal exits automatically)
viktron_sdk.force_flush()
# Or fully shut down the SDK:
viktron_sdk.stop_tracing()
```

`trace` is the primary name; `observe` is kept as an alias (same function) for existing code.

### Env-var config (no key in code)

```bash
export VIKTRON_API_KEY="vk_live_..."
export VIKTRON_ENDPOINT="https://api.viktron.ai"  # optional, this is the default
```

```python
import viktron_sdk
viktron_sdk.init()  # picks up both from the environment
```

If neither `api_key=` nor `VIKTRON_API_KEY` is set, `init()` logs one clear warning
and the SDK runs in **no-op mode**: every `@trace`/wrapped LLM call still executes
your code normally, nothing is recorded or sent anywhere, nothing raises.

### Debug mode — see traces locally

Pass `debug=True` (or set `VIKTRON_DEBUG=1`) to print every trace to stderr as it
happens — instant confirmation the SDK is wired up, no dashboard needed:

```python
viktron_sdk.init(api_key="vk_live_...", debug=True)
```

```text
[viktron] initialized — agent=my-research-agent -> https://api.viktron.ai
[viktron] llm_call model=gpt-4o in=150 out=75 cost=$0.0012 latency=320ms status=ok
```

### `wrap_openai` / `wrap_anthropic` — explicit per-client instrumentation

`init()`'s `auto_instrument_openai`/`auto_instrument_anthropic` (both on by default)
already patch every client globally. Use `wrap_openai`/`wrap_anthropic` instead when
you want to instrument one specific client instance — e.g. you only want to trace
calls through a particular client, or you called `init(auto_instrument_openai=False)`:

```python
from openai import OpenAI
from viktron_sdk import wrap_openai

client = wrap_openai(OpenAI())
response = client.chat.completions.create(model="gpt-4o", messages=[...])  # auto-traced
```

```python
from anthropic import Anthropic
from viktron_sdk import wrap_anthropic

client = wrap_anthropic(Anthropic())
response = client.messages.create(model="claude-sonnet-5", messages=[...])
```

Both work on sync and async clients (`AsyncOpenAI`/`AsyncAnthropic`), participate in
the same run tree as `@viktron_sdk.trace` (a wrapped call made inside a `@trace`d
function is recorded as its child), and cover streaming calls via the httpx-level
patch (`auto_instrument_httpx`, on by default) rather than duplicating incomplete
usage capture at the client-wrapper layer.

Cost isn't computed client-side — the backend computes `cost_usd` from model + token
counts at ingest time from its own pricing table, so it can't drift from what your
dashboard shows. It shows up in AgentIRL automatically; it's just not in the return
value of the wrapped call.

### Connect any agent with zero code — the Gateway

If your agent uses a custom transport the SDK can't auto-patch (or you just want
zero code), point its LLM `base_url` at the Viktron gateway. Your provider key
stays in `Authorization` (never stored); only your workspace key sits in the path:

```bash
export OPENAI_BASE_URL="https://api.viktron.ai/v/vk_live_xxx/openai/v1"
# keep the provider's normal path suffix (/v1 for OpenAI-wire)
# named upstreams: openai | anthropic | nous | openrouter | groq | together | deepseek
# or /passthrough to forward to any OpenAI/Anthropic/Gemini-compatible endpoint
```

`observe` shorthand types: `"llm"` → `llm_call`, `"tool"` → `tool_call`, `"guardrail"` → `guardrail_check`, `"agent"` → `agent_response`.

### CLI

```bash
# Check configuration and API connectivity
viktron doctor

# Refresh cached pricing data
viktron pricing
```

### Telemetry (send agent events to your Viktron dashboard)

```python
from viktron_sdk import ViktronTelemetry

tel = ViktronTelemetry(api_key="vk_live_...", agent_slug="my-agent")

# Record a task
tel.record_task("task-123", status="completed", duration_ms=4200, cost_usd=0.003)

# Use a context-managed span
with tel.span("run_campaign") as span:
    span.set_attribute("platform", "meta")
    result = run_campaign()
    span.set_output(result)

tel.close()  # flush remaining events
```

### Policy Guard (intercept LLM calls with governance rules)

```python
import openai
from viktron_sdk import ViktronGuard

client = ViktronGuard.wrap(
    openai.OpenAI(),
    api_key="vk_live_...",
    agent_id="sales-agent-prod",
)

# All create() calls are now policy-checked
response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Hello"}],
)
```

Works with OpenAI (sync & async), Anthropic, Cohere, and any client with `.create()` / `.generate()` call patterns.

### Framework Integrations

```python
# LangChain
from viktron_sdk.integrations.langchain import ViktronCallbackHandler
handler = ViktronCallbackHandler(api_key="vk_live_...", agent_id="my-chain")

# CrewAI
from viktron_sdk.integrations.crewai import ViktronCrewAIObserver
observer = ViktronCrewAIObserver(api_key="vk_live_...", agent_id="my-crew")

# AutoGen
from viktron_sdk.integrations.autogen import ViktronAutoGenHook
hook = ViktronAutoGenHook(api_key="vk_live_...", agent_id="my-autogen")
```

> **Run-tree note**: these three framework callbacks record telemetry through
> `ViktronTelemetry`, a separate code path from `@viktron_sdk.trace`/`wrap_openai`'s
> context-based run-tree linkage — spans they record don't currently nest under a
> `@trace`-decorated caller. If you need the framework's LLM calls to show up as
> children in a specific run tree, wrap the framework invocation itself in
> `@viktron_sdk.trace(as_type="agent")` (e.g. wrap your `crew.kickoff()` call) —
> everything that callback fires during that call will correctly attach as
> children via `wrap_openai`/`wrap_anthropic` or the global auto-patch, since
> those two DO share the run-tree context.

### OpenTelemetry frameworks — Pydantic AI, AG2, and anything that emits GenAI spans

Frameworks that trace themselves with OpenTelemetry need no Viktron-specific
hooks: point their tracer at Viktron and every agent run, model call, tool call
and token lands in your dashboard, attributed to the agent that made it.

```bash
pip install "viktron-sdk[pydantic-ai]"                    # Pydantic AI
pip install "viktron-sdk[otel]" "ag2[openai,tracing]<1"   # AG2 classic (import autogen)
pip install "viktron-sdk[ag2]"                            # AG2 1.x (import ag2)
```

```python
# Pydantic AI — one call instruments every Agent (VIKTRON_API_KEY from the env)
from viktron_sdk.integrations.pydantic_ai import instrument
instrument(agent_id="support-bot")

# AG2 classic — agents, or a group-chat pattern
from viktron_sdk.integrations.ag2 import instrument
instrument([planner, coder, user_proxy], agent_id="research-crew")

# AG2 1.x — per-agent middleware
from viktron_sdk.integrations.ag2 import telemetry_middleware
agent = Agent("planner", prompt="...", config=cfg,
              middleware=[telemetry_middleware(agent_name="planner")])

# Any other OpenTelemetry-instrumented framework
from viktron_sdk.otel import configure
provider = configure(agent_id="my-agent")   # a TracerProvider that ships to Viktron
```

Tokens and cost come from model-call spans only, so a framework that also
reports running totals on its agent span is not billed twice. Use this route
**or** `viktron_sdk.init()`'s HTTP auto-capture for a given agent, not both —
each would record the same LLM call.

### `@governed` — a policy check or human approval before any tool runs

```python
from viktron_sdk import governed

@governed("refund")   # waits for a human if a policy asks for approval
def issue_refund(order_id: int, amount: float) -> str:
    ...
```

Before the function runs, Viktron checks it against your team's policies:
allowed → it runs; denied → `ViktronPolicyViolation`; needs approval → it
appears in the dashboard's Approvals queue and the call waits (`wait=True`, the
default) until someone decides. `destructive`, `delete`, `spend`, `payment` and
`pii_access` actions always need approval unless a permission rule explicitly
allows them. The decorator keeps the function's signature, so frameworks that
build tool schemas from it still see the same parameters. Works without
`init()` — it reads `VIKTRON_API_KEY` from the environment.

## Code-context compaction (graft)

The biggest source of wasted input tokens in a coding agent is the whole-file
dump: the agent reads 800 lines of source to use 12 of them, pays for all 800,
and does it again on the next turn.

[Graft](https://github.com/trailhq/Graft) builds a structural graph of your repo
— symbols, call edges, exact `file:line` spans — locally, with no LLM and no API
key. The SDK puts that graph in front of the wire: before a request leaves your
machine it finds whole-file dumps in the prompt and replaces each one with the
spans that actually matter for the task, keeping `file:Lstart-Lend` so the model
can ask for more.

Measured on a 266-line React component with the question *"why does the cursor
trail flicker at high mouse speed?"*: **2,459 tokens -> 900, a 63% saving**, with
graft's top-ranked span being the `handleMouseMove` handler the question was
actually about.

This runs client-side because it has to — graft indexes a repo **on disk**, and
the Viktron gateway is a cloud proxy that never sees your source. The saving is
reported on the trace as `metadata.graft_context` and shows up under
`realized.sources.graft_context` in the active-optimizations dashboard.

### Setup

```bash
npm install -g @nanonets/graft   # or a project devDependency — both are found
graft build                      # deterministic, no API key, $0
```

```python
import viktron_sdk
viktron_sdk.init(api_key="vk_live_...", graft_context=True)
# or leave it out of code entirely: VIKTRON_GRAFT_CONTEXT=1
```

**Off by default.** With no graft binary, no built `graft/` index, or no repo, it
is a silent no-op. Any error leaves the request byte-for-byte unchanged, and the
final user turn — your actual instruction — is never rewritten.

| Env var | Default | What it controls |
|---|---|---|
| `VIKTRON_GRAFT_CONTEXT` | off | Master switch |
| `VIKTRON_GRAFT_REPO` | auto | Pin the repo root (containers, task runners) |
| `VIKTRON_GRAFT_BIN` | auto | Pin the graft binary |
| `VIKTRON_GRAFT_MIN_BLOB_LINES` | `40` | Smallest dump worth compacting |
| `VIKTRON_GRAFT_MIN_SAVING_TOKENS` | `200` | Discard a pack that saves less |
| `VIKTRON_GRAFT_MAX_FILES` | `8` | Cap graft calls per request |
| `VIKTRON_GRAFT_MAX_SPAN_LINES` | `40` | Cap inlined lines per span |
| `VIKTRON_GRAFT_PACK_LIMIT` | `6` | Ranked spans per file |
| `VIKTRON_GRAFT_TIMEOUT_S` | `8` | Per-invocation timeout |

Not using the auto-patch? Compact a body yourself:

```python
from viktron_sdk import graft_compact_body
stats = graft_compact_body(body)      # rewrites body["messages"] in place
if stats:
    print(stats.tokens_saved, stats.pct)
```

## API Keys

Generate your API key at [app.viktron.ai](https://app.viktron.ai) → Settings → API Keys.

Keys prefixed `vk_live_` are production; `vk_test_` are for testing.

## Links

- [Dashboard](https://app.viktron.ai)
- [Documentation](https://viktron.ai/docs)
- [GitHub](https://github.com/vikasvardhanv/viktron-agentsALR)
