Metadata-Version: 2.4
Name: audacity-sdk
Version: 0.6.0
Summary: Audacity Investments AI SDK — Bedrock-parity client for the Audacity LLM gateway
Author-email: Audacity Investments <eng@audacityinvestments.com>
License: Audacity Investments SDK License
        
        Copyright (c) 2026 Audacity Investments. All rights reserved.
        
        This software is proprietary to Audacity Investments. Permission is granted
        to download, install, and use this software solely to access services
        provided by Audacity Investments, subject to your agreement with Audacity
        Investments.
        
        You may not modify, distribute, sublicense, or create derivative works of
        this software, in whole or in part, except as expressly permitted in
        writing by Audacity Investments.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL
        AUDACITY INVESTMENTS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY
        ARISING FROM THE USE OF THE SOFTWARE.
        
Project-URL: Homepage, https://portal.audacityinvestments.com
Project-URL: Repository, https://github.com/Audacity-Investments/audacity-sdk
Keywords: audacity,llm,ai,bedrock
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Operating System :: OS Independent
Classifier: Typing :: Typed
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Dynamic: license-file

# audacity-sdk

Python client for the Audacity Investments LLM gateway.  Exposes the same
**Amazon Bedrock Converse** surface so teams migrating off Bedrock can swap
the client constructor and keep the rest of their call-sites unchanged.

---

## Installation

```bash
pip install audacity-sdk
```

Zero runtime dependencies — stdlib only. Requires Python 3.9+.

---

## Quick start

### Non-streaming (Converse)

```python
from audacity import Audacity

client = Audacity(api_key="aireserve_api_…")   # or set AIRESERVE_API_KEY

response = client.converse(
    modelId="gpt-5.4-mini",
    messages=[{"role": "user", "content": [{"text": "What is 2+2?"}]}],
    inferenceConfig={"maxTokens": 256, "temperature": 0.0},
)

print(response["output"]["message"]["content"][0]["text"])
print(response["stopReason"])   # "end_turn"
print(response["usage"])        # {"inputTokens": …, "outputTokens": …, "totalTokens": …}
print(response["metrics"])      # {"latencyMs": …}
```

### Streaming (ConverseStream)

```python
from audacity import Audacity

client = Audacity()

stream_response = client.converse_stream(
    modelId="gpt-5.4-mini",
    messages=[{"role": "user", "content": [{"text": "Write me a haiku."}]}],
)

for event in stream_response["stream"]:         # boto3 parity: response["stream"]
    if "contentBlockDelta" in event:
        delta = event["contentBlockDelta"]["delta"]
        print(delta.get("text", ""), end="", flush=True)

print()  # newline at end
```

---

## Migrating from boto3 bedrock-runtime

```python
# BEFORE — boto3
import boto3
client = boto3.client("bedrock-runtime", region_name="us-east-1")

response = client.converse(
    modelId="anthropic.claude-3-sonnet-20240229-v1:0",
    messages=[{"role": "user", "content": [{"text": "Hi"}]}],
)

# AFTER — Audacity SDK
from audacity import Audacity
client = Audacity(api_key="aireserve_api_…")   # only line that changes

response = client.converse(
    modelId="gpt-5.4-mini",                   # use Audacity model ID
    messages=[{"role": "user", "content": [{"text": "Hi"}]}],
)
```

Streaming diff:

```python
# BEFORE — boto3
stream_resp = client.converse_stream(modelId=…, messages=…)
for event in stream_resp["stream"]:
    …

# AFTER — Audacity SDK (identical call-site)
stream_resp = client.converse_stream(modelId=…, messages=…)
for event in stream_resp["stream"]:
    …
```

---

## Tool use

```python
response = client.converse(
    modelId="gpt-5.4-mini",
    messages=[{"role": "user", "content": [{"text": "What's the weather in NYC?"}]}],
    toolConfig={
        "tools": [{
            "toolSpec": {
                "name": "get_weather",
                "description": "Get current weather for a city",
                "inputSchema": {
                    "json": {
                        "type": "object",
                        "properties": {"city": {"type": "string"}},
                        "required": ["city"],
                    }
                },
            }
        }],
        "toolChoice": {"auto": {}},
    },
)

# Assistant responds with a tool call
tool_use = response["output"]["message"]["content"][0]["toolUse"]
print(tool_use["name"])   # "get_weather"
print(tool_use["input"])  # {"city": "NYC"}

# Send tool result back
response2 = client.converse(
    modelId="gpt-5.4-mini",
    messages=[
        {"role": "user",  "content": [{"text": "What's the weather in NYC?"}]},
        {"role": "assistant", "content": response["output"]["message"]["content"]},
        {"role": "user",  "content": [
            {"toolResult": {
                "toolUseId": tool_use["toolUseId"],
                "content": [{"text": "Sunny, 72°F"}],
            }}
        ]},
    ],
)
```

---

## Images (vision models)

Bedrock-style `image` content blocks are supported in user messages. Pass raw
bytes (encoded for you) or a URL (Audacity extension):

```python
with open("chart.png", "rb") as f:
    image_bytes = f.read()

response = client.converse(
    modelId="gpt-5.5",
    messages=[{
        "role": "user",
        "content": [
            {"text": "What does this chart show?"},
            {"image": {"format": "png", "source": {"bytes": image_bytes}}},
        ],
    }],
)

# Or reference a hosted image directly (not available in Bedrock):
# {"image": {"format": "jpeg", "source": {"url": "https://example.com/photo.jpg"}}}
```

`format` is one of `png`, `jpeg`, `gif`, `webp`. Use a vision-capable model.

---

## Video (Gemini models)

Bedrock-style `video` content blocks are supported in user messages. Video input
is only available on the Gemini family (`gemini-2.5-flash`, `gemini-2.5-pro`,
`gemini-3-flash-preview`); every other model rejects video with an HTTP 400.

```python
with open("demo.mp4", "rb") as f:
    video_bytes = f.read()

response = client.converse(
    modelId="gemini-2.5-flash",
    messages=[{
        "role": "user",
        "content": [
            {"text": "Summarize what happens in this video."},
            {"video": {"format": "mp4", "source": {"bytes": video_bytes}}},
        ],
    }],
)
```

`format` is one of `mp4`, `mov`, `mkv`, `webm`, `flv`, `mpeg`, `mpg`, `wmv`,
`three_gp`. Inline (`bytes`) video is base64-encoded into the request — keep
those files under ~20 MB.

For larger files (up to 1 GB), upload once and reference by URI:

```python
upload = client.files.upload("large-demo.mp4", content_type="video/mp4")

response = client.converse(
    modelId="gemini-2.5-flash",
    messages=[{
        "role": "user",
        "content": [
            {"text": "Summarize this video."},
            {"video": {"format": "mp4", "source": {"uri": upload["uri"]}}},
        ],
    }],
)
```

`client.files.upload(data, content_type=…)` accepts raw bytes or a file path
and returns `{"file_id", "upload_url", "uri", "expires_at"}`. Uploads are
**resumable**: the helper streams the file in 8 MB chunks over a GCS resumable
session and automatically resumes from the last confirmed byte after a network
drop (bounded retries). Uploaded files are transient inference inputs: they
expire after ~24 hours and are scoped to your API key's organization; the
signed upload URL itself is valid ~15 minutes.

To control video token cost, set `mediaResolution` on the request —
`"low"` processes video at ~4x fewer tokens (Gemini models; ignored elsewhere):

```python
response = client.converse(
    modelId="gemini-2.5-flash",
    mediaResolution="low",
    messages=[…],
)
```

---

## Image generation

Generate images from a text prompt with `client.images.generate`. With
`response_format="b64_json"` the image bytes come back inline:

```python
import base64

result = client.images.generate(
    model="gpt-image-1",
    prompt="A watercolor painting of a fox in a snowy forest",
    size="1024x1024",
    response_format="b64_json",
)

with open("fox.png", "wb") as f:
    f.write(base64.b64decode(result["data"][0]["b64_json"]))
```

With `response_format="url"` (the default) the gateway stores the image and
returns a signed download link that expires after ~24 hours:

```python
import urllib.request

result = client.images.generate(
    model="gemini-2.5-flash-image",
    prompt="A watercolor painting of a fox in a snowy forest",
)

urllib.request.urlretrieve(result["data"][0]["url"], "fox.png")
```

Optional parameters: `n` (1–10 images), `size` (`"WxH"`, model-dependent),
`quality` (e.g. `"standard"`, `"hd"`), and `user`. The response dict carries
`created`, `data` (each entry has `url` or `b64_json`, plus `revised_prompt`
when the provider rewrites your prompt) and optional `usage` token counts.
Errors map to the same exception classes as `converse` (401 →
`AccessDeniedException`, 429 → `ThrottlingException`, spend cap →
`ServiceQuotaExceededException`).

### Image models

| Model | Pricing |
|---|---|
| `gemini-2.5-flash-image` | token-based (≈ $0.039 / image) |
| `gpt-image-1` | token-based ($5.00 / 1M text input, $40.00 / 1M image output tokens) |

Per-image models bill a flat rate per generated image; token-based models
report token counts in the response `usage` field. Each request's cost is
recorded against your key like any other API call.

> **Reliability note.** Upstream image backends occasionally stall with a 503
> for a few minutes. There is deliberately **no automatic fallback** to a
> different image model (silently swapping models would change output style
> and quality) — the SDK already retries 503s with backoff up to
> `max_retries`, and callers should retry beyond that rather than switch
> models.

---

## Prompt caching

Place a Bedrock-style `cachePoint` block after the stable prefix you want the
provider to cache (system prompt, large documents). Everything up to the cache
point is cached provider-side on Claude models; OpenAI/Gemini models cache
automatically and ignore the marker. At most 4 cache points per request.

```python
response = client.converse(
    modelId="claude-sonnet-4-5",
    system=[
        {"text": long_system_prompt},
        {"cachePoint": {"type": "default"}},
    ],
    messages=[{
        "role": "user",
        "content": [
            {"text": big_reference_document},
            {"cachePoint": {"type": "default"}},
            {"text": "Summarise the key risks."},
        ],
    }],
)

# Cache activity is reported in usage (Bedrock names):
print(response["usage"]["cacheReadInputTokens"])   # tokens served from cache
print(response["usage"]["cacheWriteInputTokens"])  # tokens written to cache
```

A `cachePoint` with nothing before it in the same message is silently ignored.

---

## OpenAI & Anthropic wire formats (pass-through)

Prefer the OpenAI or Anthropic request shapes over Bedrock Converse? The same
client exposes both gateway-native formats directly — same key, same retry
policy, same exceptions, **no shape translation** (requests sent verbatim,
responses returned raw). Both formats work with every gateway model; the
gateway bridges the wire format for you.

```python
# OpenAI format → POST /v1/chat/completions
response = client.chat.completions.create(
    model="gpt-5.4-mini",
    messages=[{"role": "user", "content": "Hello!"}],
    max_tokens=256,
)
print(response["choices"][0]["message"]["content"])

# Streaming: raw OpenAI chunks
for chunk in client.chat.completions.create(
    model="claude-sonnet-4-6",   # any gateway model
    messages=[{"role": "user", "content": "Write a haiku."}],
    stream=True,
):
    delta = chunk["choices"][0]["delta"] if chunk.get("choices") else {}
    print(delta.get("content", ""), end="", flush=True)

# Anthropic format → POST /v1/messages (Claude Code's wire format)
response = client.messages.create(
    model="claude-sonnet-4-6",
    max_tokens=256,
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response["content"][0]["text"])

# Streaming: raw Anthropic events (message_start … message_stop)
for event in client.messages.create(
    model="claude-sonnet-4-6",
    max_tokens=256,
    messages=[{"role": "user", "content": "Write a haiku."}],
    stream=True,
):
    if event["type"] == "content_block_delta":
        print(event["delta"].get("text", ""), end="", flush=True)

# Free token counting → POST /v1/messages/count_tokens
count = client.messages.count_tokens(
    model="claude-sonnet-4-6",
    messages=[{"role": "user", "content": "How many tokens is this?"}],
)
print(count["input_tokens"])
```

Fields are passed through untouched, so anything the gateway supports works —
no SDK release needed for new request fields. (Using the official `openai` or
`anthropic` packages instead also works: point their `base_url` at the
gateway — see the developer docs.)

---

## Error handling

```python
from audacity import Audacity
from audacity.exceptions import (
    MissingApiKeyError,
    AccessDeniedException,
    ThrottlingException,
    ServiceQuotaExceededException,
    ModelStreamErrorException,
    SdkError,
)

client = Audacity()

try:
    response = client.converse(modelId="gpt-5.4-mini", messages=[…])
except client.exceptions.ThrottlingException as e:
    print(f"Rate limited: {e.message}, retry after {e.retry_after_seconds}s")
except client.exceptions.AccessDeniedException as e:
    print(f"Auth error [{e.status_code}]: {e.message}")
except client.exceptions.ServiceQuotaExceededException as e:
    print(f"Budget exhausted: {e.message}")
except SdkError as e:
    print(f"Network/decode error: {e.message}")
```

Streaming errors surface when you iterate the stream:

```python
try:
    for event in stream_response["stream"]:
        …
except client.exceptions.ModelStreamErrorException as e:
    print(f"Stream broken: {e.message}")
```

---

## Configuration & environment variables

| Constructor parameter | Environment variable | Default |
|---|---|---|
| `api_key` | `AIRESERVE_API_KEY` | — (required) |
| `base_url` | `AIRESERVE_BASE_URL` | `https://api.audacityinvestments.com` |
| `timeout` | — | `120.0` seconds |
| `max_retries` | — | `2` (3 total attempts) |

The legacy `AUDACITY_API_KEY` / `AUDACITY_BASE_URL` names still work as
fallbacks; when both are set, the `AIRESERVE_*` names win.

```python
client = Audacity(
    api_key="aireserve_api_…",
    base_url="https://api.audacityinvestments.com",
    timeout=60.0,
    max_retries=3,
)
```

---

## Retry behaviour

The SDK automatically retries transient failures with jittered exponential
backoff (capped at 20 s):

- Retried: `ThrottlingException` (429), `ModelTimeoutException` (408),
  `ServiceUnavailableException` (502/503/504), `InternalServerException` (500),
  network errors.
- Never retried: `AccessDeniedException`, `ValidationException`,
  `ResourceNotFoundException`, `ServiceQuotaExceededException`
  (including `BUDGET_EXCEEDED` at 429/402), or any 4xx except 408/429.
- Streaming: retries apply only before the first SSE byte; after that,
  connection drops raise `ModelStreamErrorException`.
