Metadata-Version: 2.4
Name: breeze-blue
Version: 0.9.0
Summary: Python SDK for the Breeze Blue Developer API.
Author-email: Breeze Blue <support@breeze.blue>
License-Expression: MIT
Project-URL: Homepage, https://breezeblue.ai
Project-URL: Documentation, https://docs.breezeblue.ai
Keywords: audio,breeze,breeze-blue,sdk,text-to-speech,tts,voice
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Multimedia :: Sound/Audio :: Speech
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: httpx<1.0.0,>=0.28.0
Requires-Dist: websockets<17.0,>=16.0
Dynamic: license-file

# Breeze Blue Python SDK

Python SDK for the Breeze Blue Developer API. Covers text-to-speech, voice
management, voice preview generation, history audio, models, and account
usage.

## Install

```bash
uv add breeze-blue
```

or:

```bash
pip install breeze-blue
```

## API Key

Create an API key in the Breeze Blue Developer Console, then export it:

```bash
export BREEZE_API_KEY=brz_...
```

The SDK sends the key with the `xi-api-key` header.

## Quickstart

```python
from breeze_blue import BreezeBlue, save

client = BreezeBlue()

audio = client.text_to_speech.convert(
    voice_id="voc_...",
    text="Hello from Breeze Blue.",
    output_format="mp3",
)

save(audio, "hello.mp3")
print(audio.content_type)
print(audio.history_item_id)
```

`BreezeBlue()` reads `BREEZE_API_KEY` by default and sends requests to
`https://api.breeze.blue`. To point at another environment, pass `base_url=...`
or set `BREEZE_BASE_URL`.

```python
client = BreezeBlue(
    api_key="brz_...",
    base_url="https://api.breeze.blue",
    timeout=120.0,
)
```

All resource methods accept `timeout=<seconds>` to override the client default
for a single call.

## Text to Speech

```python
audio = client.text_to_speech.convert(
    voice_id="voc_...",
    text="Render this line.",
)

streamed_audio = client.text_to_speech.stream(
    voice_id="voc_...",
    text="Stream this line.",
)

enhanced = client.text_to_speech.enhance(
    instruction="Calm, warm, bedtime narration.",
    language_code="en",
)
```

Use async text-to-speech for long text, reference-heavy voices, or batch
production where the caller should not hold an HTTP connection open:

```python
job = client.text_to_speech.create_job(
    voice_id="voc_...",
    text="Render this longer script.",
    output_format="mp3",
)

status = client.generation_jobs.get(job["generation_job_id"])
if status["status"] == "ready":
    audio = client.generation_jobs.download_audio(job["generation_job_id"])
    audio.save("async.mp3")
```

If the job is still active, `download_audio(...)` raises
`GenerationNotReadyError`; read `exc.retry_after` before retrying.

Use realtime text-to-speech when one WebSocket connection should handle multiple
conversation turns. Realtime audio is fixed to raw `pcm_s16le`, 24000 Hz, mono,
16-bit frames. Start consuming before appending text: audio can arrive after
`flush()` and may continue after `end_turn()`, so keep consuming until
`turn.done`:

```python
import asyncio

from breeze_blue import BreezeBlue, stream

client = BreezeBlue()


async def consume_audio(connection) -> tuple[bytes, dict]:
    audio = bytearray()
    audio_format = None
    async for event in connection:
        if event["type"] == "session.ready":
            audio_format = event["audio_format"]
        elif event["type"] == "audio":
            audio.extend(event["audio"])
        elif event["type"] == "error":
            raise RuntimeError(event.get("message"))
        elif event["type"] == "turn.done":
            if audio_format is None:
                raise RuntimeError("Missing realtime audio format")
            return bytes(audio), audio_format
    raise RuntimeError("Realtime session closed before turn.done")


async def speak() -> None:
    async with client.text_to_speech.connect_realtime(
        voice_id="voc_...",
        model_id="breeze-tts-2",
    ) as connection:
        consumer = asyncio.create_task(consume_audio(connection))
        await connection.start_turn("turn_1")
        await connection.append_text("Hello from Breeze.")
        await connection.flush()
        await connection.end_turn()
        audio, audio_format = await consumer

    stream(audio, audio_format=audio_format)


asyncio.run(speak())
```

The connection carries a single ordered message stream. Iterate the connection
directly (as above), or use exactly one of `connection.events()` (non-audio
events) and `connection.audio()` (raw PCM chunks) — running two consumers at
the same time splits the stream between them. `connection.audio()` raises
`RealtimeError` when the server reports an `error` event or the WebSocket
closes abnormally, so audio-only consumers never mistake a failure for the end
of a turn.

Between turns, update the synthesis instructions without replacing the
WebSocket. `update_instructions(...)` only writes the request; the matching
`session.updated` acknowledgement arrives on the same ordered event stream.
Wait for that acknowledgement before starting the next turn:

```python
await connection.update_instructions("Speak faster and with more energy.")
async for event in connection:
    if event["type"] == "session.updated":
        assert event["instructions"] == "Speak faster and with more energy."
        break

await connection.start_turn("turn_2")
```

Instructions must be a non-empty string of at most 1,000 characters. Only one
update may await acknowledgement, and updates are accepted only while no turn
is active. An update attempted during a turn is rejected without interrupting
that turn. Keep one stream consumer: if a long-lived consumer owns the
iterator, have it signal your turn producer when `session.updated` arrives
instead of starting a second iterator.

Not every server error ends the session. A `RealtimeError` with `recoverable`
set to `True` means the same connection remains usable. Most validation and
turn-command errors cancel the active turn, so catch the error and start a new
turn on the same connection:

```python
from breeze_blue import RealtimeError

try:
    async for chunk in connection.audio():
        handle(chunk)
except RealtimeError as exc:
    if not exc.recoverable:
        raise
    # Most recoverable turn-command errors cancelled the current turn.
    await connection.start_turn("turn_2")
    await connection.append_text("Let me try that again.")
    await connection.flush()
    await connection.end_turn()
    async for chunk in connection.audio():
        handle(chunk)
```

An active-turn `session.update` rejection is the exception: it is recoverable,
but the turn keeps running. If `connection.audio()` raises that error, call
`connection.audio()` again immediately on the same connection and keep
collecting the current turn; do not start a replacement turn.

With the low-level `connect_realtime(...)` API, the session ends after
`inactivity_timeout_seconds` (30 seconds by default) without client messages.
During idle gaps between turns, send `await connection.ping()` as a keepalive;
the server replies with a `pong` event. `connect_realtime(timeout=...)` bounds
the WebSocket handshake and defaults to 30 seconds.

For a long-running conversation, use the opt-in managed connection. It keeps
the logical connection available across bounded physical WebSockets:

```python
async with client.text_to_speech.connect_managed(
    voice_id="voc_...",
    model_id="breeze-tts-2",
) as connection:
    consumer = asyncio.create_task(consume_audio(connection))
    await connection.start_turn("turn_1")
    await connection.append_text("This conversation can continue across sessions.")
    await connection.flush()
    await connection.end_turn()
    audio, audio_format = await consumer
```

`connect_managed(...)` waits for `session.ready`, derives a safe heartbeat
interval (including while an active turn is waiting for first audio), and
requires an inbound server message within five seconds after each managed
heartbeat. A missing acknowledgement triggers an idle reconnect; during an
active turn it raises `TURN_INTERRUPTED` without replaying turn commands. The
manager marks rotation pending at the earlier of the server's safe
`max_session_seconds` deadline and a 600-second physical-session age threshold,
or when the server sends `session.expiring`. It switches only after
`turn.done` or `turn.cancelled`, so an active turn can keep the old epoch past
that threshold; the server's 1,800-second hard lifetime still applies. If a new
turn starts on the old socket while its replacement
is opening, that turn keeps the old socket and rotation resumes after its
terminal event. An idle reconnect uses fresh credentials and bounded
equal-jitter backoff (three attempts by default); the physical close is hidden
and the next visible epoch starts with another `session.ready`. A replacement
epoch sends its first heartbeat immediately, then resumes the derived cadence
after an inbound acknowledgement. If `start_turn(...)` arrives while that idle
replacement is already in progress, it waits for the same bounded transition
and sends `turn.start` once on the new epoch. Other turn commands are not
buffered. Session setup
retries transport failures and HTTP `408`, `425`, `429`, and `5xx` responses;
other `4xx` policy or configuration responses remain terminal.
After a planned rotation, the replacement is available immediately while the
old physical socket remains open for at most five seconds to forward a late
`usage.committed` event. That event is best-effort; use history and account
usage APIs as the durable source of truth.

The manager never buffers turn content or replays a command. The bounded
`start_turn(...)` wait happens before its first WebSocket write. If a physical
connection ends during an active turn, iteration raises `RealtimeError` with
`code == "TURN_INTERRUPTED"`; rebuild the turn from application conversation
state. Recoverable errors remain visible on the same socket,
known transient upstream failures trigger an idle replacement without ending
the logical audio iterator, and non-retryable policy or billing failures raise
`RealtimeError`. A positive `timeout` is a shared setup deadline for obtaining
fresh credentials, completing the WebSocket handshake, and receiving
`session.ready` on every epoch. Set `heartbeat_interval_seconds=0` to disable
the automatic heartbeat. `heartbeat_timeout_seconds` must be positive and
defaults to 5 seconds. Set `max_physical_session_seconds=0` to disable the SDK
age threshold (the server lifetime still applies), or pass a positive number
to override the default 600 seconds. Tune `rotation_margin_seconds`,
`max_reconnect_attempts`, and `reconnect_base_delay_seconds` when needed.
Treat the manager as an active realtime call, not a presence channel. Exit the
async context when the call ends, the user leaves, or the application enters a
long-lived background state; otherwise its heartbeat intentionally keeps a
server WebSocket slot occupied.

Every managed epoch creates a session through the SDK by default. If another
service brokers short-lived credentials, pass a synchronous or async
`session_factory` that returns a fresh `client_secret` or `websocket_url` on
every call. Session options such as `model_id` must then be applied by that
factory when it creates the session. Prefer an async factory for broker I/O.
A dynamically updated instruction is the exception: after its
`session.updated` acknowledgement, the manager automatically reapplies that
confirmed value before exposing the next factory-created physical epoch as
ready. SDK-created replacement sessions include the confirmed value in their
session request.
A synchronous factory runs through `asyncio.to_thread`; a manager setup timeout
or close cannot forcibly stop a worker thread that has already started.
Synchronous callbacks must therefore configure their own bounded network I/O
timeouts.

To reduce time to first audio, create the session ahead of time (for example
while waiting for user input) and connect with its `client_secret` when the
first turn starts, keeping session setup off the critical path:

```python
session = client.text_to_speech.create_realtime_session(
    "voc_...",
    model_id="breeze-tts-2",
)

async with client.text_to_speech.connect_realtime(
    voice_id="voc_...",
    client_secret=session["client_secret"],
    websocket_url=session["websocket_url"],
    direct_websocket_url=session.get("direct_websocket_url"),
) as connection:
    ...
```

When the session response includes `direct_websocket_url`, the SDK prefers its
query-free origin and authenticates with the WebSocket subprotocol. It falls
back to `websocket_url` for services that have not enabled the direct transport.

`model_id`, `language_code`, `voice_settings`, timeout, and logging parameters
are fixed when the session is created. The initial `instructions` value is
also supplied at creation, but it can later be replaced between turns with
`update_instructions(...)`. Initial values passed to `connect_realtime`
together with `client_secret` are ignored (the SDK emits a `UserWarning`).

When using the low-level `connect_realtime(...)` API, if a WebSocket is
interrupted by a network change, service deployment, or upstream realtime
worker restart, handle an `error` event with
`meta.reconnect == True`, a `session.closed` event with `reconnect == True`, or
a `RealtimeError` with `reconnect == True` by creating a new connection and
starting a new turn from your own conversation state. Active turns are not
resumed in place. `GENERATION_CAPACITY_EXCEEDED` is turn-scoped: wait for the
following `turn.cancelled`, back off using `meta.retry_after_seconds`, and start
a new turn on the same managed logical connection without replaying text.

The API uses the default text-to-speech model when `model_id` is omitted. If
you need to select a model explicitly, call `client.models.list()` and pass one
of the returned `model_id` values.

Audio responses are returned as `AudioResponse`:

```python
audio.content          # bytes
audio.content_type     # e.g. "audio/mpeg"
audio.history_item_id  # history item id when returned by the API

audio.save("speech.mp3")
```

Playback helpers are available for local scripts:

```python
from breeze_blue import play, stream

play(audio)    # ffplay, with macOS afplay fallback
stream(audio)  # mpv; accepts AudioResponse, bytes, or an iterable of byte chunks
```

For raw PCM audio, pass the `audio_format` dict returned by realtime sessions
(`session["audio_format"]` or the `session.ready` event) so `stream(...)` plays
at the right sample rate and channel count; without it, raw PCM playback falls
back to 24000 Hz mono `s16le`.

Streaming text-to-speech defaults to `pcm` to reduce time to first audio. Pass
`output_format="wav"` or `output_format="mp3"` when you need that wire format
explicitly.

`audio.stream()` is a buffered helper for `AudioResponse`.
Use the module-level `stream(audio_chunks)` helper when you already have chunked
audio data.

## Voices

Search every voice available to the account and inspect a single voice. Public
catalog results combine semantic similarity with exact, prefix, and substring
name matching; saved personal voices remain searchable by name, description,
or voice ID.

```python
voices = client.voices.search(search="calm documentary narrator")
first_voice_id = voices["voices"][0]["voice_id"]

voice = client.voices.get(first_voice_id)
settings = client.voices.get_settings(first_voice_id)

random_voice = client.voices.random()
print(random_voice["voice_id"], random_voice["name"])
```

Breeze voice creation is always two steps: produce a **preview**, let the
user accept it, then **save the preview** as a real voice. Two ways to
produce a preview:

```python
# Option A — clone preview from an audio sample
clone_preview = client.voices.create_clone_preview(
    name="Demo voice",
    file="sample.wav",
    text="This is a short preview script.",
)
generated_voice_id = clone_preview["generated_voice_id"]

# Option B — design preview from a text description (no audio)
design = client.voices.create_design_preview(
    voice_description="Warm documentary narrator with clear articulation.",
)
generated_voice_id = design["previews"][0]["generated_voice_id"]
```

`files` is also accepted with exactly one item.

Stream the preview so the user can audition it, then save the one they pick:

```python
audio = client.voices.stream_preview(generated_voice_id)
audio.save("preview.mp3")

saved = client.voices.save_preview(
    generated_voice_id=generated_voice_id,
    voice_name="Documentary narrator",
    language_code="en",
)
```

Edit, tune settings, or delete a saved voice:

```python
client.voices.edit(saved["voice_id"], name="Renamed narrator", language_code="en")
client.voices.edit_settings(saved["voice_id"], guidance_scale=1.2)
client.voices.delete(saved["voice_id"])
```

## History, Models, and Account

```python
models = client.models.list()
balance = client.account.balance()
usage = client.account.usage(days=7)
key_usage = client.account.usage(api_key_id="key_01hprod", client_type="sdk")

history = client.history.list(page_size=10)
item = client.history.get(history["history"][0]["history_item_id"])
audio = client.history.download_audio(item["history_item_id"])
```

Common response shapes are exported as `TypedDict` types from
`breeze_blue.types` and from the package root.

## Errors

All API errors inherit from `BreezeBlueError`. HTTP status codes and Breeze
error codes map to typed exceptions:

- `BadRequestError`
- `AuthenticationError`
- `ForbiddenError`
- `NotFoundError`
- `ConflictError`
- `ValidationError`
- `RateLimitError`
- `BillingInsufficientCreditsError`
- `UpstreamError`
- `ServiceUnavailableError`

```python
from breeze_blue import BreezeBlue, RateLimitError

try:
    BreezeBlue().models.list()
except RateLimitError as exc:
    print(exc.retry_after)
```

Each `ApiError` exposes `status_code`, `code`, `detail`, `meta`, and
`retry_after`.

Realtime TTS WebSocket failures raise `RealtimeError` (also a
`BreezeBlueError`), which exposes `code`, `meta`, `close_code`, and
`reconnect`.
