Metadata-Version: 2.5
Name: sorula
Version: 0.1.0
Summary: Sorula: make any text-to-speech sound human. A streaming layer around ElevenLabs, Cartesia, OpenAI, Deepgram, Azure, Google, Polly, PlayHT or your own audio.
Project-URL: Homepage, https://sorula.com
Project-URL: Documentation, https://app.sorula.com/docs
Project-URL: Source, https://bitbucket.org/sorula/repo
Project-URL: Node package, https://www.npmjs.com/package/sorula
Author: Sorula
License: Proprietary
License-File: LICENSE
Keywords: cartesia,elevenlabs,prosody,speech,streaming,text-to-speech,tts,voice
Classifier: Development Status :: 4 - Beta
Classifier: Framework :: AsyncIO
Classifier: Intended Audience :: Developers
Classifier: License :: Other/Proprietary License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Multimedia :: Sound/Audio :: Speech
Requires-Python: >=3.9
Requires-Dist: websockets>=12
Provides-Extra: audio
Requires-Dist: numpy>=1.24; extra == 'audio'
Provides-Extra: dev
Requires-Dist: numpy>=1.24; extra == 'dev'
Requires-Dist: pytest-asyncio>=0.23; extra == 'dev'
Requires-Dist: pytest>=8; extra == 'dev'
Description-Content-Type: text/markdown

<p align="center">
  <a href="https://sorula.com"><img src="https://app.sorula.com/logo.svg" alt="Sorula" width="160"></a>
</p>

<h1 align="center">Sorula for Python</h1>

<p align="center">
  <a href="https://pypi.org/project/sorula/"><img src="https://img.shields.io/pypi/v/sorula?color=6C5CE7&label=pypi" alt="PyPI"></a>
  <a href="https://pypi.org/project/sorula/"><img src="https://img.shields.io/pypi/pyversions/sorula?color=6C5CE7" alt="Python versions"></a>
  <a href="https://sorula.com"><img src="https://img.shields.io/badge/latency-from_150_ms-6C5CE7" alt="Latency from 150 ms"></a>
  <a href="https://sorula.com"><img src="https://img.shields.io/badge/website-sorula.com-000000" alt="Website"></a>
  <a href="https://app.sorula.com/docs"><img src="https://img.shields.io/badge/docs-app.sorula.com%2Fdocs-000000" alt="Docs"></a>
  <a href="https://www.npmjs.com/package/sorula"><img src="https://img.shields.io/badge/also_on-npm-CB3837?logo=npm&logoColor=white" alt="npm"></a>
</p>

<p align="center">
  <a href="#elevenlabs">ElevenLabs</a> ·
  <a href="#cartesia">Cartesia</a> ·
  <a href="#openai">OpenAI</a> ·
  <a href="#deepgram-aura">Deepgram</a> ·
  <a href="#azure-ai-speech">Azure</a> ·
  <a href="#google-cloud-text-to-speech">Google</a> ·
  <a href="#amazon-polly">Polly</a> ·
  <a href="#playht">PlayHT</a> ·
  <a href="#any-other-tts-or-your-own-audio">Any TTS</a> ·
  <a href="#phone-calls-twilio-and-other-8-khz-mu-law">Phone</a>
</p>

---

**Make any text-to-speech sound human.** Sorula sits between your TTS and
your listener: the audio goes in, the same voice comes out sounding like a
real person in a real room. Streaming, it runs from **150 ms behind** the
TTS, fast enough for live phone calls. Or hand it a finished clip.

Works with ElevenLabs, Cartesia, OpenAI, Deepgram, Azure, Google, Amazon
Polly, PlayHT, and anything else that produces audio.

## Get started

**1. Get an API key.** [Create a free account](https://app.sorula.com/signup)
(60 free minutes, no card), then make a key under
[API keys](https://app.sorula.com/keys). Keys look like `sorula_sk_...` and
are shown once.

```bash
export SORULA_API_KEY=sorula_sk_...
```

**2. Install.**

```bash
pip install sorula
```

**3. Humanize.**

```python
import sorula

client = sorula.Sorula()   # reads SORULA_API_KEY; or pass api_key="sorula_sk_..."

# a finished clip: WAV in, WAV out
human = client.humanize(open("tts.wav", "rb").read())

# a live stream: audio chunks in, human audio chunks out
for chunk in client.stream(tts_chunks, sample_rate=24000, format="s16"):
    play(chunk)
```

Your key starts with `sorula_sk_`. Anything else (an OpenAI `sk-...`, an
ElevenLabs key) is refused with a clear message, so keys can't get mixed up.

---

## Pick your TTS

Every example below is complete: paste it, set the environment variables, run
it. Each one asks the provider for raw PCM (the setting is in
`sorula.providers.<name>`), streams it through Sorula and writes a WAV.
Replace the file write with your player or your call.

### ElevenLabs

```bash
pip install elevenlabs sorula
export ELEVENLABS_API_KEY=... SORULA_API_KEY=sorula_sk_...
```

```python
import os
from elevenlabs.client import ElevenLabs
import sorula
from sorula.providers import elevenlabs as el

eleven = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
client = sorula.Sorula()

audio = eleven.text_to_speech.stream(
    voice_id="JBFqnCBsd6RMkjVDRZzb",
    model_id="eleven_flash_v2_5",
    text="Hi, it's Sarah from the clinic. I'm just calling to confirm your appointment for Thursday at ten.",
    output_format=el.OUTPUT_FORMAT,          # "pcm_24000"
)

out = bytearray()
for chunk in client.stream(audio, sample_rate=el.SAMPLE_RATE, format=el.FORMAT):
    out += chunk                             # play it / send it to the call

open("human.wav", "wb").write(sorula.audio.write_wav(sorula.audio.Pcm(bytes(out), el.SAMPLE_RATE)))
```

**Faster, with ElevenLabs word timings (~150 ms behind):** use
`stream_with_timestamps` and let the adapter pass the timings along.

```python
client = sorula.Sorula(lookahead_s=0.08)
items = eleven.text_to_speech.stream_with_timestamps(
    voice_id="JBFqnCBsd6RMkjVDRZzb", model_id="eleven_flash_v2_5", text=TEXT, output_format=el.OUTPUT_FORMAT)
for chunk in client.stream(el.iter_timestamps(items), sample_rate=el.SAMPLE_RATE):
    out += chunk
```

Using the ElevenLabs `stream-input` WebSocket instead? Add
`&output_format=pcm_24000&sync_alignment=true` to the URL and wrap the parsed
messages with `el.aiter_websocket(messages)` (see `examples/elevenlabs_words.py`).

### Cartesia

```bash
pip install cartesia sorula
export CARTESIA_API_KEY=... SORULA_API_KEY=sorula_sk_...
```

```python
import os
from cartesia import Cartesia
import sorula
from sorula.providers import cartesia as cs

cartesia = Cartesia(api_key=os.environ["CARTESIA_API_KEY"])
client = sorula.Sorula(lookahead_s=0.08)     # Cartesia gives word timings: ~150 ms behind

with cartesia.tts.websocket_connect() as ws:
    ctx = ws.context(
        model_id="sonic-3.5",
        voice={"mode": "id", "id": "6ccbfb76-1fc6-48f7-b71d-91ac6298247b"},
        output_format=cs.OUTPUT_FORMAT,      # raw pcm_s16le @ 24 kHz
        language="en",
    )
    ctx.push("Hi, it's Sarah from the clinic. I'm just calling to confirm your appointment for Thursday at ten.",
             add_timestamps=True, continue_=False)

    out = bytearray()
    for chunk in client.stream(cs.iter_responses(ctx.receive()), sample_rate=cs.SAMPLE_RATE, format=cs.FORMAT):
        out += chunk

open("human.wav", "wb").write(sorula.audio.write_wav(sorula.audio.Pcm(bytes(out), cs.SAMPLE_RATE)))
```

The SSE endpoint works the same way: `client.stream(cs.iter_responses(cartesia.tts.sse(...)), ...)`.

### OpenAI

```bash
pip install openai sorula
export OPENAI_API_KEY=... SORULA_API_KEY=sorula_sk_...
```

```python
from openai import OpenAI
import sorula
from sorula.providers import openai as oa

ai = OpenAI()
client = sorula.Sorula()

with ai.audio.speech.with_streaming_response.create(
    model="gpt-4o-mini-tts",
    voice="coral",
    input="Hi, it's Sarah from the clinic. I'm just calling to confirm your appointment for Thursday at ten.",
    response_format=oa.RESPONSE_FORMAT,      # "pcm": 16-bit mono 24 kHz
) as response:
    out = bytearray()
    for chunk in client.stream(oa.iter_bytes(response), sample_rate=oa.SAMPLE_RATE, format=oa.FORMAT):
        out += chunk

open("human.wav", "wb").write(sorula.audio.write_wav(sorula.audio.Pcm(bytes(out), oa.SAMPLE_RATE)))
```

### Deepgram (Aura)

```bash
pip install deepgram-sdk sorula
export DEEPGRAM_API_KEY=... SORULA_API_KEY=sorula_sk_...
```

```python
import os, threading
from deepgram import DeepgramClient
from deepgram.core.events import EventType
from deepgram.speak.v1.types import SpeakV1Text
import sorula
from sorula.providers import deepgram as dg

deepgram = DeepgramClient(api_key=os.environ["DEEPGRAM_API_KEY"])
client = sorula.Sorula()
source = sorula.CallbackSource()             # Deepgram pushes audio via events

def synthesize():
    with deepgram.speak.v1.connect(model="aura-2-thalia-en", **dg.CONNECT_KWARGS) as conn:   # linear16 @ 24 kHz
        conn.on(EventType.MESSAGE, dg.on_message(source))   # audio -> source, Flushed -> close
        conn.on(EventType.ERROR, source.fail)
        conn.start_listening()
        conn.send_text(SpeakV1Text(text="Hi, it's Sarah from the clinic. I'm just calling to confirm your appointment for Thursday at ten."))
        conn.send_flush()
        source.wait_closed(timeout=60)
        conn.send_close()

threading.Thread(target=synthesize, daemon=True).start()
out = bytearray()
for chunk in client.stream(source, sample_rate=dg.SAMPLE_RATE, format=dg.FORMAT):
    out += chunk

open("human.wav", "wb").write(sorula.audio.write_wav(sorula.audio.Pcm(bytes(out), dg.SAMPLE_RATE)))
```

### Azure AI Speech

```bash
pip install azure-cognitiveservices-speech sorula
export AZURE_SPEECH_KEY=... AZURE_SPEECH_REGION=... SORULA_API_KEY=sorula_sk_...
```

```python
import os
import azure.cognitiveservices.speech as speechsdk
import sorula
from sorula.providers import azure as az

client = sorula.Sorula(lookahead_s=0.08)     # Azure reports word boundaries: ~150 ms behind

cfg = speechsdk.SpeechConfig(subscription=os.environ["AZURE_SPEECH_KEY"], region=os.environ["AZURE_SPEECH_REGION"])
cfg.speech_synthesis_voice_name = "en-US-AvaMultilingualNeural"
cfg.set_speech_synthesis_output_format(getattr(speechsdk.SpeechSynthesisOutputFormat, az.OUTPUT_FORMAT_NAME))  # Raw24Khz16BitMonoPcm
cfg.set_property(speechsdk.PropertyId.SpeechServiceResponse_RequestWordBoundary, "true")
cfg.set_property(speechsdk.PropertyId.SpeechServiceResponse_RequestPunctuationBoundary, "true")

synthesizer = speechsdk.SpeechSynthesizer(speech_config=cfg, audio_config=None)
source = sorula.CallbackSource()
az.attach(synthesizer, source)               # audio + word boundaries -> source
synthesizer.speak_text_async("Hi, it's Sarah from the clinic. I'm just calling to confirm your appointment for Thursday at ten.")

out = bytearray()
for chunk in client.stream(source, sample_rate=az.SAMPLE_RATE, format=az.FORMAT):
    out += chunk

open("human.wav", "wb").write(sorula.audio.write_wav(sorula.audio.Pcm(bytes(out), az.SAMPLE_RATE)))
```

### Google Cloud Text-to-Speech

```bash
pip install google-cloud-texttospeech sorula
export GOOGLE_APPLICATION_CREDENTIALS=... SORULA_API_KEY=sorula_sk_...
```

```python
from google.cloud import texttospeech
import sorula
from sorula.providers import google as gg

tts = texttospeech.TextToSpeechClient()
client = sorula.Sorula()
voice = texttospeech.VoiceSelectionParams(language_code="en-US", name="en-US-Chirp3-HD-Aoede")
TEXT = "Hi, it's Sarah from the clinic. I'm just calling to confirm your appointment for Thursday at ten."

# finished clip: LINEAR16 is a WAV, humanize() takes it as is
response = tts.synthesize_speech(
    input=texttospeech.SynthesisInput(text=TEXT), voice=voice,
    audio_config=texttospeech.AudioConfig(audio_encoding=texttospeech.AudioEncoding.LINEAR16, sample_rate_hertz=gg.SAMPLE_RATE))
open("human.wav", "wb").write(client.humanize(response.audio_content))

# streaming
config = texttospeech.StreamingSynthesizeRequest(streaming_config=texttospeech.StreamingSynthesizeConfig(
    voice=voice, streaming_audio_config=texttospeech.StreamingAudioConfig(
        audio_encoding=texttospeech.AudioEncoding.PCM, sample_rate_hertz=gg.SAMPLE_RATE)))
text = texttospeech.StreamingSynthesizeRequest(input=texttospeech.StreamingSynthesisInput(text=TEXT))
responses = tts.streaming_synthesize(iter([config, text]))

out = bytearray()
for chunk in client.stream(gg.iter_streaming(responses), sample_rate=gg.SAMPLE_RATE, format=gg.FORMAT):
    out += chunk
```

### Amazon Polly

```bash
pip install boto3 sorula
export AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... AWS_DEFAULT_REGION=... SORULA_API_KEY=sorula_sk_...
```

```python
import boto3
import sorula
from sorula.providers import polly as pl

polly = boto3.client("polly")
client = sorula.Sorula()
VOICE = dict(VoiceId="Joanna", Engine="neural",
             Text="Hi, it's Sarah from the clinic. I'm just calling to confirm your appointment for Thursday at ten.")

words = pl.words_from_speech_marks(polly.synthesize_speech(**VOICE, **pl.SPEECH_MARK_KWARGS))   # word timings, no audio
audio = polly.synthesize_speech(**VOICE, OutputFormat=pl.OUTPUT_FORMAT, SampleRate=pl.SAMPLE_RATE_STR)  # raw PCM @ 16 kHz

def source():
    yield words                              # timings first: Polly knows them up front
    yield from pl.iter_audio(audio)

out = bytearray()
for chunk in client.stream(source(), sample_rate=pl.SAMPLE_RATE, format=pl.FORMAT):
    out += chunk

open("human.wav", "wb").write(sorula.audio.write_wav(sorula.audio.Pcm(bytes(out), pl.SAMPLE_RATE)))
```

### PlayHT

```bash
pip install pyht sorula
export PLAY_HT_USER_ID=... PLAY_HT_API_KEY=... SORULA_API_KEY=sorula_sk_...
```

```python
import os
from pyht import Client
from pyht.client import Format, TTSOptions
import sorula
from sorula.providers import playht as ph

play = Client(user_id=os.environ["PLAY_HT_USER_ID"], api_key=os.environ["PLAY_HT_API_KEY"])
client = sorula.Sorula()

options = TTSOptions(
    voice="s3://voice-cloning-zero-shot/775ae416-49bb-4fb6-bd45-740f205d20a1/jennifersaad/manifest.json",
    format=getattr(Format, ph.FORMAT_NAME),  # FORMAT_RAW: headerless 16-bit PCM
    sample_rate=ph.SAMPLE_RATE,
)
chunks = play.tts("Hi, it's Sarah from the clinic. I'm just calling to confirm your appointment for Thursday at ten.",
                  options, voice_engine="Play3.0-mini-http")

out = bytearray()
for chunk in client.stream(ph.iter_chunks(chunks), sample_rate=ph.SAMPLE_RATE, format=ph.FORMAT):
    out += chunk

open("human.wav", "wb").write(sorula.audio.write_wav(sorula.audio.Pcm(bytes(out), ph.SAMPLE_RATE)))
```

### Any other TTS, or your own audio

Sorula needs nothing but mono PCM chunks (16-bit or float32) and the sample
rate. Ask your engine for raw PCM (most call it `pcm`, `linear16` or `raw`)
and use whichever shape matches your code:

```python
# 1. an iterable of bytes (a generator, a list, a file read in blocks)
for chunk in client.stream(my_chunks, sample_rate=24000, format="s16"): ...

# 2. an async iterable
async for chunk in client.astream(my_async_chunks, sample_rate=24000): ...

# 3. callbacks, when the engine calls you
source = sorula.CallbackSource()
engine.on_audio(source.push)                 # bytes
engine.on_done(source.close)
for chunk in client.stream(source, sample_rate=24000): ...

# 4. a finished file or buffer
client.humanize_file("tts.wav", "human.wav")
pcm_out = client.humanize(pcm_bytes, sample_rate=24000, format="s16")
```

If your engine reports word timings, pass them too and Sorula runs ~150 ms
behind: yield `sorula.Word("Hello,", 0.10, 0.42)` (keep the punctuation)
between the chunks, or `source.push_word(text, start_s, end_s)`.

### Phone calls (Twilio and other 8 kHz mu-law)

Ask the TTS for 24 kHz PCM, humanize, then convert for the line:

```python
from sorula.audio import Pcm, to_mulaw_8000
for chunk in client.stream(audio, sample_rate=24000, format="s16"):
    payload = base64.b64encode(to_mulaw_8000(Pcm(chunk, 24000))).decode()
    twilio_ws.send(json.dumps({"event": "media", "streamSid": sid, "media": {"payload": payload}}))
```

`from_mulaw_8000` goes the other way if your TTS only speaks `ulaw_8000`.

---

## Reference

### `sorula.Sorula(...)`

```python
sorula.Sorula(
    api_key="sorula_sk_...",   # or SORULA_API_KEY
    room="kitchen",            # "kitchen" | "living_room" | "office" | "bedroom" | "none"
    intensity=1.5,             # how strong the effect is; 1.0 = natural
    strength=1.0,              # 0..1 blend between the original delivery and Sorula's
    lookahead_s=0.16,          # streaming: how much audio Sorula hears ahead; 0.08 with word timings
)
```

Any of these can also be given per call: `client.stream(..., room="none")`.

| method | in | out |
| --- | --- | --- |
| `humanize(audio, sample_rate=None, format="s16", mode="clip", words=None)` | WAV bytes, or raw PCM + rate | WAV bytes, or raw PCM |
| `humanize_file(src, dst)` | WAV path | WAV path |
| `stream(source, sample_rate, format="s16")` | iterable of bytes / `Word`s | iterator of bytes |
| `astream(source, ...)` | (async) iterable | async iterator |
| `session(sample_rate, format, mode)` | manual: `send`, `words`, `end`, `async for` | |

`mode` is `"clip"` for finished audio (default in `humanize`) or `"live"`
(default when streaming). `format` is `"s16"` or `"f32"`; rates from 8 to
48 kHz, 24 kHz recommended.

### Manual session

```python
async with client.session(sample_rate=24000, format="s16") as s:
    await s.send(pcm)                       # any size, as often as you like
    await s.words([sorula.Word("Hello,", 0.1, 0.4)])
    await s.end()
    async for chunk in s: ...
    s.result                                # {"in_s": ..., "out_s": ..., "words": ...}
```

### `sorula.audio`

Dependency-free helpers: `read_wav`, `write_wav`, `Pcm`, `s16_to_f32`,
`f32_to_s16`, `resample_s16`, `mulaw_encode`, `mulaw_decode`,
`to_mulaw_8000`, `from_mulaw_8000`.

### Requirements

Python 3.9+, `websockets`. `numpy` is optional (`pip install sorula[audio]`)
and speeds up the conversions. `examples/` has each snippet above as a
runnable file.
