Metadata-Version: 2.4
Name: thinkingsdk
Version: 0.1.5
Summary: AI crash debugging for Python in production
Author-email: srikar appalaraju <srikar2097@gmail.com>
License: MIT
Project-URL: Homepage, https://github.com/srikarappal/thinkingsdk-python
Project-URL: Website, https://thinkingsdk.ai
Project-URL: Issues, https://github.com/srikarappal/thinkingsdk-python/issues
Keywords: crash reporting,error tracking,exceptions,debugging,observability,monitoring,AI,production
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Topic :: Software Development :: Debuggers
Classifier: Topic :: System :: Monitoring
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.7
Classifier: Programming Language :: Python :: 3.8
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Operating System :: OS Independent
Requires-Python: >=3.7
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: requests>=2.25.0
Requires-Dist: pyyaml>=5.4
Requires-Dist: psutil>=5.9
Requires-Dist: contextvars>=2.4; python_version < "3.7"
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: pytest-asyncio>=0.21; extra == "dev"
Requires-Dist: pytest-cov>=4.0; extra == "dev"
Requires-Dist: black>=22.0; extra == "dev"
Requires-Dist: flake8>=5.0; extra == "dev"
Requires-Dist: mypy>=1.0; extra == "dev"
Requires-Dist: psutil>=5.9; extra == "dev"
Provides-Extra: performance
Requires-Dist: psutil>=5.9; extra == "performance"
Provides-Extra: keyring
Requires-Dist: keyring>=24.0; extra == "keyring"
Dynamic: license-file

# ThinkingSDK

AI crash debugging for Python in production. One line catches every uncaught exception in your live app, ships it to ThinkingSDK's analysis service, and gives you back the root cause and a concrete fix, not just another stack trace.

[![PyPI version](https://badge.fury.io/py/thinkingsdk.svg)](https://pypi.org/project/thinkingsdk/)
[![Python](https://img.shields.io/badge/python-3.7%2B-blue)](https://www.python.org/downloads/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)

```bash
pip install thinkingsdk
```

```python
import thinkingsdk as thinking

thinking.start(
    api_key="sk_live_...",
    server_url="https://api.thinkingsdk.ai",  # hosted service; the SDK otherwise defaults to localhost
)
```

That is it. Deploy. When your app throws an unhandled exception, in a request handler, a worker thread, or a background job, ThinkingSDK captures it with full context, analyzes it, and shows you the diagnosis in your dashboard at [thinkingsdk.ai](https://thinkingsdk.ai). Get your project key there.

## What you get for every crash

- **Root cause in plain language.** Why it happened, grounded in the actual stack and the local variables at each frame, not just where it threw.
- **A concrete fix.** The specific code change to make, not a generic "handle the exception."
- **Full context.** Stack trace, locals per frame, and the execution path into the failure.
- **Smart grouping.** Repeated crashes are deduplicated, so one real bug is one issue, not a thousand alerts.

## How it works

ThinkingSDK installs exception hooks at startup:

- `sys.excepthook` for the main thread and `threading.excepthook` for worker threads, so no uncaught exception escapes unseen.
- Captured events are sent asynchronously in batches on a background sender, off your request path, so reporting a crash never blocks or slows the failing request.
- Only runtime event data leaves your process. Your source code is never uploaded.

Deeper call tracing and performance capture are available via config, off by default to keep overhead near zero.

## Performance

ThinkingSDK is built to stay off your application's hot path. The capture path is cheap and bounded; the network and the AI analysis happen elsewhere. What makes that true, in the client itself:

- **Background daemon thread**, sends off your request path; capture just enqueues in microseconds.
- **Bounded ring buffer** (`deque(maxlen=…)`), drops oldest when full instead of back-pressuring, so a flood can't stall your app or grow memory.
- **Batching** (50 events, or ~2s), flat network cost, not one request per event.
- **Circuit breaker**, pauses sending after repeated backend failures, so a down service can't become retry pressure on you.
- **Bounded retries + exponential backoff + hard request timeout**, all confined to the background thread.
- **Priority sampling**, exceptions/errors always captured, routine events sampled by rate.
- **Call-stack dedup**, a hot error loop collapses to one analyzed issue, not thousands of sends.
- **Idle-until-throw hooks**, `excepthook` adds no per-call/per-line cost on the happy path; deeper tracing is opt-in.

## Framework and library integrations

Awareness for the stack you already run:

- **Web:** FastAPI, Flask, Django
- **Data:** SQLAlchemy, psycopg2, PyMongo, Redis
- **Standard library:** logging

## Production example (FastAPI)

```python
import thinkingsdk as thinking
from fastapi import FastAPI

thinking.start(api_key="sk_live_...", server_url="https://api.thinkingsdk.ai")

app = FastAPI()

@app.get("/orders/{order_id}")
def get_order(order_id: str):
    # If this raises in production, ThinkingSDK captures it with the request
    # context and returns an AI root cause plus a fix in your dashboard.
    return load_order(order_id)
```

When `load_order` blows up on a malformed id, you do not get a bare `KeyError` buried in your logs. You get:

```
KeyError in get_order -> load_order at order_store.py:42

Root cause: load_order indexes self._orders[order_id] directly, but order_id
arrives from the URL unvalidated, so any unknown id raises KeyError instead of
returning a 404.

Suggested fix:
    order = self._orders.get(order_id)
    if order is None:
        raise HTTPException(status_code=404, detail="order not found")
    return order
```

## Configuration

`thinking.start()` accepts:

- `api_key`: your project key from [thinkingsdk.ai](https://thinkingsdk.ai)
- `server_url`: set this to `https://api.thinkingsdk.ai` for the hosted service (or via `THINKINGSDK_SERVER_URL`). It defaults to `http://localhost:8000` for local dev, so crashes won't reach the hosted dashboard unless you set it.
- `config`: a dict of tuning options

```python
thinking.start(
    api_key="sk_live_...",
    server_url="https://api.thinkingsdk.ai",
    config={
        "capture_exceptions": True,
        "capture_performance": False,
        "sample_rate": 1.0,   # e.g. 0.1 to sample 10% on a high traffic service
    },
)
```

### Caught exceptions (off by default)

The SDK reports crashes: unhandled exceptions, caught through `sys.excepthook`,
`threading.excepthook` and the framework error handlers. Exceptions your code catches and
handles are not reported, because a `raise` is not a failure. `any()` short-circuiting closes a
generator and raises `GeneratorExit`, SQLAlchemy's type cache uses `try/except KeyError` as its
miss path, every iterator ends with `StopIteration`. One real FastAPI request raised 981
exceptions with nothing wrong.

If you do want caught-exception telemetry, turn it on and measure it on your own workload
first. It is expensive by nature: every `raise` in the process pays a frame walk.

```python
thinking.start(api_key="sk_live_...", config={"capture_caught_exceptions": True})
```

```bash
export THINKINGSDK_CAPTURE_CAUGHT_EXCEPTIONS=true
```

### Environment variables

```bash
export THINKINGSDK_API_KEY=sk_live_...
export THINKINGSDK_SERVER_URL=https://api.thinkingsdk.ai   # required for the hosted service (defaults to localhost)
```

## Self hosting

The analysis service (the AI engine and dashboard) runs as a separate component. `server_url` selects which one: `https://api.thinkingsdk.ai` for the hosted service, or your own deployment. The SDK's built-in default is `http://localhost:8000` (local dev), so set `server_url` for anything else. See [thinkingsdk.ai](https://thinkingsdk.ai) for details.

## License

MIT

## Caught exceptions are opt-in (0.1.5)

`sys.monitoring.events.RAISE` fires on every `raise`, not only on failures, so subscribing to it
instrumented ordinary control flow: generator short-circuits, cache misses, iterator exhaustion.
Measured at 2703x on caught exceptions, and 5.2s of server time on a single FastAPI request that
raised 981 times with zero errors.

The tracer is now off unless `capture_caught_exceptions` is set. Crash reporting is unchanged: the
excepthooks are the crash path and are always installed, so unhandled exceptions arrive with their
traceback, locals, breadcrumbs and repository context as before.

`LoggingIntegration` also no longer forces `DEBUG`. It uses its own `INFO` default, so a debug-chatty
application does not turn every record into a breadcrumb.

## Resource safety (0.1.4)

The SDK suppresses capture throughout its own background upload context, including HTTP
and TLS dependencies. Application exceptions continue to be captured. Diagnostic values
are bounded before formatting; custom object `repr` and `str` methods are not invoked.
Exception messages, nested values, traceback frames and exception chains are truncated
when they exceed capture limits.

Default buffering limits are 8 MiB of serialized queued reports, 256 KiB per report,
2 MiB of retained duplicate samples and 64 KiB per duplicate sample. A batch has a
1 MiB payload budget. These are payload limits, not a promise that total process memory
stays below their sum. Oversized reports and overflow entries are dropped, with queue
and deduplicator statistics exposing drops. Inputs are detached so caller mutation
cannot enlarge retained payloads.

The first occurrence of a crash is queued immediately. Repeats with the same exception
type and stack location are aggregated over a fixed 15 minute window into one bounded
sample and a count. Timestamps and changing locals do not create retained variations.
The existing `deduplicated_pattern` envelope carries the repeat count in `data.frequency`;
it excludes the first event already sent. Backend consumers must add that frequency,
not count the envelope as one occurrence. This release does not change backend counting
or provide server idempotency for retried HTTP requests.

Failed uploads use the existing finite HTTP retries and circuit breaker plus an
interruptible exponential delay between batches, capped at 60 seconds. Delivery remains
best effort: exhausted retries, process exit and buffer pressure can lose reports.
These changes do not add persistent disk storage.

```python
thinkingsdk.start(
    api_key="YOUR_API_KEY",
    config={
        "queue": {
            "max_bytes": 8 * 1024 * 1024,
            "max_event_bytes": 256 * 1024,
        },
        "deduplication": {
            "window_size_ms": 900000,
            "flush_interval_ms": 900000,
            "max_bytes": 2 * 1024 * 1024,
            "max_sample_bytes": 64 * 1024,
        },
        "sender": {"max_batch_bytes": 1024 * 1024},
    },
)
```
