Metadata-Version: 2.5
Name: snoopscan
Version: 0.2.2
Summary: Python client for the SnoopScan web scraping API
Project-URL: Homepage, https://snoopscan.com
Project-URL: Documentation, https://snoopscan.com/docs
License-Expression: MIT
License-File: LICENSE
Keywords: agents,crawler,llm,markdown,scraping,web-scraping
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Internet :: WWW/HTTP
Classifier: Topic :: Text Processing :: Markup :: Markdown
Requires-Python: >=3.10
Requires-Dist: httpx>=0.27
Description-Content-Type: text/markdown

# snoopscan

Python client for the SnoopScan scraping snoop.

MIT licensed. The server is AGPL-3.0; a client library must not be, or every
application that imports it inherits the copyleft.

```bash
pip install snoopscan
```

```python
from snoopscan import SnoopScan

snoop = SnoopScan(api_key="sk_...")

page = snoop.scrape("https://example.com/")
print(page.markdown)

for link in snoop.map("https://example.com/", limit=100):
    print(link.url)

job = snoop.crawl_and_wait("https://example.com/", limit=50)
print(job.completed, "of", job.total)
```

## Platforms

A site that publishes its data as JSON is asked, not crawled:

```python
catalogue = snoop.products("https://store.example.com")      # Shopify, WooCommerce, Squarespace, Magento
for p in catalogue["products"]:
    print(p["title"], p["price"], p["currency"], p["available"])

posts = snoop.posts("https://blog.example.com")               # WordPress, Substack, Squarespace, Discourse; else the feed
print(posts["source"], len(posts["posts"]))
```

Every page's `metadata.platform` says what built it, and a `scrape()` of a
Shopify, WooCommerce or Amazon product page carries `product` beside the
markdown. Amazon shows the honest client no price — pass `tier="browser"`.

## Monitors

```python
m = snoop.create_monitor("Pricing", ["https://example.com/pricing"], intervalMinutes=60,
                         webhook="https://hooks.example.com/snoop")
check = snoop.run_monitor(m["id"])          # a check now: {"counts": {...}, "pages": [...]}
snoop.monitor_checks(m["id"])                # recent checks
snoop.delete_monitor(m["id"])
```

Each page in a check is `same`, `changed` (with a git diff), `new` or `error`.
The webhook `monitor.check.completed` fires only when a check has something to
say.

## Company & domain

Two lookups that answer questions a single page can't:

```python
lead = snoop.company("acme.com")                # firmographics + contacts, from the site itself
print(lead["company"]["name"], lead["company"]["headcount"])
for email in lead["contacts"]["emails"]:
    print(email["email"], email["role"], email["onDomain"])

info = snoop.domain("acme.com")                  # registration, DNS, backlinks — not a page fetch
print(info["registration"]["registrar"], info["dns"]["mx"])
print(info["backlinks"]["referringDomains"], "domains link here, per our own crawl graph")
```

`company()` takes `contacts=False` to skip contact discovery and return only
firmographics. `domain()`'s three lookups are each opt-out —
`registration=False`, `dns=False`, `backlinks=False` — since a caller asking
about a domain usually wants all of it, not a form to fill in.

## Base URL

Defaults to `http://localhost:8099`, the engine's own dev port. Point it
elsewhere with the `SNOOP_BASE_URL` environment variable, or per client:

```python
snoop = SnoopScan(api_key="sk_...", base_url="https://api.example.com")
```

## Errors

Every failure raises `SnoopScanError` carrying the API's machine-readable
code, so callers can branch on what actually happened:

```python
from snoopscan import SnoopScanError

try:
    page = snoop.scrape(url)
except SnoopScanError as exc:
    if exc.is_blocked:
        ...          # BLOCKED — the target refused us; retrying as-is will not help
    elif exc.code == "FETCH_FAILED":
        ...          # the target could not be reached; retrying may help
    elif exc.code == "INVALID_REQUEST":
        ...          # our request was wrong; fix it, do not retry
```

## Cost

Every response carries what it cost to produce — the tier that answered, every
tier attempted, proxy bytes, browser milliseconds, and whether it came from
cache. A cache hit reports the accounting of the fetch that filled it, so
`cost.tier` is never null on a page that was really fetched once.

## Development

From a checkout of the engine repo:

```bash
uv pip install -e sdk/python --python .venv/bin/python
```
