Metadata-Version: 2.5
Name: dclimate-tabular-py
Version: 0.2.1
Summary: Content-addressed tabular/telemetry manifest for climate data on IPFS (dclimate-tabular/1)
Project-URL: Homepage, https://github.com/dClimate/tabular-py
Project-URL: Repository, https://github.com/dClimate/tabular-py
License: MIT
Requires-Python: >=3.12
Requires-Dist: blake3>=0.4.1
Requires-Dist: dag-cbor>=0.3.3
Requires-Dist: httpx>=0.28.1
Requires-Dist: multiformats[full]>=0.3.1.post4
Requires-Dist: py-hamt>=3.6.0
Requires-Dist: pyarrow>=18.0.0
Description-Content-Type: text/markdown

# tabular-py

Python reader for `dclimate-tabular/1` — content-addressed entity/telemetry data on IPFS.

The Python counterpart to [`tabular-js`](https://github.com/dClimate/tabular-js), reading the
same format from the same CIDs. Built on [`py-hamt`](https://github.com/dClimate/py-hamt),
which supplies the HAMT and the content-addressed store the same way
`@dclimate/ipld-index` does for JS.

```
tabular-js  ->  @dclimate/ipld-index      (HAMT, CAS, range reads)
tabular-py  ->  py-hamt                   (same, already existed)
```

## Status

**Reader only.** Publishing, compaction, and rollup stay in `tabular-js` — the ETL that
writes these datasets is JS, and a second writer would be a second thing to keep
byte-identical for no current gain. Everything needed to *read* a dataset published by
`tabular-js` is here.

**`dclimate-tabular/1` only.** Roots written under `/0` are refused with the remedy rather
than misparsed; there is no dual-read shim. `/1` replaced the station model with the
**entity model**: an entity is whatever a dataset is keyed by, and need not be a place, so
a dataset of derivative contracts is now expressible. See the [changelog](CHANGELOG.md).

## Install

```bash
uv pip install dclimate-tabular-py
```

## Usage

```python
import asyncio
from tabular_py import GatewayRangeSource, EntityDataset

async def main():
    source = GatewayRangeSource("https://ipfs-gateway.dclimate.net")
    ds = await EntityDataset.open(source, root_cid)

    # nearest station to a point, then a year of readings
    near = await ds.nearest(34.05, -118.24, max_km=100)
    rows = await near.time_range("2024-01-01", "2024-12-31").elements("PRCP").rows()
    for row in rows[:5]:
        print(row.entity_id, row.ts, row.values)

asyncio.run(main())
```

The chainable selection API mirrors `tabular-js`:

```python
ds.select("USW00023174")           # explicit entity ids
ds.circle(34.05, -118.24, 50)      # within 50 km
ds.rectangle(33.0, -119.0, 35.0, -117.0)
ds.polygon([[(lon, lat), ...]])
await ds.nearest(lat, lon)         # async: reads the geo index
ds.time_range(start, end)
ds.elements("PRCP", "TMAX")
ds.where(gt("TMAX", 300))          # pushed down to fragment statistics
```

Selections are immutable — each call returns a new `EntityDataset`, so a base dataset can
be reused across queries.

Terminal operations:

| Call | Returns |
|---|---|
| `await ds.rows()` | `list[ResultRow]` |
| `await ds.to_records()` | `list[dict]` — `entity_id`, `time`, `values` |
| `await ds.to_records("TMAX")` | `list[dict]` — `entity_id`, `time`, `value` |
| `await ds.to_arrow()` | `pyarrow.Table` |
| `await ds.plan()` | `QueryPlan` — what *would* be fetched, without fetching |
| `await ds.list_entities()` | `list[EntityInfo]` |
| `ds.columns()` | `list[EntityColumn]` — the dataset's vocabulary, with units |
| `await ds.columns_for(id)` | `list[EntityColumn]` — what one entity reports |
| `await ds.gaps_for(id)` | `list[DataGap]` — windows known to be *unknown* |

## How reads stay small

A query never scans the dataset. Three things prune before any Parquet byte is fetched:

1. **The entity index** (a HAMT keyed by entity id) resolves named entities directly.
2. **The geo projection** answers region queries by reading one or two shard blocks
   instead of walking every entity.
3. **Fragment statistics in the manifest** — per-column min/max and null counts — let a
   predicate skip whole fragments unread.

What survives is fetched with HTTP range requests against the exact column-chunk byte
ranges the manifest records, so a query for one column of one year moves kilobytes.

### One deviation from `tabular-js`, and why

`tabular-js` synthesizes Parquet `FileMetaData` client-side from the manifest and reads a
fragment with **zero** footer fetches. PyArrow exposes no public `FileMetaData`
constructor, so that trick does not transfer.

Instead this reader fetches the footer by its manifest-recorded
`footer_offset`/`footer_length` in a single ranged GET, verifies it against the
manifest's `footer_digest`, and hands the parsed metadata to PyArrow. Cost is one extra
range request per fragment — and it buys a corruption check `tabular-js` does not
perform. Footers are cached per fragment CID, so a repeated query pays it once.

## Development

```bash
uv sync
uv run pytest                     # unit tests, no network
uv run pytest -m network          # conformance against the live gateway
uv run ruff check . && uv run mypy tabular_py
```

### How this is tested against `tabular-js`

`tests/test_golden_cids.py` is the cross-language contract. It builds the same small
structures `tabular-js` builds in its own `test/wire.test.ts`, and asserts they encode to
the **same CIDs**. Because a CID is a hash of the encoded bytes, agreement means the two
libraries produce byte-identical blocks for identical inputs — checkable offline, against
no published dataset at all.

If a vector mismatches, fix the encoder rather than the vector. When an encoding changes
deliberately, change it in `tabular-js` first and copy the new CID across, so the two are
never quietly updated to match each other.

The rest of the suite runs against a synthetic `/1` dataset assembled in
`tests/helpers/build.py` — both index variants, an entity with no position, a pre-epoch
timestamp, declared gaps, and real ZSTD Parquet fragments read through byte ranges. It is
a test helper, not a writer: it composes the public `*_to_wire` encoders and does no
fragment planning, bucketing, compaction, or rollup.

`tests/test_conformance.py` runs against the four published `/1` datasets — GHCNd
(132,437 entities, 1.15 B rows), SCAN, SNOTEL, and NDBC. It covers the one thing the
offline suite cannot: that both libraries agree end to end on data neither of them wrote.
Each root is fetched, decoded, and re-encoded back to its own CID, which is the strongest
available statement that nothing was dropped, reordered, or silently defaulted.

When an ETL republishes, the roots move — update the CIDs and re-derive the asserted
figures rather than loosening the assertions.
