Metadata-Version: 2.4
Name: brontide
Version: 0.1.1
Summary: Forecasting the research frontier: which regions of the arXiv embedding cloud are emerging vs saturating.
Author: Bobby Zhang
License: MIT
Project-URL: Homepage, https://github.com/numberingeometry/brontide
Project-URL: Repository, https://github.com/numberingeometry/brontide
Project-URL: Issues, https://github.com/numberingeometry/brontide/issues
Keywords: forecasting,scientometrics,embeddings,time-series,arxiv,research-trends
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Classifier: Topic :: Scientific/Engineering :: Visualization
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy>=1.26
Requires-Dist: pandas>=2.0
Requires-Dist: pyarrow>=14
Requires-Dist: scipy>=1.11
Requires-Dist: matplotlib>=3.8
Requires-Dist: pillow>=10
Provides-Extra: pipeline
Requires-Dist: scikit-learn>=1.3; extra == "pipeline"
Requires-Dist: statsforecast>=1.7; extra == "pipeline"
Requires-Dist: umap-learn>=0.5; extra == "pipeline"
Requires-Dist: torch>=2.0; extra == "pipeline"
Requires-Dist: transformers>=4.40; extra == "pipeline"
Requires-Dist: adapters>=0.2; extra == "pipeline"
Provides-Extra: naming
Requires-Dist: anthropic>=0.40; extra == "naming"
Provides-Extra: dev
Requires-Dist: ruff>=0.6; extra == "dev"
Requires-Dist: pytest>=8; extra == "dev"
Dynamic: license-file

<!-- GENERATED by scripts/make_pypi_readme.py from README.md. DO NOT EDIT. -->
![](https://raw.githubusercontent.com/numberingeometry/brontide/v0.1.1/assets/brontide-logo.svg)

# brontide — forecasting the research frontier

> *brontide* (n.) — the low rumble of thunder from a storm still beyond the horizon.

Which regions of the arXiv embedding cloud are **emerging**, and which are **saturating**? 2.7M papers → SPECTER2 embeddings → 1,024 microregions → 24-month share forecasts, held to a **pre-committed, multi-cutoff out-of-time harness** with a published negative result.

![The arXiv research frontier igniting, 2007–2026](https://raw.githubusercontent.com/numberingeometry/brontide/v0.1.1/figures/map-evolution.gif)

**2,705,992 papers · 2007-01 → 2026-06 · 234 months · 1,024 microregions → 64 Ward super-regions · 768→50 exact PCA · 24-month horizon · 3 masked cutoffs**

`PyTorch` · `transformers` + `adapters` (SPECTER2 + proximity adapter) · `scikit-learn` (MiniBatchKMeans) · `NumPy`/`SciPy` (exact streamed-covariance PCA, Ward linkage) · `UMAP` · `statsforecast` (AutoETS / AutoARIMA / AutoTheta) · `pandas`/`pyarrow` · Python 3.11

---

## Headline result — the criterion was not met

The go/no-go was frozen at repo scaffold on **2026-07-22**, before any data was touched, and scored on **2026-07-23**. It failed, and the failure is reported verbatim:

> **Roster beat SmoothedMomentum on precision@50 in 0/3 cutoffs (need ≥2).
> Verdict: region-level emergence is momentum-dominated at 24m.**

![precision@50 by model and cutoff against the chance rate](https://raw.githubusercontent.com/numberingeometry/brontide/v0.1.1/figures/validation-precision.png)

**But the verdict needs its scope note, and I found the reason myself before anyone read it.** At two of three cutoffs the cross-validation selected `SmoothedMomentum` — the baseline — so the test evaluated `0.240 > 0.240` and `0.300 > 0.300`, false by identity. At the third, an asymmetric minimum-train guard gave the baseline **zero CV folds**. Meanwhile a fixed AutoETS scored **0.613 mean precision@50** against a **0.101** chance rate, versus momentum's 0.433 with 3.5× the dispersion.

So the null is about **model selection, not absent signal**: the harness selected on point accuracy (MASE + Spearman) and was scored on the top 5% of a ranking. Those are not the same loss. Every non-Naive model clears chance at every cutoff.

The frozen sentence stands unedited. The full defect disclosure, the four defensible readings of that sentence (one of which flips the verdict to 2/3), the chance-rate arithmetic, and what the result does and does not license are in **[LIMITATIONS.md](https://github.com/numberingeometry/brontide/blob/v0.1.1/LIMITATIONS.md)**.

## Verify it yourself in two seconds

Every published number regenerates from committed artifacts — no raw corpus, no GPU, no network:

```sh
git clone https://github.com/numberingeometry/brontide && cd brontide
pip install -e .
python scripts/verify_results.py
```

```
re-scored 3 cutoffs from committed artifacts (2017-12, 2020-12, 2022-12)
...
negative criterion: 0/3 (need >=2) -> region-level emergence is momentum-dominated at 24m
OK: all 18 published rows reproduce exactly.
```

The pre-registration ordering is checkable too — thresholds at `0e6d16a` (2026-07-22), results appended at `5c69145` (2026-07-23), with every intervening edit an addition:

```sh
git log --follow --format='%h %ad %s' --date=short -- artifacts/validation.md
git diff --numstat 0e6d16a 7a9edb9 -- artifacts/validation.md   # 18 insertions, 0 deletions
```

## The interactive dashboard

**[`dashboard.html`](https://github.com/numberingeometry/brontide/blob/v0.1.1/dashboard.html)** — one self-contained file (no server, no CDN, no network). Scrub the year to watch the frontier fill in, hover for the region under the cursor, click any region for its share history and forecast, and sort all 64 super-regions by forecast growth.

```sh
python scripts/run_dashboard.py && open dashboard.html   # or just open the committed file
```

## What the map actually looks like

![The arXiv frontier with the largest super-regions labelled](https://raw.githubusercontent.com/numberingeometry/brontide/v0.1.1/figures/frontier-map.png)

Region names are LLM-generated from each super-region's c-TF-IDF terms and sample titles ([`name_regions.py`](https://github.com/numberingeometry/brontide/blob/v0.1.1/src/brontide/name_regions.py), output committed to [`artifacts/region_names.json`](https://github.com/numberingeometry/brontide/blob/v0.1.1/artifacts/region_names.json) so nothing downstream needs an API key). They are **cosmetic**: names never enter the forecast, the emergence statistic, or the scored validation.

The hierarchy is unsupervised all the way down — nothing told it that computation and the physical sciences are different:

![Ward linkage over the 1,024 microregions](https://raw.githubusercontent.com/numberingeometry/brontide/v0.1.1/figures/region-hierarchy.png)

## Forecast against ground truth

At the 2020-12 cutoff, with everything after it masked — the cutoff where CV picked the baseline itself:

![History, forecast band, and realized path for six super-regions](https://raw.githubusercontent.com/numberingeometry/brontide/v0.1.1/figures/forecast-vs-realized.png)

And what the model says at the production origin (2026-06):

![Super-regions ranked by forecast growth](https://raw.githubusercontent.com/numberingeometry/brontide/v0.1.1/figures/emerging-regions.png)

## Pipeline

| # | Stage | Module | Where | Output |
|---|---|---|---|---|
| 1 | Ingest | `ingest.py` | PC (~1 h) | `data/papers.parquet` (id, month, categories, title, abstract) |
| 2 | Embed | `embed.py` | **PC GPU (~14 h measured)** | float16 memmap `N×768`, SPECTER2 + proximity adapter, resumable |
| 3 | Reduce | `reduce.py` | PC | exact PCA 768→50 (streamed centered covariance + `eigh`; deterministic, batch-order invariant); UMAP 50→2 for display only |
| 4 | Regions | `regions.py` | PC (~30 m) | MiniBatchKMeans **K=1024** + Ward tree → 64 super-regions |
| 5 | Series | `series.py` | PC | EB-smoothed region-**share** monthly series + novelty stream |
| 6 | Forecast | `forecast.py` | PC/laptop | statsforecast roster, 24m horizon, rolling-origin CV |
| 7 | Emergence | `emergence.py` | laptop | rank by ĝ; emerging / declining / saturating |
| 8 | Validate | `validate.py` | laptop | 3 cutoffs, precision@50 + Spearman + CSI vs the baseline |
| 9 | Labels | `labels.py`, `name_regions.py` | laptop | c-TF-IDF terms, then LLM names per super-region |
| 10 | Export | `export_map.py`, `export_figs.py`, `export_dashboard.py` | laptop | map binaries, 7 figures, dashboard |

Every numeric knob is frozen in **[`src/brontide/config.py`](https://github.com/numberingeometry/brontide/blob/v0.1.1/src/brontide/config.py)** and mirrored, dated, into **[`artifacts/validation.md`](https://github.com/numberingeometry/brontide/blob/v0.1.1/artifacts/validation.md)** *before* any scoring run. Determinism is load-bearing: the harness refits PCA, k-means, κ, and the novelty threshold at each cutoff on ≤T₀ data only, so an order-dependent solver would make the masked runs irreproducible.

### Three tiers of reproduction

| Tier | Command | Cost | Needs |
|---|---|---|---|
| Verify the headline | `python scripts/verify_results.py` | ~2 s | committed artifacts only |
| Rebuild figures + dashboard | `python scripts/run_figs.py && python scripts/run_dashboard.py` | ~1 min | committed artifacts only |
| Smoke-test the pipeline | `python scripts/smoke_test.py --sample 50000` | minutes | a corpus snapshot |
| Full rebuild | `run_ingest → run_embed → … → run_export` | ~16 h, ~10 GB | GPU + corpus snapshot |

```sh
conda env create -f environment.yml && conda activate brontide
pip install -e .
```

Every stage takes `--sample N` so the whole chain can be smoke-tested before a full run. `data/` is gitignored and PC-only; see [`data/README.md`](https://github.com/numberingeometry/brontide/blob/v0.1.1/data/README.md) for what belongs there and how to rebuild it.

## Data and model provenance

- **Corpus** — arXiv metadata (titles + abstracts), 2007-01 → 2026-06. arXiv metadata is distributed under CC0; the underlying papers remain under their authors' licenses and arXiv's Terms of Use. Check the licence stated on the snapshot you download before redistributing.
- **Embeddings** — [`allenai/specter2_base`](https://huggingface.co/allenai/specter2_base) with the `allenai/specter2` proximity adapter (Singh et al., SciRepEval). Apache-2.0; see the model cards.
- **What is committed here** — region-level *aggregates* (share series, forecasts, emergence ranks), a 149,994-point display sample of 2-D coordinates, and 256 sample titles for labelling. No abstracts and no per-paper records are redistributed.
- Region names are LLM-generated; the exact prompt is in `name_regions.build_prompt`.

## Status

- Pipeline complete on the full corpus; all three cutoffs scored 2026-07-23.
- Map export is real (149,994 sampled points, `synthetic: false`) and committed.
- Figures, dashboard, and the reproduction script land 2026-08-11.
- Open: a test suite and CI. The name is settled (see below).
- **Not built:** the agentic grader and the persistent-homology layer that earlier drafts described. Neither exists in `src/`; both were dropped rather than quietly carried.

Next, pre-registered separately before any scoring: novelty-stream incremental skill, and 80% interval coverage / PIT for the three auto-models. Both signals are already computed at all three cutoffs and have never been scored — a clean pre-registration is still available there.

## Prior art

Large-scale research maps already exist. **Klavans, Boyack & Murdick** (PLOS ONE 2020) and the deployed **CSET/ETO Map of Science** own scale and deployment — 38–43M documents, ~15× the corpus here — and **PreScience v2** (arXiv:2602.20459v2, Ai2 + UChicago) defines the share-change target publicly. This project claims **no new formulation** and no priority; scale is not a claim. What it adds is the pre-committed multi-cutoff out-of-time harness and an honest null.

Full comparison, the PreScience v1/v2 version trap, and the honest scoping of the word "pre-registered" are in **[RELATED-WORK.md](https://github.com/numberingeometry/brontide/blob/v0.1.1/RELATED-WORK.md)**.

> **Renamed 2026-08-11.** This project was developed under the internal codename *lacuna*, which collides with arXiv:2606.26246 (Mila) — an unrelated static research map. The name was chosen from a vetted candidate set; `brontide` is free on PyPI and unclaimed in ML. Commits before the rename refer to the old name.

## Layout

```
src/brontide/    ingest embed reduce regions series forecast emergence validate
               labels name_regions export_map export_figs export_dashboard  + config.py
scripts/       run_<stage>.py thin runners · verify_results.py · smoke_test.py
artifacts/     committed outputs + validation.md (the pre-registration record)
figures/       generated — regenerate with scripts/run_figs.py
decisions/     dated ADRs, immutable (index in decisions/README.md)
blueprint/     claims.yaml → LEDGER.md, plus the check.py gate
data/          gitignored — big intermediates, PC only
```

`LEDGER.md` is generated from [`blueprint/claims.yaml`](https://github.com/numberingeometry/brontide/blob/v0.1.1/blueprint/claims.yaml); `python blueprint/check.py` is the gate.

## License

MIT — see [LICENSE](https://github.com/numberingeometry/brontide/blob/v0.1.1/LICENSE).
