Metadata-Version: 2.4
Name: hypothesis-generation
Version: 0.5.0
Summary: Deterministic premise invention and hypothesis generation without LLMs.
Author: Chris Dimurro
License-Expression: Apache-2.0
Project-URL: Homepage, https://github.com/cdimurro/hypothesis-generation
Project-URL: Repository, https://github.com/cdimurro/hypothesis-generation
Project-URL: Issues, https://github.com/cdimurro/hypothesis-generation/issues
Project-URL: Documentation, https://github.com/cdimurro/hypothesis-generation#readme
Keywords: science,hypothesis,research,deterministic,agents
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
License-File: NOTICE
Provides-Extra: test
Requires-Dist: pytest<9,>=8; extra == "test"
Requires-Dist: ruff<0.16,>=0.15; extra == "test"
Requires-Dist: jsonschema<5,>=4; extra == "test"
Requires-Dist: build<2,>=1; extra == "test"
Requires-Dist: setuptools>=77; extra == "test"
Dynamic: license-file

# Hypothesis Generation

**Generate novel, testable hypotheses without an LLM.**

Hypothesis Generation is a deterministic, programmable engine for scientists,
engineers, mathematicians, and research software teams. It constructs candidate
hypotheses from explicit observations, concepts, relations, constraints,
mechanisms, analogies, and domain knowledge.

The engine runs locally. It has no runtime model dependency, makes no network
requests, and does not sample from a pretrained distribution. The same request,
engine version, extension set, domain pack, and corpus snapshot produces the
same result.

## Why it is different

Language models generate text from learned statistical patterns. This project
uses typed transformations, finite formal reasoning, constraint checks, bounded
search, and quality-diversity selection to construct hypotheses that were not
supplied as templates.

That makes each result reproducible and inspectable:

- domain facts come only from declared inputs, traceable corpora, formal rules,
  or reviewed domain packs;
- invalid candidates can be rejected by explicit type, logic, graph,
  dimensional, contradiction, and mechanism checks;
- requested portfolios are selected for structural coverage instead of padded
  with paraphrases;
- novelty diagnostics are always relative to an identified corpus; and
- derivations, scores, provenance, tests, and archives are available when a
  caller requests them.

The engine proposes candidates for investigation. Experiments, prior-art
review, safety review, and scientific judgment remain external responsibilities.

## Install

Python 3.10 or later is required. The core has no runtime dependencies.

```powershell
python -m pip install -e .
```

The repository version can also be installed directly:

```powershell
python -m pip install "hypothesis-generation @ git+https://github.com/cdimurro/hypothesis-generation.git"
```

## Quick start

The common API requires a phenomenon and a count:

```python
from hypothesis_generation import generate_hypotheses

hypotheses = generate_hypotheses(
    "urban heat islands remain hotter overnight after heat waves",
    count=8,
)

for hypothesis in hypotheses:
    print(hypothesis)
```

The default output is a list of English hypotheses. Backend artifacts remain
hidden unless explicitly requested.

The equivalent command-line request is:

```json
{
  "phenomenon": "urban heat islands remain hotter overnight after heat waves",
  "count": 8
}
```

```powershell
python -m hypothesis_generation --input request.json --pretty
```

## Add scientific structure

More explicit inputs produce more specific hypotheses:

```python
from hypothesis_generation import ConceptSpec, DiscoveryRequest, RelationSpec, discover

request = DiscoveryRequest(
    phenomenon="nighttime cooling slows after a heat wave",
    count=4,
    concepts=(
        ConceptSpec("impervious surface fraction", "DRIVER"),
        ConceptSpec("stored subsurface heat", "MEDIATOR"),
        ConceptSpec("nighttime cooling rate", "OUTCOME"),
    ),
    relations=(
        RelationSpec(
            "impervious surface fraction",
            "increases",
            "stored subsurface heat",
        ),
    ),
)

result = discover(request)
print(result.as_dict())
```

Requests can also declare known facts, observations, contexts, constraints,
prior hypotheses, analogy sources, causal structures, quantities, domain packs,
and a frozen prior-art corpus. Undeclared synonyms and domain facts are never
guessed.

See [the discovery API](docs/DISCOVERY.md) for request fields, response
projection, search budgets, persistence, and exhaustion behavior.

## How generation works

1. **Normalize the request.** Convert declared knowledge into typed concepts,
   relations, constraints, mechanisms, quantities, and provenance records.
2. **Invent premises.** Apply typed transformations and bounded multi-step
   programs to propose mechanisms, missing bridges, alternatives, boundary
   conditions, analogies, interventions, and counterexamples.
3. **Generate candidates.** Explore fourteen complementary reasoning modes.
4. **Test and constrain.** Reject candidates that fail applicable formal or
   declared scientific checks.
5. **Check prior art.** Compare candidates with a frozen local corpus using
   exact, lexical, BM25, relation, and structural diagnostics.
6. **Select a portfolio.** Optimize quality and structural coverage under a
   declared ranking profile.
7. **Render the result.** Return English hypotheses by default and only the
   requested artifact fields when details are needed.

The fourteen public modes cover causal, abductive, temporal, systems, regime,
heterogeneity, measurement, invariant, contrastive, intervention, null-model,
relational, analogical, and compositional reasoning. They are search strategies
over one integrated premise-invention system, not sentence templates.

## Large hypothesis campaigns

Persistent campaigns can generate exact portfolios of up to 10,000 unique
hypotheses when the admitted search pool is large enough. Results are returned
through deterministic, tamper-evident pages. If the requested total cannot be
produced without duplicates or cosmetic rewrites, the campaign reports search
exhaustion before returning a partial result.

Portfolio profiles include:

- `BALANCED`
- `DIVERSITY_FIRST`
- `NOVELTY_FIRST`
- `PLAUSIBILITY_FIRST`
- `TESTABILITY_FIRST`
- bounded caller-declared custom weights

Profiles change transparent ranking objectives; they do not load a learned
ranking model.

## Prior-art comparison

The optional local prior-art index supports exact overlap, phrase containment,
token similarity, Okapi BM25, declared aliases, relation overlap, metadata
filters, date-bounded corpus views, and SQLite/FTS5 storage.

These diagnostics answer whether a candidate overlaps the supplied corpus.
They do not claim exhaustive or universal novelty.

## Program it for an application

There are two extension paths:

- **Domain packs** add reviewed, data-only concepts, relations, rules, sources,
  aliases, constraints, and generative vocabulary.
- **Extension bundles** add deterministic application code through
  content-addressed manifests that bind source files, dependencies, and runtime
  policy.

Starter domain-pack scaffolds are available for engineered-system reliability,
materials/process engineering, and systems biology:

```powershell
hypothesis-generate --scaffold-domain-pack materials-process-engineering `
  --output-directory packs/materials-process
```

Scaffolds contain no facts or approvals and remain `AWAITING_REVIEW` until the
completed digest receives the required reviews. See [domain packs](docs/DOMAIN_PACKS.md)
and [extensions](docs/EXTENSIONS.md).

Optional scientific-computing and formal-solver adapters can be supplied by a
caller. They are never discovered, installed, or invoked implicitly.

## Evaluation

The repository includes structural, historical, cross-domain, external, and
negative-control benchmarks. Selected pinned results are:

| Evaluation | Result | Scope |
|---|---:|---|
| Release qualification | 9/9 controls | Runtime boundary, campaigns, ranking, novelty, extensions, packs, evaluation, and legacy behavior |
| HypoSpace | 441/441 admissible targets | 88 Boolean, causal-graph, and voxel cases |
| HypoBench synthetic | 1,450/1,500 held-out rows | Nine classification cases |
| HypoBench real-world | 2,003/3,100 held-out; 1,431/2,593 OOD | Seven text-classification tasks |
| DiscoveryBench expanded | 8/24 strong; median sealed-test R-squared 0.727 | Systematic synthetic task sample |
| SRSD-Feynman portfolio | 37/50 strong; median test R-squared 0.876 | Recursive symbolic search with sealed test rows |
| External Premise Challenge | 11/12 generated and admitted; 5/12 public top-10 | Dated historical diagnostic across six fields |

These measurements apply only to the pinned files, grammars, budgets, and
evaluation rules. They are not estimates of general scientific truth or expert
preference. External metrics that require an LLM judge are reported as not run.

See [the benchmark guide](benchmarks/README.md), [evaluation kit](docs/EVALUATION_KIT.md),
and [scientific limitations](docs/SCIENTIFIC_LIMITATIONS.md).

## Release status

Version 0.5 is a public beta. The deterministic core, packaging, compatibility,
and benchmark contracts are qualified; the complete suite passes 538 tests.
Version 1.0 is reserved for stable contracts informed by independent users,
reviewed reference domain packs, blind expert evaluation, and prospective
evidence in multiple fields.

The legacy v0.3 typed mathematical-model API remains supported. See
[release readiness](docs/RELEASE_READINESS.md), [release qualification](docs/RELEASE_QUALIFICATION.md),
and the [release plan](docs/RELEASE_PLAN.md).

## Documentation

Start with the [documentation index](docs/README.md). Important references are:

- [Architecture and runtime boundary](docs/DESIGN.md)
- [Discovery API](docs/DISCOVERY.md)
- [Domain packs](docs/DOMAIN_PACKS.md)
- [Application extensions](docs/EXTENSIONS.md)
- [Agent protocol](docs/AGENT_PROTOCOL.md)
- [Scientific capabilities and limitations](docs/SCIENTIFIC_LIMITATIONS.md)
- [Release qualification](docs/RELEASE_QUALIFICATION.md)

## Contributing

Read [CONTRIBUTING.md](CONTRIBUTING.md) before adding an operator, family,
domain pack, adapter, or benchmark. Security guidance is in
[SECURITY.md](SECURITY.md).

Hypothesis Generation is released under the [Apache License 2.0](LICENSE).
