Metadata-Version: 2.4
Name: quickthink
Version: 0.2.2
Summary: Compressed planning scaffold for local LLMs — latency-aware routing and structured output reliability
Author-email: Rolando Bosch <roli@hermes-labs.ai>
License-Expression: Apache-2.0
Project-URL: Homepage, https://hermes-labs.ai
Project-URL: Repository, https://github.com/hermes-labs-ai/quickthink
Project-URL: Documentation, https://github.com/hermes-labs-ai/quickthink#readme
Project-URL: Changelog, https://github.com/hermes-labs-ai/quickthink/blob/main/CHANGELOG.md
Project-URL: Issues, https://github.com/hermes-labs-ai/quickthink/issues
Keywords: llm,local-llm,ollama,inference,scaffold,routing,llm-routing,plan-then-answer,latency-aware-routing,agent,structured-output
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: httpx>=0.27.0
Requires-Dist: typer>=0.12.0
Provides-Extra: dev
Requires-Dist: pytest>=8.2.0; extra == "dev"
Requires-Dist: ruff>=0.4.0; extra == "dev"
Dynamic: license-file

# quickthink

[![CI](https://github.com/hermes-labs-ai/quickthink/actions/workflows/ci.yml/badge.svg)](https://github.com/hermes-labs-ai/quickthink/actions/workflows/ci.yml)
[![License: Apache-2.0](https://img.shields.io/badge/License-Apache_2.0-blue.svg)](LICENSE)
[![Python](https://img.shields.io/badge/python-3.9%2B-blue.svg)](pyproject.toml)

**quickthink is a local-first CLI and Python library that wraps Ollama-backed LLM calls with a compressed plan-then-answer scaffold and latency-aware routing.** It adds a short, validated planning step before the answer for prompts that look multi-step, and routes simple prompts straight through to the model.

It currently ships as a lightweight scaffolding layer for local LLMs with three modes:
- `lite` (default): one-pass inline plan prefix + answer in a single generation
- `two_pass`: separate plan call then answer call
- `direct`: no planning pass, raw prompt to model

The plan can be logged as metadata while hidden from normal UI output.

Part of the [Hermes Labs reliability stack](https://github.com/hermes-labs-ai). quickthink shapes the inference call; sibling tools cover other layers — for example, [LintLang](https://github.com/hermes-labs-ai/lintlang) statically lints agent-config files, which is complementary to (not a substitute for) quickthink's runtime planning scaffold.

## Where it fits

quickthink is useful when you need:
- **local LLM routing** for local-first inference pipelines
- **small model optimization** for constrained hardware and low-latency workflows
- **latency-aware inference** via routing, bypass, and planning-budget controls
- **structured output reliability** through strict planning grammar and eval gates
- **Ollama middleware** for practical local deployment
- **agent runtime compatibility** for CLI and automation-driven execution contexts

## What this is / what this is not

What this is:
- A local middleware layer for Ollama-backed LLM calls.
- A small CLI for planned-answer generation, routing diagnostics, and local benchmarking.
- A canonical eval harness for reproducible project-level quality checks.

What this is not:
- Not a hosted API service.
- Not a model training framework.
- Not a replacement for full agent orchestration platforms.

## Why

Small/local models are fast but often underperform on multi-step tasks.
`quickthink` adds a strict planning pass (6-16 keyword tokens by default) to improve response quality without full verbose reasoning traces.

## Features

- Ollama-first integration
- Model profiles: `qwen2.5:1.5b`, `mistral:7b`, `gemma3:27b`
- Three execution modes: `lite` (default), `two_pass`, `direct`
- Preset routing profiles: `fast`, `balanced`, `strict`
- Lane policy: `default` or `strict_safe` (routes strict-format tasks to direct path)
- Hidden plan by default, optional plan display/logging
- Bypass mode for short prompts (latency control)
- Adaptive routing (`skip`, `12-token`, `max-token` planning lanes)
- Strict plan grammar: `g:<...>;c:<...>;s:<...>;r:<...>`
- Local eval UI server (`quickthink ui`) at `http://127.0.0.1:7860`
- Canonical eval harness: run → judge → validate → report

## 5-minute quickstart

Prerequisite: install and start [Ollama](https://ollama.com/) locally.

```bash
# 1) Install the released package
python -m pip install "quickthink==0.2.2"

# 2) Pull one supported model
ollama pull qwen2.5:1.5b

# 3) Run your first command
quickthink ask "Give me a 3-step plan to learn SQL basics" --model qwen2.5:1.5b
```

If this command works, your local setup is ready.

For development, clone the repository and install the editable development extras:

```bash
git clone https://github.com/hermes-labs-ai/quickthink.git
cd quickthink
python -m pip install -e '.[dev]'
```

## Documentation Map

- Docs index: `docs/README.md`
- First-time setup: `docs/GETTING_STARTED.md`
- Common failures and fixes: `docs/TROUBLESHOOTING.md`
- Known limitations: `docs/KNOWN_LIMITATIONS.md`
- Quick demo script: `docs/demo/QUICK_DEMO.md`
- OSS readiness scorecard: `docs/release/OSS_READINESS_SCORECARD_2026-02-25.md`
- OSS standards alignment (with external references): `docs/release/OSS_STANDARDS_ALIGNMENT_2026.md`
- Agent operating notes: `AGENTS.md`

## Repository Layout

```text
src/quickthink/         Runtime package (CLI, engine, prompts, routing, UI server)
scripts/eval_harness/   Canonical evaluation pipeline (run/judge/validate/report)
scripts/evals/          Legacy smoke/demo helpers (non-canonical)
scripts/demo/           One-command local demo runner
docs/evals/             Prompt sets, rubrics, harness specs, deployment gate notes
docs/release/           Release process and repository audit notes
tests/                  Unit tests for runtime and harness safety checks
```

See full architecture + publishability audit:
`docs/release/REPO_STRUCTURE_AND_PUBLISHABILITY_AUDIT_2026-02-20.md`.

## Canonical vs Legacy Scripts

Canonical project workflows:
- `scripts/eval_harness/*`: maintained evaluation pipeline for run/judge/validate/report.
- `scripts/demo/quickstart.sh`: canonical end-to-end local smoke/demo flow.

Legacy helpers (kept for compatibility and ad-hoc smoke checks):
- `scripts/evals/*`: non-canonical helpers; do not treat as release gate source of truth.

When in doubt, use `scripts/eval_harness/*` and `scripts/demo/quickstart.sh`.

## Usage

List supported profiles:

```bash
quickthink list-models
```

List preset routing profiles:

```bash
quickthink list-presets
```

Show officially supported compatibility models:

```bash
quickthink compatibility
```

Ask with compressed planning:

```bash
quickthink ask "How would a cow round up a border collie?" --model qwen2.5:1.5b --preset balanced
```

Show plan in terminal:

```bash
quickthink ask "How would a cow round up a border collie?" --model mistral:7b --show-plan
```

Switch to two-pass mode:

```bash
quickthink ask "How would a cow round up a border collie?" --mode two_pass --show-route --show-plan
```

Show routing diagnostics:

```bash
quickthink ask "Design a robust parser with tradeoffs and a JSON output schema" --show-route --show-plan
```

Skip planning entirely (`direct` mode):

```bash
quickthink ask "What is the capital of France?" --mode direct
```

Inspect routing and the exact prompt(s) without calling Ollama (`--dry-run` works offline):

```bash
quickthink ask "Design a retry strategy for a flaky payments API: compare exponential backoff versus a circuit breaker, list the tradeoffs, and return a JSON schema for the config" --mode two_pass --dry-run
```

Optional continuity hint (tiny, off by default):

```bash
quickthink ask "Continue the previous structure" --continuity-hint "ctx:prior_goal,format_json"
```

Strict-format-safe lane policy (routes strict format tasks to direct path first):

```bash
quickthink ask "json only: {\"ok\":true,\"why\":\"short\"}" --lane-policy strict_safe --show-route
```

Benchmark with strict-safe lane policy:

```bash
quickthink bench "Answer with YES or NO only: Is 2+2=4?" --lane-policy strict_safe --runs 3
```

Log plan + metrics as JSONL metadata:

```bash
quickthink ask "Design a tiny retry strategy" --log-file ./logs/quickthink.jsonl
```

Benchmark all three modes (lite, two_pass, direct):

```bash
quickthink bench "Design a robust parser for CSV with malformed quotes" --model qwen2.5:1.5b --runs 3
```

## One-Command Quickstart Demo

Run full local demo setup and artifact generation:

```bash
bash scripts/demo/quickstart.sh
```

It does:
- Python env + package install
- `ollama pull` for supported models
- Sample A/B/C eval run
- Result validation
- Markdown/HTML report generation
- Compatibility snapshot update

For a one-minute terminal walkthrough command set, see `docs/demo/QUICK_DEMO.md`.

Optional environment flags:
- `QUICKTHINK_PRESET=fast|balanced|strict`
- `QUICKTHINK_LIMIT=<n>` (number of prompts from canonical set)
- `QUICKTHINK_RUNS=<n>`
- `QUICKTHINK_RUN_JUDGE=1` (switch judge backend from `rule` to `ollama`)
- `QUICKTHINK_JUDGE_MODEL=<model>`

## Troubleshooting

For common setup/runtime failures and fixes, see `docs/TROUBLESHOOTING.md`.

## Reports

Canonical report flow:

```bash
python3 scripts/eval_harness/run_suite.py \
  --prompt-set docs/evals/prompt_set.jsonl \
  --out docs/evals/results/run-<timestamp>.jsonl \
  --manifest-out docs/evals/results/manifest-<timestamp>.json \
  --runs 3

python3 scripts/eval_harness/judge_suite.py \
  --prompt-set docs/evals/prompt_set.jsonl \
  --results docs/evals/results/run-<timestamp>.jsonl \
  --out docs/evals/results/judged-<timestamp>.jsonl \
  --backend rule

python3 scripts/eval_harness/validate_judged_results.py \
  --path docs/evals/results/judged-<timestamp>.jsonl

python3 scripts/eval_harness/report_suite.py \
  --runs docs/evals/results/run-<timestamp>.jsonl \
  --judged docs/evals/results/judged-<timestamp>.jsonl \
  --out-json docs/evals/results/report-<timestamp>.json \
  --out-md docs/evals/results/report-<timestamp>.md \
  --out-html docs/evals/results/report-<timestamp>.html
```

Legacy helpers in `scripts/evals/*` remain available for smoke/demo use only.

## Compatibility Matrix

- Supported models are fixed to:
  - `qwen2.5:1.5b`
  - `mistral:7b`
  - `gemma3:27b`
- Experimental evaluations may include additional models (for example `llama3.2:latest`) in
  deployment-gate or variant-gate workflows. Treat those as research lanes unless promoted
  into `SUPPORTED_MODELS` in runtime config.
- Regenerate matrix + snapshot with:

```bash
python3 scripts/evals/compat_matrix_snapshot.py
```

Launch local web UI (for eval/scaffolding testing):

```bash
quickthink ui
```

Then open `http://127.0.0.1:7860` if it does not open automatically.

UI lane control:
- `Lane policy` dropdown supports `default` and `strict_safe` for single-prompt runs and 3-mode comparisons.

UI eval safety gates:
- Preflight is required before any eval run (`validate_prompt_set.py` must return `status=OK`).
- Run-file ingestion is blocked unless `validate_results.py` returns `status=OK`.
- UI displays validator output and dataset SHA256 for reproducible/comparable runs.

## Public Repo Scope

Included in this public repository:
- runtime source code (`src/quickthink`)
- reusable evaluation harness (`scripts/eval_harness`, `docs/evals` prompt/spec files)
- tests and release process notes

Excluded from public tracking:
- internal multi-agent comms logs
- generated eval result dumps and ad-hoc local traces
- private experiment workspaces under `experiments-local/`

## Branching

- Keep version tracks isolated in `codex/*` branches.
- Merge to `main` only after benchmarks and notes are updated.
- See `docs/VERSION_NOTES.md` for version-to-version differences.

## Maintainer Commands

Install (editable + dev):
```bash
python -m venv .venv
source .venv/bin/activate
pip install -e '.[dev]'
```

Test:
```bash
PYTHONPATH=src .venv/bin/pytest -q
```

Lint (same commands as CI):
```bash
.venv/bin/ruff check src/ tests/ scripts/
python -m compileall src tests scripts
```

Release docs + checklist:
```bash
make release-check VERSION=x.y.z
```
Follow:
- `docs/release/RELEASE_CHECKLIST.md`
- `docs/release/RELEASE_PROCESS.md`
- `docs/release/SUPPLY_CHAIN_BASELINE_2026.md`

## Limitations / What it does not do

Grounded in how the code actually behaves:

- It does not improve every answer. Whether the planning scaffold helps is model- and task-dependent; run the eval harness before claiming improvements.
- It does not add an LLM of its own. The routing, plan grammar, and validation/repair logic are plain Python; the answer (and, in `two_pass` mode, the plan) still come from your Ollama model, so `two_pass` adds one extra model call versus `direct`.
- It does not verify correctness. When a generated plan fails the grammar check, the engine substitutes a fixed fallback plan (`g:solve;c:constraints;s:direct_reasoning;r:verify_output`); this keeps the format valid but does not make the answer correct.
- It only targets the three pinned models in `SUPPORTED_MODELS` (`qwen2.5:1.5b`, `mistral:7b`, `gemma3:27b`). Other models may run but are untuned.
- It is local-first only: it talks to a local Ollama HTTP endpoint and is not a hosted API, a training framework, or an agent-orchestration platform.
- Hidden planning is still logged for transparency; keep `--log-file` output auditable if you rely on the plan.

## License

Apache-2.0

---

## About Hermes Labs

[Hermes Labs](https://hermes-labs.ai) is an AI reliability engineering studio for product and engineering teams shipping production agents and LLM applications. We find the structural AI failures standard evals miss, then harden retrieval, memory, agents, and the language layers around production AI systems with runtime controls and defensible evidence.

Browse the [open-source catalog](https://hermes-labs.ai/open-source) or contact [roli@hermes-labs.ai](mailto:roli@hermes-labs.ai).
